7,169 Matching Annotations
  1. Aug 2026
    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The factors that create and maintain diversity in host-associated microbiomes remain poorly understood. A better understanding of these factors will help in the efforts to leverage the adaptive potential of the microbiome to help solve pressing problems in health and agriculture.

      Experimental evolution provides a promising path forward as we can track the causes and consequences in the emergence of novel variants, but experimental evolution remains underutilized in host-microbiome interactions. Here, Gracia-Alvira utilizes a long-term experimental evolution study in Drosophila simulans under hot and cold temperature regimes to identify strain-level variation in an important fly bacterium, Lactiplantibacillus plantarum. They identify three strains of L. plantarum, which are most prevalent in their respective three temperature regimes, suggesting that these are locally adapted bacteria. Then, using a combination of genomics, in vitro, and in vivo, Gracia-Alvira et al attempt to understand the factors that led to the differentiation of the hot and cold L. plantarum and their impacts on the fly host.

      Strengths:

      This is an excellent use of experimental evolution to track the emergence of novelty in the microbiome. The genomic analyses are all solid and appropriate for the data sets. It is especially striking that the comparisons with the other, independent experimental evolution studies in different labs (and across continents between Portugal and South Africa) show a consistent response to temperature. Many have disregarded the microbiome as it is something that is too sensitive to seemingly innocuous variables (particularly in the fly microbiome), such that we cannot find generalities. However, this finding highlights the potential for experimental evolution to uncover these dynamics. The question of how strains emerge and are maintained is timely and is one of the key open questions in host-microbiome evolution currently.

      Weaknesses:

      (1) The framing in the title and throughout the discussion about "subspecies competition" does not match the data that was collected. The subspecies competition requires actually tracking the competitive outcomes between the hot, cold, and unevolved L. plantarum. In the in vivo work, I can see that mixes of the strains were made, but they did not track whether the cold strain outcompeted the hot strain in vivo under cold conditions, for example.

      We thank the reviewer for the honest concern and take this opportunity to defend our claim of "subspecies competition used across the manuscript. As the reviewer states, subspecies competition requires tracking the competitive outcomes between the three clades, and this is what we did by sampling and sequencing across ten years of experimental evolution (Figures 4 and S3). For this reason, we point that the subspecies competition assessment comes from the direct observation of changes in relative abundance across the time series, and not from the follow-up experiments in vivo or in vitro.

      While Figure 4 is suggestive that there is ongoing competition in the hot temperature regime, this is not necessarily shown in the cold, which is dominated by the C clade. It could also be that the bacteria cannot survive in the flies at the different temperatures. The growth curve assays hint that the bacteria can grow, but the plate reader couldn't actually maintain the 18 {degree sign}C temperature (line 455). So all of this evidence is very indirect and insufficient to say that strain competition is driving these patterns.

      We thank the reviewer for the alternative hypothesis that could explain the observed subspecies dynamic. We rule out that dominance of clade C in the cold occurs because the other two clades cannot grow in this regime based on three pieces of evidence:

      (1) In the time series, clades H and U decrease, but never disappear (Figures 4 and S3), even showing some peaks of abundance in specific replicate populations (Figure S3).

      (2) We isolated individuals belonging to clade H in the cold-evolved populations, as shown in figure 2. This is a direct evidence that clade H prevails in the cold-evolved populations, although in low abundance.

      (3) We did grow the three taxa in fly food Petri dishes incubated at both temperature regimes, observing growth in all cases.

      We will include the food growth experiment in the revised manuscript as further supporting evidence for growth in both regimes.

      (2) The in vivo results are interesting in that there appears to be a fitness cost of clade C, but the explanation is underdeveloped. I say under-developed because in Figure 4, the cold L. plantarum remains much higher throughout adaptation to the hot temperature regime than the hot L. plantarum in the cold regime. The hot L. plantarum is low abundance throughout the cold regime. I felt like this observation was not explained, but it seems relevant to understanding the strain dynamics.

      We acknowledge that a strong fitness cost of clade C is observed in axenic D. melanogaster. In the native host, D. simulans, with reduced microbiome, we observed delayed development that could even be an advantage depending on the situation, as pointed out by reviewer 3 in the recommendations.

      Even if we assume that flies colonized with clade C are less fit in the experimental evolution, another caveat is whether the flies can actively select for the L. plantarum clade. Under this assumption, a clade that imposes a fitness cost to the fly (clade C) should be selected against over time because the flies colonized by this clade will have less offspring or develop later than the rest. Alternatively, as the microbiome is shared among all the individuals in the population, the host might not be able to “purge” the pernicious clade, and L. plantarum dynamics might be controlled solely by the relative fitness between clades in the given experimental treatment. We will discuss this hypothesis in the revision as a way to explain the relationship between the abundance of each clade and the effect on the host.

      I will also note that this is not the first time that L. plantarum or other Lactobacillus have been shown to exert fitness costs to Drosophila. Gould, PNAS, 2018, shows that both Lactobacillus plantarum and Lactobacillus brevis in mono-association have lower fitness (measured through Leslie matrix projections using lifespan and fecundity) than axenic flies. Many studies of wild Drosophila fail to find Lactobacillus, or it is low abundance (e.g., Chandler, PLoS Genetics, 2014; Wang, Environmental Microbiology Reports, 2018; Henry & Ayroles, Molecular Ecology, 2022; Gale, AEM, 2025). This might help provide useful context for the in vivo results.

      We thank the reviewer for the references. These observations are compared to our phenotypic results and discussed in the revised version of the manuscript.

      (3) The data in Figure 4 are compelling to focus on the L. plantarum variants. However, I can see from the methods that the competitive mapping included only other strains of Wolbachia.

      We appreciate the thorough reading of the methods by the reviewer. The competitive mapping comprised two steps: first we discarded the reads that mapped to Drosophila, Wolbachia and additional potential contaminants from sequencing facitilies (human, dog...). This step leaves the reads originated from whole the external microbiome of the flies, including L. plantarum. The second competitive mapping step recruits the reads that map any clade of L. plantarum.

      It is not clear how other members of the microbiome changed in response to the temperature regimes. As I note in point #2, given that Lactobacillus is often rare, it is not clear what the rest of the microbiome looks like over the course of adaptation. Indeed, it seems like Mazzucco & Schlotterer, PRSB, 2021 did a broader analysis of the microbiome and found that Acetobacter is by far the most common bacterium (I think this data is also part of the data shown here?). Expanding on why or why not in this context is important and will improve this study, particularly if the focus is on connecting these evolutionary dynamics to ecological competition to explain the emergence of strain diversity.

      We acknowledge that the rest of the Drosophila microbiome is not addressed in this study, as we wanted to focus the storyline around the intraspecific dynamics found in L. plantarum. We consider that a complete characterization of the whole Drosophila microbiome would unnecessarily elongate the paper and thus we treat it as a constant biotic factor.

      We must point out that our dataset is not the one reported by Mazzucco & Schlötterer, which was done in D. melanogaster, rather than D. simulans. Nevertheless, both experiments share the same infrastructure, temperature regimes and fly maintenance.

      We have included a list of taxa that were isolated from the populations, as well as to report L. plantarum prevalence and abundance across the experiment in order to provide context of the microbiome, beyond L. plantarum, to the readership.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Gracia-Alvira et al. investigated how environmental temperature affects competition among members of the microbiome, with a focus on intraspecific diversity, using the Drosophila model.

      Notably, the authors identified three clades of Lactiplantibacillus plantarum from a natural population of Drosophila simulans collected in Florida. They tracked the dynamics of these three bacterial clades under two temperature conditions over the course of more than ten years. Using comparative genomics and phylogeny, they showed that these three bacterial clades likely adapted to their host independently in a temperature-specific manner. Further, by combining in vitro culture and in vivo mono-association assays, they demonstrated the functional divergence of these three bacterial clades phenotypically, including their growth dynamics and effects on host fitness. Lastly, they performed pathway analysis and speculated on key genomic variance supporting such functional divergence.

      Strengths:

      The laboratory evolutionary experiment in response to cold or hot environmental temperature is impressive, given its more than ten years of experimental time period. This collection of achieved microbiome samples paired with the fly host data can be a valuable resource for the field.

      Weaknesses:

      The laboratory evolutionary experiment can be limited due to its artificial experimental setup. For example, wild flies rely on a more diverse set of food sources and are constantly exposed to new bacterial inoculations, whereas under laboratory conditions, flies live in a more restricted ecosystem. In addition, environmental temperatures differ among different locations, but they also involve seasonal changes within the same region. This manuscript can be strengthened with further discussions that elaborate on these limitations.

      As the reviewer has correctly noted, our experimental setting is not exempt from limitations. Lab-reared flies are fed with a defined standard diet. Furthermore, although the system is not completely closed to bacterial migration, this is limited as replicate populations are not allowed to mix during the maintenance of the flies. For this reason, we consider our laboratory setting as a compromise between observing wild populations, which undergo all biotic and abiotic stresses but cannot be manipulated, and evolving the bacteria in absence of the host, or in gnobiotic hosts, in which biotic interactions are not fully considered. We will extend on this in the new version of the manuscript.

      Moreover, the extent of host effects involved in these experiments remains ambiguous, because it is unclear whether these Lactiplantibacillus plantarum mostly reside within fly guts or on Drosophila medium. The laboratory evolutionary experiment possibly favored better colonizers on Drosophila medium under either cold or hot temperatures, which subsequently can saturate fly guts. As fully dissociating these variables can be experimentally tedious, the authors may want to comment more on these aspects in the discussion. Or they may want to consider some measurements. For example, measuring the growth rate of these bacteria on Drosophila medium under different temperatures, in addition to the current MRS culture experiments, or measuring the portion of the Lactiplantibacillus on Drosophila medium versus these stably colonizing fly guts.

      The reviewer's point was briefly addressed in the Results chapter: "Phenotypic differences in liquid culture".

      Reviewer #3 (Public review):

      Summary:

      The study presents an analysis of 297 pangenomes derived from 20 populations of Drosophila simulans, at 19 time points for fast-reproducing individuals in a hot environment, or at 10 time points for slow-reproducing individuals in a cold environment, over a period of more than 10 years. The authors select a particular microbial component of the pangenomes and study the dynamics of Lactiplantibacillus plantarum strains in two environments. They discover that the revealed operational taxonomic units could be divided into three phylogenetic clades, which have their own genomic and genetic features, different adaptive capabilities that depend on the environment, and have a distinct impact on the fitness of the host.

      Strengths:

      The authors prove that bacterial microbiome components are sensitive to the environment and could rapidly (years) be fixed in eukaryotic populations. This study establishes a tractable model that potentially enables the study of variability of the physiological influence of distinct strains of an important commensal species, Lactiplantibacillus plantarum, on the Drosophila host. It is clearly shown that this single species consists of several phylogenetically and functionally diverse strains. The authors did not limit their interest to their own model, but rather they have integrated a comparative approach by analysing phylogenetic relationships among 92 described L. plantarum strains.

      Overall, the study is novel and delivers important discoveries of a longitudinal, well replicated experiment, generating a substantial amount of genomic data. It highlights an important dimension of research that environmental selection operates at the subspecies level.

      Weaknesses:

      Even though the authors show only one particular example by conducting their longitudinal experiment, they honestly acknowledge failures important for interpretation of the biological significance of the results (gnotobiotic mono-association experiments was done with D. melanogaster, but not D. simulans) and therefore they state limitations of their conclusions (weaker effects in the non-axenic flies are due to the presence of other taxa or to higher-order interactions with other members of the microbiome). These interactions could significantly affect bacterial growth, metabolism, and physiological influence on the host.

      We agree with the reviewer in that the use gnobiotic animals is a limitation, as by "tuning" the flies' microbiome we are modifying the interactions between members, which can potentially change the phenotypic outcome. Nevertheless, we use it as a complementary approach, rather than the only inference in our study.

      The authors exploit the results of their experiment to speculate about a wide range of evolutionary phenomena, like within-species competition, ecological adaptation and evolution of the host, fitness advantage of bacteria to the host, the benefits of parasitism or mutualism, the domestication of the microbiome, etc. At the end, they conclude that their study "highlights that even subspecies diversity plays a key role in adaptation to environmental temperature". However, the potential mechanisms of such adaptation are barely discussed, so that the focus of the study shifts from the temperature-induced changes in microbial population structures toward metabolism-related adaptations of clade representatives that enable them to diversify their carbon and nitrogen sources. The role of the temperature factor remains elusive.

      We acknowledge that our study does not fully resolve the mechanism by which a different clade ends up dominating each temperature regime. The MRS liquid experiment was an attempt to answer whether differences in optimal growth temperature could explain the temperature-specific abundance of the two clades. Our experiments showed, however, that this was not the case. Beyond this point, it is hard to disentangle the role of the temperature, as it could also act indirectly on the bacteria, for example, through the host or the food.

      A second observation in our time series was that a third clade, U, was unfit in both regimes despite starting the experiment in high abundance. For this reason we also studied what made this clade less fit. Based on our analyses, we propose that the decrease of clade U was driven by the shift to a laboratory diet, shared by all experimental populations.

      In addition to that, the paper has a clearly minimalistic experimental approach to address functional properties of the revealed L. plantarum strains, so that their own fitness, or their relationship with the Drosophila host, is characterised superficially. Therefore, the authors' discourse can be speculative rather than factual (especially when the authors use the expression "likely" to share their guesses in the "Results" section). Nevertheless, these minor drawbacks do not underscore the novelty of the discovered phenotypes and the importance of their further investigation.

      We consider the reviewer's concern and toned down the phrasing when reporting our findings in the revised version of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) One solution to resolve the "competition" issue would be to check that the L. plantarum strains are established at similar or different titers in the in vivo work. Fly phenotypes can be sensitive to microbial load (Keebaugh, iScience, 2018), which might explain some of the counterintuitive in vivo results. In line 227, the authors mention that "bacterial load" is contributing to the magnitude of the effect, but I don't see the data reported anywhere. If this is from the in vitro assays, then the authors need to show that in vitro predicts in vivo L. plantarum abundance.

      Bacterial load inoculated in the in vivo experiment was normalized to OD=0.05 (~5*10<sup>6</sup> CFUs/ml) for the three clades at the beginning of the experiment. Thus, all vials were inoculated with the same titer of L. plantarum. Only the genotype varied between treatments. However, it is possible that, once inoculated, each clade grew to a different titer (as they have different growth rate and carrying capacity).

      Statement in line 227 comes from the differing results in transfers 1 and 2. In transfer 1 we inoculated a fixed load of ~2.5*10<sup>5</sup> CFUs. In transfer 2, however, we inoculated no bacteria to the food, and the flies seeded the vial. Our statement comes from the assumption that bacterial load in transfer 2 has to be lower than in transfer 1 as bacteria seed the vial solely by defecation of the parents.

      Following the reviewer's suggestion, in the revised version of the manuscript we have included a new experiment in which we quantified the bacterial load of each clade in individual flies.

      (2) Tracking the competitive outcomes is tricky, though it could be done with whole genome sequencing. An alternative would be to label the strains with fluorescent proteins (e.g., Obadia Current Biology 2018 has done this in Lactobacillus) and track fluorescence to better understand the results of the "mix" treatment in Figure 6.

      We appreciate the feedback of the reviewer, but consider this rather labour-intensive approach as an interesting option for future follow-up experiments.

      (3) That being said, my main concern with this is the "competition" claim. If the paper were reframed appropriately, this paper could still make an important contribution to the evolution of host-microbiome interactions, but the authors would need to consider what they can and cannot do with this interesting dataset.

      The "competition" claim comes from the changes in relative abundance observed in the time-series data, not from any of the follow-up experiments. Thus, we consider the use of the term "competition" appropriate.

      (4) The text on the figures is very small and hard to read.

      We increased the size of the text in all figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Have you conducted the in vitro culture experiments following the "cold" conditions?

      We have conducted the experiment in "cold" conditions, but with some modifications to the experimental settings, as the plate reader did not have cooling capacity. Instead, we grew a subset of the isolates (four per clade) in glass vials at constant 20 °C, and measured their OD twice a day. We have included the results in the revised version of the manuscript.

      (2) How many technical and biological replicates were measured for the in vitro culture experiments (Figure 5)? Please add this information to the figure legend and method.

      We measured the growth of four isolates from clade U, nine isolates from clade C and sixteen isolates from clade H. Each isolate was grown three times.

      We have included this information, as requested by the reviewer.

      (3) Making the labels in Figures 2, 3, 5, and 6 bigger would be helpful.

      We have increased the font size of all figures.

      Reviewer #3 (Recommendations for the authors):

      (1) Line 268: "Based on our results in experimentally evolved fruit flies, we propose that within-species competition, thus far largely overlooked, could contribute to ecological adaptation and evolution of the host". Overstatement should be avoided, since the evolution of the host was not directly studied here.

      Our results show that reproductive traits of the host differ upon colonization with each clade. Although we don't test the host's evolution, we speculate that flies differing in their offspring number and developmental time might differ in their overall fitness. Finally, we consider the Discussion section as the right place for speculation and development of hypotheses that can be tested in future work.

      (2) Line 258: "These differences do not explain the clade-specific selection, but reflect the different evolutionary histories of the clades". The temperature factor and its possible role in clade selection would be better discussed at least a little bit.

      In this paragraph we described potential metabolic differences between clades using comparative genomics. We did not find enrichment in a function or group of functions that could explain the different dynamics between clades H and C in the temperature regime.

      In the revised version of the manuscript we highlight that we did not find temperature-specific differences from this analysis.

      (3) Line 252: "...This could explain why clade U, which displayed a high growth rate and carrying capacity in liquid culture". The statement could be further developed with a caution. Even if the isolates that belong to the clade U are outcompeted by H or C, it should be noted that the strain U cannot be used as a true reference for fitness, since it could possess its hidden adaptive properties, not being simply "a loser". Such a hypothesis could explain the maintenance of this strain in the wild.

      We agree with the reviewer in that fitness is relative to the selective environment. Clade U is less fit than H and C in our specific experimental conditions, but it could outcompete them in other conditions, such as wild flies or MRS liquid medium. In the revised version of the manuscript we have rephrased this statement to clarify that we specifically refer to clade U's fitness under the new laboratory conditions.

      (4) In a cold environment, association with the clade C induces developmental delay and produces less progeny, which potentially allows the host to survive in case of harsh conditions and potential food limitation. Could the authors speculate and not exclude that this phenotype could be potentially adaptive? It would be curious to check in further studies whether flies associated with C strains are more stress-resistant, for example.

      We thank the reviewer for this alternative hypothesis. In our manuscript we used the Darwinian definition of fitness; reproductive success of an organism in the focal environment. And thus, both higher progeny per female and shorter developmental time would be beneficial in direct competition with other individuals. It is true that delayed developmental time, or less progeny, could be advantageous in specific cases. This could be the case for D. simulans inoculated with clade C. However, we consider that the fecundity levels observed in D. melanogaster upon inoculation with Clade C (average of 0.06 offspring/female*day in the cold) are too low to sustain a population.

      We have included this hypothesis in the Results section.

      (5) It would be highly recommended to add an experiment to complete the story by measuring the quantity of bacteria in the medium and in the flies. This will resolve the hypothesis (Line 785): "Thus, the ability to exploit this ubiquitous source of carbon and nitrogen could be very advantageous in the fly microbiome context, but would not affect the fitness in liquid culture".

      Following the reviewer's recommendation we included two additional experiments. We measured the bacterial load per fly in the native host, D. simulans, inoculated with the three clades. We also compared the clades' growth speed in solid fly food (without host). In the former experiment, we found similar bacterial loads upon inoculation with clades U and clade H. In contrast, in the latter we found delayed growth of clade U relative to H and C in the food. Thus, chitobiose consumption does not seem to provide an advantage in the fly gut to clade H. We attribute the fitness advantage of H and C to their advantage growing on the laboratory fly food, regardless of the host.

      Both experimental results have been included in the revised version of the manuscript, and the comparative genomics paragraph and discussion have been modified in consequence.

      (6) The chapter "Extended clade-specific differences in KEGG metabolic pathways" could be presented in the main text as it contains important results. These results are mentioned in the chapter "Functional divergence on the genomic level", which looks rather humble when it stands alone as it currently does.

      We appreciate the interest of the reviewer in this supplementary chapter. To keep the length of the manuscript digestible for a broad set of readers, we decided to only include in the main text the functional differences that could play a role in adaptation to the new laboratory environment.

      We consider that a full description of the metabolic differences between the three clades has to be published, as it might be relevant for researchers interested in L. plantarum metabolism. However, it does not fully follow the storyline, as the differences reported in the supplementary, such as nitrate respiration or synthesis of molybdenum cofactors, might not be involved in the clade-specific selection observed in the time series.

      (7) Line 773: "Clades C and H encode a shared genetic repertoire related to sugar/riboflavin metabolism that is lacking in clade U". This indeed allows us to hypothesise that the fixation of these clades in fly populations was due to their improved metabolic capabilities. However, the analysis of fitness shows similarity in flies associated with clades H and U, meaning that sugar/riboflavin metabolism in H does not provide an obvious adaptive trait to flies. Moreover, one could say that sugar metabolism in clade C is maladaptive not only for flies, but also for bacteria in liquid cultures. It is recommended to more clearly state the respective limitations of the study.

      Here we have to make a distinction between bacterial fitness and host fitness. The three clades differ in their (bacterial) relative fitness, as evidenced by the time-series dynamics (Figure 4). In the cited statement we hypothesize that a more versatile sugar metabolism repertoire could increase the bacterial fitness of clades H and C (relative to clade U) in the sugar-rich laboratory diet.

      This is independent of the fitness effect that L. plantarum could have in the host. Finally, as it was discussed in the recommendation 3, fitness is specific to the environment. Clade C is the least fit in liquid MRS in hot conditions, but the fittest in cold experimental conditions.

      (8) The authors should better explain why growth in MRS was not performed in a cold temperature regime to further support or refute the hypothesis that capacity and inflection time could partially explain the higher fitness of bacterial strains from clade U.

      We did not perform this experiment in cold conditions due to technical limitations of the plate reader, that does not have cooling capacity. Nevertheless, following the reviewers' suggestion, we have included in the revised version of the manuscript a new MRS growth experiment in cold-like conditions (constant 20 °C).

      (9) When mentioning that L. plantarum can "increase larval fitness of Drosophila melanogaster relative to germ-free flies" (line 196), the authors should specify in which specific conditions this phenotype was observed, and how relevant the mentioned phenotypes are to the current study.

      Following the reviewer's recommendation, we have modified the paragraph in order to clarify the conditions used in other papers and those used in our work. The references cited in this section (PMID: 21907145, 29290388 and 28062579) report that L. plantarum increases the host fitness in protein-poor diets (12 g/l of dried yeast or less), but not in high-protein diet (50 g/l of yeast or higher). Since our experimental diet contains an intermediate amount of protein (24.3 g/l of dried yeast) we were agnostic of whether L. plantarum would benefit the host or not in our conditions. Regarding the phenotypes, we chose two reproductive traits that are affected by changes in the microbiome according to the literature. Developmental time is directly affected by L. plantarum in the aforementioned papers. Offspring number is another fitness component affected by Drosophila microbiome (PMID: 30510004).

      (10) Provide a reference for line 205: "In axenic D. melanogaster none of the L. plantarum clades provided a fitness advantage to the host relative to germ-free controls, contrary to the effects reported in the literature". If the conditions were different from those in the studies referred to, then it would be of no use to compare the fitness advantage (for example, in Reference 24 another type of diet was used).

      Already covered in recommendation 9.

      (11) Please provide more context to this statement (Line 210): "The high content of dried yeast 24.3 g/l in the fly food used in our experiment likely provided already sufficient amounts of essential amino acids, which negated the growth-promoting effects of L. plantarum". It is not clear why amino acids are taken into account, and what the evidence is for the fact that the amount of essential amino acids was sufficient to abolish growth-promoting effects.

      The whole paragraph was modified in order to clarify the relationship between protein input and nutritional fitness benefit of L. plantarum.

      (12) Please provide measurements of bacterial quantity which would support the statement (Line 215): "The fitness reduction was stronger in the first transfer of flies, likely due to a higher bacterial load".

      Upon request of the reviewer, we have estimated the bacterial load per individual fly in D. simulans. Additionally, we have specified the CFUs inoculated in the vials in transfer 1.

      (13) Correct the typo (line 220): "However, the developmental time was significantly extended after inoculation with clade C at cold temperature (Dunn's test, p < 0.05 05 for all significant comparisons)".

      Done.

      (14) Specify more precisely the temperature conditions referred to in line 226: "In summary, we observed that clade C, which is dominant in the cold-evolved populations, decreases host fitness when axenic flies are inoculated". Does it decrease fitness both in hot and cold environments?

      For the axenic flies, we did find a decrease in fitness in both regimes, yes. We specified it in the revised version of the manuscript.

      (15) Please provide evidence for line 227, or otherwise rephrase it: "The magnitude of this effect varies depending on the environmental temperature, the bacterial load, and the presence of other microbial taxa".

      Novel evidence was provided regarding the role of bacterial load on host fitness.

      (16) Correct the following statement, so that it reproduces the results of the original work (reference 19, line 229): "In a low-protein diet, strains that were not isolated from Drosophila enhanced larval growth relative to germ-free individuals, whereas another Drosophila-associated strain did not have any effect".

      This statement was removed from the revised version. This reference was cited in the discussion to state that: " the nutritional symbiosis in L. plantarum is strain-specific".

      (17) Please provide a rationale for using KEGG Orthologs. Why was this database chosen as an appropriate one, even though it is known to be a non-exhaustive metabolomic resource?

      KEGG is a well-known metabolic database that is widely used in comparative genomics (PMID: 40177264) and built in state-of-the-art software for microbial ecology such as Anvi'o (PMID: 33349678). Other similar gene-to-function databases are less focused on metabolic pathways, such as COG or GO, or limited to specific enzymatic activities, like CAZy. Furthermore, the hierarchical organization of KEGG Orthologs in modules and pathways allowed us to map clade-specific orthologs to the broad metabolic context. For these reasons, we considered KEGG to be the best option for this analysis.

      (18) Line 250: "Therefore, we speculate that the ability to exploit this ubiquitous source of carbon and nitrogen in the lab-maintained fruit flies, could be a strong target of selection in the lab environment". This statement concludes the "Results" section but would be more appropriate for the Discussion section, since the authors do not provide any experimental evidence that could support this statement.

      We have modified this chapter, as covered in recommendation 5.

      (19) Line 276: "However, the intraspecific richness of L.plantarum in our flies was three times higher than that estimated in human gut microbiomes". Note that there are other recent studies which show the presence of several OTUs within L.plantarum isolates (for example PMID: 41484402).

      We thank the reviewer for the reference. We comment on it in the revised manuscript.

      (20) Line 287: "Our finding shows that the well-characterized nutritional symbiosis between Drosophila and L. plantarum depends on the bacterial genotype and cannot be generalized to the entire species". Note that such a conclusion has already been previously stated (for example, PMID: 30008290 and 28993620).

      We thank the reviewer for the references. Indeed, these papers show that some L. plantarum strains are beneficial for the host while others are neutral. Furthermore, as commented by Reviewer #1 in the public review, L. plantarum has been shown to reduce the host's fitness relative to axenic flies (Gould, PNAS, 2018).

      Our observations are novel in two ways. (1) The fecundity observed in D. melanogaster, 0.06 offspring/female/day in average, is lethal (in Gould et al. 2018 fecundity never decreased below 1 offspring/female/day). (2) Clade C outcompetes the other clades in the cold, despite being detrimental for the host.

      We have modified the Discussion to account for the previous work.

      (21) Line 343: "In addition, we obtained L. plantarum genomes from two other experimental evolution studies. Two genomes from the South African experiment and seven genomes from the Portugal experiment". Merge two sentences into one.

      Done.

      (22) Line 360: "At sampling, the age of the flies varied between four and eight days for the hot environment and between nine and 16 days for the cold environment". Please comment on the fact that different age of flies (different physiology) is not the reason for bacterial community differences.

      During maintenance, flies are sampled at different ages because the temperature affects their developmental time. We cannot rule out the hypothesis that age difference drives microbiome differences. Temperature could affect clade competition directly (differences in optimal temperature between clades) or indirectly, by affecting either the host (e.g. changes in Drosophila developmental time alters L. plantarum fitness), the surounding microbiome, or the food (e.g. increased metabolic activity in the hot regime changes nutrients profile). We ruled out the direct effect of temperature with growth experiments in liquid MRS medium and solid fly food, but disentangling the indirect effects is not feasible.

      (23) The majority of figures have low-quality labels that are not legible due to the small size of the font. Please improve.

      Done.

      (24) Figure 1 - Correct the legend: There is no "10" label on the picture. Probably by 10, the authors mean "Generation", while by x10 - number of isogenic replicates.

      Done.

      (25) Figure 2 - No numbers at nodes are indicated, whereas it is announced in the legend that they represent bootstrap support values. In addition, it is recommended to show a reference pangenome in the middle panel to clearly refer to the total size of the possible black bar.

      We added high bootstrap support as coloured nodes in figures 2, 3 and S2.

      We do not understand the reference pangenome request. In the middle panel, each black/white bar corresponds to an orthologous gene that can be either present or absent in each of the genomes. These orthologs were sorted based on hierarchical clustering of the their patterns of abundance (columns present in the same set of genomes, together), not by synteny. Thus, a reference pangenome would be simply a black bar.

      (26) Figure 2: It would be advantageous to add a figure that represents the frequency of each strain in each replicate (at the last time point, for instance). It would explain why some “blue” strains appear to be within the “red” cluster. Otherwise, it is confusing to find cold-evolved bacterial strains in hot-evolved fly populations.

      The frequency of each clade in each replicate is shown in figure 4. We think that it would be more confusing to follow the suggestion of the reviewer, as the isolates were sampled at different time points of the experiment. We would not like to call it a confusion that "blue" strains appear in the "red" cluster, but rather the logical consequence of the color code used in figure 2, which corresponds to the temperature regime in which the isolate was sampled (regardless of its clade). Whereas in the following figures colour represents the clade. It is thus possible to find clade H isolates in the cold temperature regime, as this clade is in low frequency but not completely absent in this regime.

      (27) Figure 3 – Add a label for the X-axis.

      We rotated the tree to be able to increase the genome IDs. We have added the label to the Y-axis.

      (28) Figure 4 - Please indicate how the clade relative abundance was assessed.

      Clade relative abundance was inferred by mapping competitively the short reads against the three clades’ reference sequences. It is specified in the legend now.

      (29) Figure 6 - Total number of F1 flies eclosed normalised by day (during which period?). What do T1 and T2 correspond to?

      During the respective number of days that females were allowed to lay eggs: one day in the hot settings and two days in the cold settings in transfer 1. One day and three days, respectively, in transfer two.

      T1 and T2 correspond to the first and second transfers, as described in the Materials and Methods. In first transfer, flies laid eggs in vials pre-inoculated with a set load of L. plantarum. After egg laying, same adults were then transferred to a sterile set of vials and allowed to lay eggs again (second transfer). Bacterial load in these vials was solely seeded by the parents.

      In order to avoid any confusion, in the revised version of the manuscript we have modified figure 6 to show transfer 1 for both Drosophila species, and moved transfer 2 dataset to supplementary figure S5.

      (30) Figure S2 - label the X-axis.

      We guess the reviewer means Y-axis. Done.

      (31) Figure S3 demonstrates the real data and its variability, so it would be better used instead of Figure 4 (which seems to be just a derivative from Figure S3, not a separate dataset and separate type of analysis).

      As the reviewer suggested, we have replaced figure 4 with figure S3.

      (32) Figure S4: Improve plot title: (e.g., C:H:U = 3:3:3).

      Done.

      (33) Figure S6: It is stated that N = 10; however, some datasets do not have 10 points represented. Please specify why. Also, please specify the meaning of "T1/T2".

      For the inoculation experiment in Drosophila simulans, we had nine replicates per treatment, not ten. This has been corrected in the figure and in the Materials and Methods section.

      T1 and T2 correspond to the first and second transfers, already covered in recommendation 29.

      (34) Table S3: provide legend for values (1 - present in all strains, but 0.04 - what does it mean?).

      It means that 4% of the genomes from this clade harbour the specific gene. We have specified it in the legend of the revised table.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Uphoff et al. propose a structural and mechanistic model in which the multidomain ECM protein SVEP1 enables Angiopoietin (ANG) binding to the orphan receptor TIE1, thereby promoting downstream receptor phosphorylation and signaling. Using AlphaFold-based modeling, the authors predict that the CCP20 domain of SVEP1 binds to TIE1, creating a composite surface that facilitates Angiopoietin association and TIE1 activation. The resulting ternary model (SVEP1-TIE1-ANG) offers a structural rationale for how SVEP1 converts TIE1 into a functional, ligand responsive receptor. Additional models and biological assays suggest roles for other domains of SVEP1, such as CCP5-EGF-L7, although these interactions are predicted with low confidence. The authors interpret these findings as the first structural framework for how SVEP1 enables ANG-TIE1 signaling.

      Strengths:

      (1) The central hypothesis - that SVEP1 enables ANG binding to the orphan receptor TIE1 - is biologically compelling and addresses an important question in vascular biology.

      (2) The AlphaFold-predicted ternary complex (SVEP1-TIE1-ANG) is plausible, high-confidence, and structurally consistent with prior functional data (e.g., poly-Ala scanning from Sato-Nishiuchi et al.).

      (3) The authors' model offers a potential explanation for the previously observed role of SVEP1 in enhancing ANG signaling through TIE1, and may represent the first structural insight into TIE1's transition from orphan to ligand-activated receptor.

      (4) The potential clinical implication - that a combinatorial ligand (ANG+SVEP1) can activate TIE1- could have translational relevance for vascular leak and inflammatory disease.

      Weaknesses:

      (1) Lack of structural validation and mechanistic follow-up: Despite the promising AlphaFold model, there are no figures of the predicted interface, no residue-level interactions shown, no ipTM values reported, and no experimental follow-up to test the interface. PAE plots are incorrectly used as confidence justifications, which is not appropriate for complex predictions.

      We have appended the data showing AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity. We also added ipTM scores and confidence plots for the predicted complexes.

      (2) Biophysical validation is missing: No surface plasmon resonance (SPR), ITC, or biochemical assays are included to confirm ternary complex formation or quantify binding kinetics. Given the manuscript's structural focus, this is a major gap. For instance, an SPR experiment where ANG is immobilized, and TIE1 binding is measured {plus minus} SVEP1, would directly test the model. And allow direct comparison to ANG-TIE2.

      We have addressed this question and performed ELISA assays to measure binding affinities between SVEP1 and TIE1 in presence or absence of ANG1 or ANG2, thus confirming that the affinity is increased in the presence of ANG1 or ANG2.

      (3) Missed opportunity for mutagenesis-driven validation: The manuscript does not include any interface-targeted mutations, despite clear opportunities. For example, mutating T2595 in SVEP1 (to R) or mutating the TIE1-specific residues (residues PL 202-203 to LF) could strongly test the model and potentially reveal dominant-negative behaviors. E.g. A T2595 mutant should block ANG binding but not TIE1 binding.

      We have depicted figures of the interfaces including P202-L203 and included the TIE1 P202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study, for the following reason: Modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2, thus reducing its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both TIE1 and ANG2, the data is not included in the manuscript.

      (4) Overinterpretation of weak models: The additional AlphaFold model involving the CCP5-EGFL7 domains binding TIE1 has extremely low confidence (ipTM < 0.15) when reexamined by this reader and should not be emphasized. There is no biophysical evidence or binding data (SPR) to support this interaction, and its inclusion detracts from the much stronger CCP20 model.

      We agree with this point made by both reviewers and have removed the data on CCP5-EGFL7 from the manuscript.

      (5) Language around modeling is overstated and potentially misleading: Terms like "unequivocal," "high-affinity," or "affirms strong binding" in reference to AlphaFold predictions are inappropriate. These are hypotheses -not confirmations - and must be tested at the biochemical level. This should be clarified throughout the manuscript to ensure non-experts do not misinterpret modeling confidence as binding affinity.

      We agree with the reviewer, and have adjusted the wording.

      (6) Negative stain EM data is not informative due to low resolution and lack of defined interfaces; unless replaced by higher-resolution Cryo-EM, this should be omitted. Better would be co-gel filtration, AUC, or SEC-MALLs with ANG-SVEP1-TIE1.

      We have now appended the data by adding gold-labelled TIE/ANG proteins, thus enhancing clarity.

      (7) Disjointed narrative: The manuscript presents a compelling mechanism involving CCP20-driven ANG binding to TIE1, but then becomes fragmented by introducing the low-confidence CCP5-EGFL7 model and speculative higher-order polymerization models that are not experimentally supported.

      We agree with this point and have have centered the manuscript around CCP20. We removed data concerning CCP5-EGFL7 as suggested by both reviewers.

      Reviewer #2 (Public review):

      Uphoff and colleagues present the results of a study focused on characterizing the binding of SVEP1 to TIE1 along with Angiopoietin-2. Starting with computational prediction of SVEP1 binding to TIE1, the authors identify the region of SVEP1 that serves as a high-affinity ligand for TIE1. Advanced studies identify a weak secondary binding site within SVEP1 that appears to be sufficient but not necessary for its interaction with TIE1 based on in vivo rescue experiments. The most novel contribution of the manuscript seems to be the identification of angiopoietin-1 and -2 as co-factors that seem to enhance the binding of SVEP1 with TIE1 and impact downstream AKT signaling. They propose a complex in which SVEP1 binds to TIE1 and ANG2.

      Although the first set of results is essentially confirmatory, the identification of ANG-2 as a "cofactor" enhancing the binding of SVEP1 to TIE1 and associated downstream signaling (i.e., Figures 3 and 4) is novel and is of interest. However, the manuscript and its conclusions would greatly benefit from some clarifying details and additional experiments to ensure rigor and support specific claims.

      We have addressed the reviewers concerns and significantly appended the manuscript. Most importantly, we provide structural validation of AlphaFold models reporting interfaces, residue-level interactions and ipTM values. We have included new biophysical validation of binding kinetics of SVEP1 and TIE1 in the presence or absence of ANG1 or ANG2. Furthermore, we have removed the data on CCP5-EGFL7 from the manuscript in order to retain focus on the CCP20 domain.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The AlphaFold modeling for the CCP20-based interactions is strong (as determined by this reader rerunning and getting ipTM values and visually inspecting the interactions “because this is not in the manuscript”). As presented, the manuscript stops at the hypothesis-generation stage. Validation is needed to fulfill the paper's title and claims. The structure-function link is not demonstrated, despite an obvious and achievable experimental path (mutagenesis, SPR, kinetics).

      Additional Context and Suggestions

      (1) Show and label AlphaFold-predicted interfaces, including residues, hydrogen bonds, and surface complementarity.

      We have appended the data in the new supplementary figures 1.1, 1.2, 2.1 and 2.2

      (2) Provide ipTM scores and confidence plots for each predicted complex.

      We have added the values and plots in the new supplementary figures.

      (3) Perform SPR assays with ANG-coated surfaces and measure binding of TIE1 {plus minus} SVEP1. Compare to TIE2 binding for context.

      We performed the proposed experiment using an ELISA assay to measure binding affinities between SVEP-1 and TIE1 in presence or absence of ANG2 and included these data in the manuscript in figure 2.

      (4) Test interface mutants: e.g., T2595R in SVEP1 (should impair ANG recruitment but not TIE1 binding), or PL→LF muta on in TIE1 (should disrupt SVEP1 binding).

      We have depicted figures of the interfaces including P202-L203 and included the TIEP202L L203F mutant, as well as the previously described SVEP1 (E2568A - G2569A) mutant in our experimental data. The T2595 mutant was not included in the current study. Our modeling suggested that replacing T2595 with an Arg will cause steric and charge clashing with 469GKL471 of ANG1 and 467NKFN470 of ANG2 thus reduce its binding to ANG2 although T2595 does not interact with ANG1/2. A SVEP1 protein comprising CCP15 to the C-terminus with the T2594R mutation shows reduced binding to ANG2, but also reduced binding to TIE1. As the mutation hinders interaction with both, TIE1 and ANG2, the data is not included in the manuscript.

      (5) Clarify in the Introduction that SVEP1 is a large, multidomain ECM protein to help readers contextualize the domain names early on. Do not use terms like CCP before defining them.

      We added: “Svep1 encodes a 3571 amino acid long extracellular matrix protein containing different domains such as Willebrand factor type A domain (vWF), ephrin-receptor like domains, complex control protein (CCP) domains, and Hyalin repeats at the N-terminus. The C-terminus mainly consists of CCP and EGF domains. Svep1 is expressed in mesenchymal cells, but not in endothelial cells, and functions non-cell-autonomously (Karpanen et al. 2017; Morooka et al. 2017)”.

      (6) Replace or remove negative-stain EM unless higher-resolution cryo-EM data are available.

      We would like to retain the EM data, but have now replaced the negative-stain EM by adding gold-labelled TIE/ANG Proteins to verify that the proteins we show are the ones we expect. The reason we would like to retain the data is that TIE1 has been considered for so many years as an orphan receptor, and thus we consider it appropriate to demonstrate SVEP1/TIE1 binding using multiple methods.

      (7) Reframe claims around modeling to avoid overstatement. For example: "The model suggests a plausible mechanism consistent with prior biochemical data" is more appropriate than "unequivocal".

      We agree with the reviewer, and have adjusted the wording.

      (8) Consider narrowing the focus: the CCP20-TIE1-ANG model is a strong story on its own. The CCP5EGFL7 model and polymerization hypotheses are not essential and may dilute the impact.

      Since this point was made by more than one reviewer, we have removed the data on CCP5EGFL7 from the manuscript.

      (9) Properly define TIE1: Tyrosine kinase with Ig and EGF domains.

      We have corrected the full protein name for TIE1 and added: “TIE1 and Tie2 exhibit a high degree of homology with a globular head domain consisting of three immunoglobulin-like (Ig) domains and three epidermal growth factor-like (EGF) modules and a short stalk formed by three fibronectin type III repeats, while the N-terminal two Ig domains of Tie2 harbor the angiopoietin binding site (Macdonald et al. 2006).” We also added: “D1 and D2 refer to the two N-terminal Ig domains, D3 refers to the three EGF domains and D4 to the third Ig domain of TIE1 or TIE2.”

      (10) Use RTK, not tyrosine kinase receptors TKR.

      Tyrosine Kinase receptor was replaced by receptor tyrosine kinases (RTKs)

      (11) PDBs (.cif and .json files) of the models must be supplied for the readers so they don't need to rerun the AlphaFold jobs.

      We are providing all PDBs with the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) In several locations, the authors state that alphafold detects an "unequivocal" and/or "high affinity" interaction. Experts on computational structure prediction can weigh in, but I am not sure it is accurate to say that alphafold predicts affinity. Quantitative estimates of the prediction confidence or other parameters of the alphafold output are not provided.

      Thank you for the comment, we agree with the reviewer and have adjusted the wording. We have also appended the data showing interfaces, ipTM scores and confidence plots for the predicted complexes.

      (2) In Figure 1C, what concentrations of TIE1 and TIE2 are being used in the SPR experiments shown?

      We have added the concentrations of the proteins to the methods section

      (3) In Figure 1C, what is the affinity constant (KD) of the interaction between SVEP1 and TIE1 and SVEP1 and TIE2?

      We have added the values of affinity constants to the manuscript text.

      (4) In Figure 1C, the authors immobilize a "70kD" fragment of human SVEP1 to determine the interaction between SVEP1 and TIE1. How does the affinity they measure between the SVEP1 fragment and TIE1 compare to the affinity between immobilized full-length SVEP1 with TIE1?

      The largest molecule we used in any assay is not full-length SVEP1, but consisted of a C-terminal SVEP1 protein spanning from the first EGF domain to the C-terminus (approximately 295 kDa) as previous studies have shown that SVEP1 is proteolytically cleaved N-terminal to the first EGF domain, generating a protein of this size. In our ELISA assays, this molecule has a higher affinity to TIE1 than the 70 kD fragment. We have included data for the 70 kDa as well as for the 295 kDa fragment in the manuscript. ELISA assays have used larger SVEP1 fragments (as indicated in the figure legends), the SPR assay was performed with the 70 kDa fragment.

      (5) How does the alphafold structure prediction for the "70kD" fragment of hSVEP1 compare to the prediction of the same 70kD fragment from the full-length protein prediction?

      All modellings using different SVEP1 fragments including the 70 kD and full-length version identify the putative binding site at position CCP20. While AlphaFold3 can predict the correct domain folds in both the 70kDa fragment and full-length SVEP1, the orientation of these domains is highly variable due to flexible linkers between each domain module. This flexibility effects the output confidence metrics, thereby hampering our interpretation of the models. Therefore, we conducted the structural prediction of the complexes with smaller fragments and not the full-length SVEP1.

      (6) What are the amino acids for the 70kD fragment?

      The relevant amino acids are 2261-2890. The accession number and amino acids of each protein have been listed in the key resources table.

      (7) In Figure 1D, what is being depicted by the red stars? This is not explained in the text or figure legend.

      We have replaced figure 1D.

      (8) By itself, Figure 1D is not terribly informative and in my opinion does not support the statement that the authors "were able to directly visulalize the attachment of SVEP1 and TIE1." As a minimum, the authors should repeat the same set of images with SVEP1 and TIE2, but other approaches, such as labeling, could be performed.

      We replaced figure 1D with new data and gold-coated protein enhancing clarity. We think it beneficial to demonstrate SVEP1/TIE1 binding using multiple methods as TIE1 has been considered as an orphan receptor for so many years. We have not performed these experiments with TIE2 proteins as we were not able to show binding of TIE2 to SVEP1 with other assays.

      (9) What is being stained in Figure 1D? Full-length SVEP1/TIE1? Or fragments of these proteins?

      We replaced figure 1D by a new assay with labeled proteins using the 150 kDa version of SVEP1 and the ectodomain of TIE1 as well as ANG1 or ANG2 (new figure 1D and new supplementary figure 2.3). TIE1/ANG proteins were gold-labelled. The protein fragments used in this assay are described in detail in the methods section.

      (10) The authors discover CCP6-EGFL7 as a low-affinity binding region of SVEP1 for TIE1. Is this region in physical proximity to CCP20 (the high-affinity binding region for TIE1) based on alphafold prediction? How would the authors think this is binding TIE1?

      We have removed this data set (see comment to reviewer 1’s request).

      (11) What is the affinity constant (KD) for CCP6-EGFL7 with TIE1?

      We have removed this data set (see comment to reviewer’s 1 request).

      (12) The authors claim that ANG1/ANG2 increase affinity between SVEP1 and TIE1 based on immunoblotting. Immunoblots are semi-quantitative at best. If the claim is higher affinity, I think the authors should measure this by SPR and determine the KD between immobilized SVEP1 with TIE1 in the absence and presence of ANG1 and/or ANG2.

      We conducted ELISA assays (figure 2) showing that the affinity is increased in the presence of ANG1 or ANG2 and agree with this reviewer that this strengthens the data.

      (13) In Figure 3, can the authors explain why ANG1/2 does not pull down with SVEP1/TIE1?

      We noticed that upon transfection of TIE1 into HEK cells, ANG1/2 is almost undetectable anymore in the total lysate. Thus, we believe that after the pull down we are below the detection limit.

      (14) In Figure 3, what is "TL"? I assume total lysate, but this is not specified.

      Thank you, we now specify TL as total lysate.

      (15) In Figure 3 "TL" panel (again, I assume this is total lysate), why are the ANG1/2 immunoblots so weak when co-transfected with TIE1?

      We consider it likely that in the presence of TIE1, ANG1/2 proteins are internalized and digested. Another reason for low signals could be that upon transfection of two plasmids, the amount of ANG1/2 protein is reduced as the cell has limited capacity for transcription and translation.

      (16) In Figure 3B, why is the SVEP1 fragment now 150kD when 70kD fragment was previously used? What domains are contained in this 150kD fragment?

      We now better define the domains of the 150kD SVEP1 fragment. The 150 kD fragment was the one produced first and available in high amounts in our laboratory and thus used for functional assays. The 70 kDa fragment together with ANG2 also induces phosphorylation of AKT, but it was not used in as many conditions/replicates as the amounts we had available were lower.

      (17) In Figure 4A, signaling with SVEP1 by itself and ANG2 by itself should be shown to support the claims being made.

      We added the lines for SVEP1 and ANG2, and also the quantification. SVEP1 itself already affects the phosphorylation of TIE1, most likely because hdLECs produce ANG2 by themselves. We show this with the ANG2 blocking antibody for pAKT.

      (18) For pAKT, what are the concentations of proteins being used and the times of incubation?

      This information is provided in the Materials and Methods section. We added the concentration of the anti-ANG2 antibody, which was missing.

      (19) It seems that p-AKT and AKT are being blotted on different gels. If this is correct, loading controls need to be shown for p-AKT blot.

      We added HSC70 as a loading control for both blots.

      (20) It is interesting that anti-ANG2 antibody inhibits SVEP1-induced p-AKT signaling. As the authors may know, SVEP1 has been identified as a receptor for PEAR1, which also leads to downstream p-AKT signaling, which seems to be independent of ANG2. Do the LECs being used here express PEAR1? If these cells express PEAR1, how do the authors think ANG2 silencing will eliminate SVEP1-associated p-AKT signaling?

      hdLECs express PEAR1. However, we show that p-AKT signaling is attenuated after siRNA KO of TIE1. Thus, the downstream signaling is dependent on TIE1 (Figure3).

      (21) Again, experts on computational structure prediction can weigh in, but I am not sure how to interpret the prediction of the 2:2:2 stoichiometry for the theoretical SVEP1/TIE1/ANG1-2 complex. Are there quantitative estimates of the confidence that can be provided? Did the authors attempt to model this complex with different stoichiometries? It is difficult to know how relevant this model is without any experimental results supporting this result.

      Since 1:1:1 is the smallest possible triple complex, it is our starting point. We can model a 2:2:2 version, but anything larger than this AlphaFold will not run. Furthermore, we now provide quantitative estimates of confidence with the pLDDT, PAE, pTM, ipTM scores for all models including the 2:2:2 complexes.

      Minor comments:

      (1) The authors could consider including a reference to alphafold on line 102.

      We have added a reference for AlphaFold2 and 3

      (2) The authors should refer to surface plasmon resonance (SPR) assays by this term as opposed to using the brand name Biacore.

      We agree with the reviewer and have changed the term Biacore to SPR.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the reviewers for their time and for their valuable inputs throughout the review process. We wish to clarify, one final time, the primary scope and empirical grounding of our work for prospective readers.

      Our study was designed to evaluate whether microsaccades track (in a correlative manner) covert visual-spatial attentional shifting, maintenance, or both. We did so within a single dedicated paradigm, across a large sample (N = 48 human participants). Despite remaining criticisms concerning per-participant event counts and microsaccade classification criteria, the key observation remains a striking dissociation (of the link between microsaccades and covert attention) during the initial shifting and the subsequent maintenance of visual-spatial attention. Moreover, we note how the robust effect observed during shifting (but not maintaining) attention, mitigates residual concerns regarding microsaccade sparsity or signal-to-noise ratio.

      We thus remain confident in the empirical foundation of our work and we invite readers to examine the full paper, supplementary materials, and open-access data to evaluate these findings independently.


      The following is the authors’ response to the original reviews.

      We sincerely thank the reviewers and the editors for their careful evaluation of our article and for their valuable input. Building on these suggestions, we were able to further corroborate our main conclusions, make our article more comprehensive, and thereby substantially strengthen the manuscript.

      We have one additional point of our own: we noticed that in our original submission, we had smoothed the data more than intended. Having caught this, we have now reduced the smoothing employed by 2.5 times compared to the original amount of smoothing (the exact smoothing values have also been added to the methods section). Importantly, however, while this has affected how the results look, this has not affected any of our original conclusions.

      General summary

      We would like to first respond to the major points brought forward by both the editorial summary and the public reviews. As we understand, the two main points that were raised regard: (1) the novelty and theoretical importance of our work and (2) the (in)completeness of our results. We start by providing our response to both of these main points below.

      Novelty and theoretical relevance of the work

      Regarding the novelty of our work, we believe the reviews and, by extension, the editorial summary underappreciated the main theoretical value of the question we addressed. Our work set out to investigate whether microsaccades track covert attentional shifting, attentional maintenance, or both. We fully recognise that there are ample prior studies that investigated and reported a link between microsaccades and covert attention, but also underscore how other studies report seemingly contradicting evidence by reporting that there is no such link. One such example is a recent paper by Willett & Mayo in PNAS (2023). Prompted by the recent hypothesis that this seemingly conflicting evidence may be due to prior work investigating attention ‘in different stages’ (van Ede, PNAS, 2023), we set out to address precisely this using a dedicated task that we designed for this purpose. As acknowledged by the summary and public reviews, this helps to reconcile seemingly opposing views in the literature. In our view, such reconciliation has substantial theoretical value.

      While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’). To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. 

      Finally, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Completeness of results

      Regarding the completeness of our results, the editorial summary points to “the absence of independent measures, single-trial analyses, and neutral-condition controls needed to substantiate the central claims”. While the raised points are valuable, they pertain to issues that are tangential to our primary question and stem from unfortunate misunderstandings of key analytical choices, as we now better clarify. We consider our results complete and comprehensive with regards to the main question our studies set out to answer.

      First, regarding the portrayed “need” for independent measures to define the ‘shift window’ of interest, we clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research. 

      Second, regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the single-trial level in our discussion, and prompted by the valuable feedback, we have expanded on this important contextualisation as part of our revision.

      Finally, regarding the portrayed “need” for a neural-attention control condition, we agree that inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ also at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that challenges our main conclusions. Having clarified this, we embraced this valuable reflection and revised the article to ensure that we do not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation, but refer to ‘the presence of an attentional modulation’ instead.

      Taken together, the explicit design and analysis choices that we made align with the theoretical aims of our study, and the central question we set out to address. The raised points are valuable and we are grateful to have been able to leverage them to improve our article, but we hope to have also clarified how they do not render our findings “incomplete” (as currently portrayed) with regards to the key goal of our article.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a study examining the relationship between microsaccades and covert attention. This question has been widely investigated, with numerous studies showing that during sustained fixation, when subjects covertly attend to a peripheral stimulus, microsaccades tend to be biased toward the attended location. Here, the authors ask whether this microsaccade bias reflects a shift of covert attention or the maintenance of covert attention. They conclude that the bias is primarily driven by attention shifts, a finding that also helps reconcile the seemingly conflicting results of prior research, where the bias was questioned in paradigms that largely involved attention maintenance rather than shifting.

      Strengths:

      The paradigm and conclusions appear sound and supported by the results. A large sample size was used.

      We thank the reviewer for this clear and supportive summary of our work.

      Weaknesses:

      Weaknesses are mostly related to how the authors enforced fixation in the task, and clarifications are needed regarding some methodological details. A more direct comparison of the effect in the two experimental conditions is missing.

      We thank the reviewer for raising these valuable points. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”). Regarding the fixation points, we will address them in our point-by-point replies below.

      Reviewer #2 (Public review):

      Summary:

      This study aims to test the hypothesis that microsaccades are linked to the shifting of spatial attention, rather than the maintenance of attention at the cued location. In two experiments, participants were required to judge an orientation change at either a validly cued location (80% of the time) or an invalidly cued location (20% of the time). This change was presented at varying intervals (ranging from 500 to 3,200 ms) after cue onset. Accuracy and reaction times both showed attentional benefits at the valid versus invalid location across the different cue-target intervals. In contrast, microsaccade biases were time-dependent. The authors report a directional bias primarily observed around 400 ms after the cue, with later intervals (particularly in Experiment 2) exhibiting no biases in microsaccade direction towards the cued location. The authors argue that this finding supports their initial hypothesis that microsaccade biases reflect shifts in attention, but that maintaining attention at the cued location after an attention shift is not correlated with microsaccade direction.

      Strengths:

      The results are straightforward given the chosen experimental design. The manuscript is clearly written, and the presentation of the study and its visualisations are both of a high standard.

      We thank the reviewer for this clear summary of our work.

      Weaknesses:

      The major weakness of this paper is its incremental contribution to a widely studied phenomenon. The link between attention and microsaccades has been the subject of extensive research over the past two decades. This study merely provides a limited overview of the key insights gained from these papers and discussions. In fact, it attempts to summarise previous work by stating that many experiments found a link, while others did not, and provides only a relatively small number of references. To make a significant contribution, I believe the authors should evaluate the field more thoroughly, rather than merely scratching the surface.

      We thank the reviewer for this valuable reflection. For an elaborate response to the perceived novelty, please see our general summary reply above. In addition, we have added a more thorough evaluation of the field to the introduction (page 2, find relevant paragraph below). We hope that this will provide more context for the manuscript and strengthen its contribution to the field.

      Revised paragraph from introduction:

      “This link between microsaccades and covert visual-spatial attention has been demonstrated repeatedly. Early studies linked the direction of microsaccades to the deployment of covert attention [14, 15] and these findings were later replicated and extended. For example, it has been demonstrated in both humans [14–25] and non-human primates [26, 27]; during both externally directed perceptual attention [14, 15, 17, 20, 21, 24–27] and internally directed attention within visual working memory [16, 18, 19, 22, 23]; and in both perception and action tasks following directional cues [28]. Several studies have further linked the directional microsaccade bias to task performance [18, 21, 25, 28, 29]. For example, following spontaneous microsaccades, perception of visual targets presented in the same direction is better [25], and visual discrimination benefits may start already prior to microsaccade execution [21]. Recent evidence further suggests that microsaccades may even play a causal role in shaping the perception of peripheral stimuli [30].”

      The authors then present a potential solution to the conflicting past findings, arguing that attention should be considered a dynamic process that can be broken down into an attention shift and a sustained attention phase. Although the authors present this as a novel concept, I cannot think of anyone in the field who considers spatial attention to be a static entity. Nevertheless, I was curious to see how the authors would attempt to determine the precise timing of the attention shift and manipulate the different stages individually. However, the authors only varied the interval between the onset of the attention cue and the test stimulus, failing to further pinpoint their dynamic attention concept.

      The current version of the experiment, therefore, takes a correlational approach, similar to initial studies by Engbert and Kliegl (2003) and Hafed and Clark (2002). Meanwhile, we have learned a great deal about the link between microsaccades and attention. Below, I will list just a few of these findings to demonstrate how much we already know. It is important to note that, while the present study cites some of these papers, it does not provide a clear overview of how the current study goes beyond previous research.

      (1) Yuval-Greenberg and colleagues (2014) presented stimuli contingent on online-detected microsaccades. A postcue indicated the target for a visual task, and the target could be congruent or incongruent with the microsaccade direction. The authors showed higher visual accuracy in congruent trials. The authors cited that paper, but it is still important to emphasize how this study already tried to go beyond purely correlational links on a single trial level.

      (2) The Desimone lab (Lower et al., 2018) showed that firing rates in monkey V4 and IT were increased when a microsaccade was generated in the direction of the attended target.

      (3) However, attention can modulate responses in the superior colliculus even in the absence of microsaccades (Yu et al., 2022)

      (4) Similarly, Poletti, Rucci & Carrasco (2017) observed attentional modulations in the absence of microsaccades, or comparable attention effects irrespective of whether a microsaccade occurred or not (Roberts & Carrasco, 2019).

      Thus, in light of these insights, I believe the current study only adds incrementally to our understanding of the link between microsaccades and spatial attention.

      We thank the reviewer for this insightful comment, and for pointing out several important studies on this topic. While we appreciate that our reported insights may resonate and appear plausible to those working on this topic, we are not aware of any prior studies that directly addressed whether the link between covert attention and microsaccades may fundamentally depend on the ‘stage’ of attentional deployment (‘shift’ vs. ‘maintain’).

      To fill this key gap and address this timely issue, we developed a dedicated experiment designed to evaluate the relationship between microsaccades and the different stages of attention within a single paradigm. We did so by varying the cue-target intervals to uniquely incentivise early shifting (by having short intervals), while also being able to assess microsaccade biases during subsequent maintenance (in the longer trials). To our knowledge, no previous task has jointly examined these components in this manner. Moreover, our inclusion of two widely adopted approaches to fixational control provides yet another source of novelty. Together, we believe that these features position our work as a substantive advance that reconciles seemingly opposing theoretical views.

      Regarding the use of single-trial analyses, we want to emphasise that single-trial predictability is not where our theoretical question resides. We start from the perspective that the relationship between covert visual-spatial attention and microsaccades is inherently probabilistic. Our aim is not to address or question this. Rather, our aim is to determine whether this probabilistic relationship behaves similarly during attentional shifting and maintenance— an issue our analyses directly and appropriately address. In addition, we also explicitly discuss how the link between microsaccades and attention is fundamentally probabilistic at the singletrial level in our discussion. Prompted by the reviewer’s valuable feedback, we have expanded on this important contextualisation in our discussion section (page 8: “Therefore, even if microsaccades may more reliably track shifting than maintaining attention, as our current findings show, our findings should not be taken as evidence that microsaccades reliably track attentional shifts at the single-trial level.”). We also incorporated the valuable reference suggestions in our revised manuscript, including in the revised paragraph in our introduction where we provide a more extensive overview of the prior literature, as shown in response to the preceding comment and in the discussion where we discuss the relationship between microsaccades and attention on a single-trial level.

      In general, it is important to have an independent measure of the dynamics of an attention shift. I think a shift of 200-600 ms is quite long, and defining this interval is rather arbitrary. Why consider such a long delay as the shift? Rather than taking a data-driven approach to defining an interval for an attention shift, it would be more convincing to derive an interval of interest based on past research or an independent measure.

      We thank the reviewer for their question. We wish to clarify how our main analysis is completely agnostic to predetermined time windows, as we employ a cluster-based permutation approach to assess our rich time-resolved data across the full time axis. For the complementary analyses that address the ‘shift’ and ‘maintain’ windows more directly, we use a priori defined windows that are based on ample prior literature (from prior literature studying microsaccade biases, as well as from prior literature on the time course of top-down attention as studied through SOA manipulations). Accordingly, even these ‘zoomed in’ analyses rely on time windows that are empirically grounded in prior research.

      The present analyses report microsaccade statistics across all trials, but do not directly link single-trial microsaccades to accuracy. Similarly, reaction times and accuracy were analyzed only with respect to valid vs. invalid trials. Here, it would be important to link the findings between microsaccades and performance on a single-trial level. For instance, can the authors report reaction times and accuracy also separately for trials with vs. without microsaccades, and for trials with congruent vs. incongruent microsaccades?

      We thank the reviewer for their sincere interest in our findings and for the great suggestion of an additional analysis. We have now investigated whether trials with a congruent, incongruent or no microsaccade in the shift window (where congruent or incongruent was determined as based on the first microsaccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to stress that our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we have decided not to include these analyses.

      The study would benefit greatly from including a neutral condition to substantiate claims of attentional benefits and costs. It is highly probable that invalid trials would also demonstrate costs in terms of reaction times and accuracy. It would be interesting to observe whether directional biases in microsaccades are also evident when compared to a neutral condition.

      We thank the reviewer for this valuable reflection. We agree that the inclusion of a neutral attention condition could be informative for disentangling the ‘benefits’ versus ‘costs’ of attentional cueing. However, such disambiguation is tangential to our central aim. Rather, our behavioural data primarily serve to verify attentional ‘allocation’ at later cue-target intervals. Observing a difference between valid and invalid cues suffices for this central aim. We also note how inclusion of a neutral condition would have reduced trial-numbers and statistical power for our critical conditions of interest. Accordingly, we do not see this as a limitation that in any way challenges our main conclusions.

      Prompted by this reflection, we have ensured to not mention selective ‘benefits’ or ‘costs’ of our cueing manipulation throughout the article, but refer to this only as ‘the presence of an attentional modulation’ instead (such changes were made on pages 3 and 9, and we kept this phrasing consistent in our additions on pages 5 and 22).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The results resemble recent findings by Brandolani et al. (2025), who also showed that the microsaccade bias-in a similar task and using a comparable analysis approach-was restricted to a narrow time window, primarily around the time of the attention shift. The authors should reference this work, discuss similarities and differences, and, given their larger sample size, consider whether they observe a similar correlation between response times and microsaccade rate.

      We thank the reviewer for pointing out this useful reference to us. We have included this article in our introduction, when sketching the current state of the field (page 2) and also mentioned this related article in the discussion (page 8).

      In addition, prompted by this comment, we have investigated whether we observe a similar correlation between response times and microsaccade rate in valid trials. We have investigated this for both possible timeframes: (1) the ‘shift’ timeframe and (2) the ‘maintain’ timeframe. For both experiments, there was no consistent correlation between response times and overall microsaccade rate, as shown in Author response image 1.

      Author response image 1.

      Relationship between the reaction time and the overall saccade rate. This figure shows the relationship between the average reaction time and the average overall saccade rate for both the ‘shift’ period (from 200 to 600 ms after cue onset) and the ‘maintain’ period (from 600 to 1400 ms after cue onset). Each dot represents one participant. Throughout the entire figure, the following significance levels were used: *: p< 0.05, **: p < 0.01, ***: p < 0.001, ****: p < 0.0001.

      (2) I could not find information on the average number of trials per condition and the average number of microsaccades per subject per condition. Ideally, these numbers should be reported (e.g., in a supplementary table). Since the analysis is based on microsaccade direction, knowing how many microsaccades each subject contributed per condition is critical. Microsaccade rates vary substantially across individuals, and subjects with very few events may add noise to the analysis, as proportions of toward/away microsaccades become unreliable.

      We thank the reviewer for pointing out that this relevant information was missing. We have now included these numbers in a supplementary table as suggested (page 17).

      (3) Relatedly, it was unclear whether the time-course analyses were based on collapsing all microsaccade events across subjects or on subject-level averages. In the Methods, the authors state that "the permutation distribution of the largest cluster size was acquired by randomly permuting the trial-average data at the group level 10,000 times," but this is ambiguous. Please clarify.

      We thank the reviewer for pointing out this ambiguity. We have changed the methods section to reflect more clearly that we first obtain time-courses of the microsaccade rate per participant, and subject these time courses to second level statistics using cluster-based permutation analysis (page 13). We have also reworded the sentence you quoted to remove any ambiguity (page 13): “A permutation distribution of the largest cluster size was acquired by randomly permuting the condition labels of each participant’s trial-averaged time course data (i.e. randomly flipping the sign of the difference in rate of toward vs. away saccades) 10,000 times and identifying the size of the largest clusters observed in these randomised data after each permutation.”

      (4) The authors analyze only downward microsaccades, but the cutoff definition is not specified. Presumably, this includes all directions between 180{degree sign} and 360{degree sign}, which may also include nearly horizontal events. This should be clearly stated. In addition, the rate of upward microsaccades should still be shown, divided into up-left and up-right quadrants to parallel the toward/away analysis. This would provide informative context on whether upward microsaccade rates change systematically over time.

      We thank the reviewer for pointing out that this information was missing, and for suggesting this valuable additional analysis. In the methods section, we now explicitly state the angular cutoffs used for the main analysis (page 12). Additionally, we have added a supplementary figure that shows the time course of upwards microsaccades over time (page 21), please see Supplementary Figure S4.

      (5) If the dataset contains enough microsaccadic events per subject, it would be useful to test more conservative angular cutoffs for defining "toward" versus "away."

      We thank the reviewer for this insightful suggestion. We have included an additional analysis, where only microsaccades were included with a direction within a 45° angle around the exacttoward and exact-away directions. This replicated our main finding. The results from this analysis are now included in the supplementary materials (page 24), please see Supplementary Figure S8.

      (6) Figure 2C: It is unclear what the "Center" and "Border" lines represent. The figure is also potentially confusing because it shows microsaccade amplitudes rather than landing positions. Small amplitudes may still bring gaze close to the target; this distinction should be clarified.

      We thank the reviewer for pointing this out. We have changed the “centre” and “border” labels to include more information (pages 6 and 19). We have also added an in-text clarification of the distinction between saccade amplitude and landing position (page 5: “Note that Figure 2C does not show saccade landing positions. While it is theoretically possible for multiple small unidirectional saccades to lead to a larger change in gaze position, a complementary analysis shows that fixation was maintained during the period of peak microsaccade rate in both experiments (see Supplementary Figure S6).”). In addition, in response to the related comment below, we have now also added heatmaps of gaze showing that gaze overall remained close to fixation in our tasks.

      (7) From the Methods, it appears that in Experiment 1, there was no automatic criterion for discarding trials in which gaze deviated from fixation. In Experiment 2, trials were terminated if gaze left a 2{degree sign} window, but given that the target was only 5{degree sign} from fixation, this seems a relatively loose criterion. It would be important to show the distribution of gaze positions during the task to assess whether fixation control was adequate.

      We thank the reviewer for this great suggestion. We have now added a figure to the supplementary materials (page 23) that shows the probability density of gaze position throughout the ‘shift’ period, for left cued trials and right cued trials separately, please see Supplementary Figure S6. We hope that this will further show that even in Experiment 1, fixational control was successful. We also show the difference between left cued and right cued trials, which again shows a gaze bias towards the cued item.

      (8) Did the authors examine whether there was a response time benefit (e.g., RT in congruent microsaccade trials minus RT in incongruent microsaccade trials, as in Brandolani et al., 2025) or an accuracy benefit when microsaccades were directed toward the target?

      We thank the reviewer for their interest in our findings and for the great suggestion of an additional analysis. As we discussed also in response to the general summary from reviewer #2 above, we have now investigated whether trials with a congruent, incongruent or no saccade in the shift window (where congruent or incongruent was determined as based on the first saccade within the shift window) have, on average, different reaction times. This analysis did not show significant differences between these three conditions (congruent microsaccade, incongruent microsaccade, no microsaccade).

      In interpreting this observation, we would like to again stress how our experiment is not particularly well-suited to this analysis, as the amount of time between cue onset and the target events are highly variable across trials. Because of this clear drawback, we decided not to include these analyses. However, please note that we did now include the outcomes of another analysis that more directly targeted the relation between the spatial modulations in microsaccades and task performance, as we turn to below. 

      (9) Was there a relationship between the size of the attentional effect and the magnitude of the microsaccade bias?

      We thank the reviewer also for this insightful question. We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (10) The criteria for minimum microsaccade amplitude and duration are not specified. This should be clarified. I recommend excluding events smaller than ~5 arcmin, as these are likely noise-especially since eye tracking was monocular. Monocular "microsaccades" can be spurious, but this can be determined only with binocular tracking. It is also unclear whether subjects used chin/head rests. A main-sequence plot in the supplementary material would be helpful.

      We thank the reviewer for pointing this out. We have included a main-sequence plot in the supplementary materials (page 23). The main-sequence plot can also be found in Supplementary Figure S7 and suggests that our saccade-detection algorithm worked well with detected saccades following the main sequence. We have also stated more clearly in the methods section that subjects used a chinrest (page 11).

      (11) Please specify the asterisk convention in figure captions (i.e., what * vs. ** vs. *** correspond to in terms of p-values).

      We thank the reviewer for pointing out that these significance levels were not mentioned in every figure caption, so we have added this information to every figure caption where they were missing (page 4, 6, 7 and 19).

      (12) The fact that stricter fixation criteria reduced the size of the effect suggests the possibility that gaze drift toward the target might have conferred an eccentricity advantage in this discrimination task. A direct comparison of the effect in the two experiments would be valuable. The authors should comment on this. It would be informative to plot the average gaze position around the time of peak microsaccade rate in both experiments. Reanalyzing the data post hoc with a stricter trial-selection criterion (e.g., excluding trials where gaze deviated more than 1{degree sign} from fixation) could also be very valuable, as it would systematically test how fixation control influences the observed microsaccade-attention relationship. This would be informative for the community studying this topic.

      We thank the reviewer for these valuable reflections. We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

      Regarding the average gaze position around the time of peak microsaccade rate: in response to reviewer #1, under point (7), we have included Supplementary Figure S6 that shows the probability distribution of gaze position throughout the ‘shift’ period (the same figure is found on page 23 in the article), which shows that fixational control was successful in both experiments. This period is also the period of peak microsaccade rate in both experiments.

      We wholeheartedly agree that systematically investigating the effect of fixational control is important for the field as a whole, and this is also precisely why we set out to perform the same experiment in two different experimental settings with regards to fixational control, and why we decided to include the results from both experiment variants side-by-side in our article.

      Reviewer #2 (Recommendations for the authors):

      In addition to my general concerns in the public review, I have the following recommendations.

      (1) Did the authors distinguish between the initial and subsequent microsaccades during their analysis? Is it possible to produce multiple microsaccades when shifting attention, or do the authors only consider the first microsaccade to be linked to an attention shift?

      We thank the reviewer for pointing out this ambiguity. We have now stated more clearly in the methods section that we consider all microsaccades for our analyses (page 12: “Crucially, we did not restrict our analyses to initial saccades; rather, all detected saccades were included. This allowed us to examine oculomotor behaviour during later trial phases, where initial saccades are unlikely to occur.”). We also believe this methodological choice is important, as otherwise it would be conceivable that no microsaccade bias can be found during the ‘sustain’ period, simply because no ‘first’ microsaccades occur anymore.

      (2) Two microsaccades had to be separated by at least 100 ms. This is an unusually long delay.

      Could the authors please specify how many microsaccades were discarded using this criterion?

      This inter-saccade-interval is quite large on purpose, as we want to minimise the probability of counting the same microsaccade twice. We have re-analysed the data with a minimum ISI of 50 ms, and this led to an increase of found saccades of a, respectively, 5.1% and 1.9% increase for Experiments 1 and 2. However, two participants in Experiment 1 led to a much higher increase in saccades than all other participants (these participants had z-scores of 3.9 and 2.4 for the number of additionally found saccades with an ISI of 50 ms; all other z-scores for Experiment 1 were between -0.5 and 0.5). When those two participants were removed, in Experiment 1 only 1.6% more saccades were found.

      (3) If I understand correctly, the authors did not use staircase procedures to eliminate differences in task difficulty between participants. Could the authors demonstrate how task difficulty relates to the link between microsaccades and performance? For example, is the time course of the microsaccade direction bias correlated with performance?

      We thank the reviewer for this suggestion (that overlaps with a comment of Reviewer 1). We have investigated the relationship between the magnitude of the microsaccade bias during the ‘shift’ period and the behavioural benefit. We have done this separately for a response time benefit and an accuracy benefit. Experiment 1 shows a significant correlation for both reaction times and accuracy with the magnitude of the microsaccade bias, but for Experiment 2 both of these relationships did not survive. Because this relationship did not prove robust across both experiments, but is nonetheless a set of findings our readers will likely be interested in, we have included this figure in the supplementary materials (page 22). Please see Supplementary Figure S5.

      (4) The authors reported using equiluminant stimuli. Could the authors please specify the exact luminance?

      We thank the reviewer for pointing out this missing information. We have now included this information in the methods section (page 11: “, with a luminance of 88.5 cd/m<sup>2</sup>.”). We have also included the luminance of the background (page 11: “luminance: 29.0 cd/m<sup>2</sup>”).

      (5) Could the authors please provide a full polar plot showing all microsaccade directions, and colour-code those included in the analysis?

      We thank the reviewer for this great suggestion on how to present our results even more clearly and comprehensively. We have included a supplementary figure showing the full polar histograms (with 20 radial bins), for all three timeframes of interest: the whole trial, the ‘shift’ period and the ‘maintain’ period (page 20). As requested, the saccades included in the main analyses are colour-coded. See Supplementary Figure S3

      (6) Can the authors please directly compare the main effects between experiment 1 and experiment 2 (Figure 2B)?

      We thank the reviewer for this great suggestion (that was also made by reviewer 1). We have now added a direct statistical comparison of the main effect in the two experimental conditions (page 5: “When directly comparing this spatial saccade bias between experiments, we observed that the effect was significantly larger in Experiment 1 than in Experiment 2 from 535 to 745 ms (cluster p=0.013) and 1153 to 1309 ms after cue onset (cluster p=0.029)”).

    1. Author Response:

      We are grateful for the careful and extensive reviews, and are pleased that the reviewers found the work of broad interest to sensory processing. Please find our proposal for revision based on public reviews:

      Reviewer 1

      1) Request for dose-response curve for DL-TBOA and leak current or RMP. We can provide this, at least for the initial phase of the curve relevant to the concentrations used for synaptic experiments. Prolonged exposure to higher concentrations leads to very large cationic currents (through AMPAR) which appear to be damaging to membrane integrity.

      2) We will increase the N for glial vs neuronal block with the blockers we already used; this seems more practical than doing new experiments with different concentrations of TFB-TBOA. 

      3) We can include data to test the effect of blockers or small depolarizations on excitability.

      Reviewer 2

      1) We differentiated experiments with “25-50 uM” from 200 uM DL-TBOA because the higher concentration clearly led to massive AMPAR activation and depolarization block, as shown in Fig 1. We then chose lower concentrations to minimize background current while allowing glutamate build-up during exocytosis.  We felt we were clear on this point. 

      As to reversibility and “off target effects” like synaptic changes, we will provide this information. See also response to Reviewer 1, comment 1.

      2) See response to Reviewer 1, comment 3.

      3) We are certain that increasing stimulus strength increases the number of stimulated fibers, and this is well accepted. The stimulus electrode is placed in the auditory nerve root, well away from recorded cell and synapses, minimizing current spread to synapses. We can compare PPR for weak and strong stimuli in our current dataset to confirm no effects on release probability. As to variations in the intrinsic properties of myelinated auditory nerve fibers and their sensitivity to stimulation, there is no information about this, and do not understand how it would be relevant, particularly in as much as we report a negative result: no difference in blocker effect with small or large numbers of fibers active. The 3 main types of myelinated auditory nerve fiber, Type 1a,b,c, are known to respond to different sound thresholds, but that is a synaptic issue in the inner ear, and apparently not related to the myelinated axon bundle.

      4) We appreciate the reviewer's caution about a role for neuronal transporters and will revise accordingly.  We cited molecular evidence for expression of subtypes in the pre and postsynaptic neurons and in glial cells. Of course, given how ubiquitous such expression is across the brain, we suspect the kinds of experiments we provided offer more direct evidence for function.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Joint Public Review:

      Summary:

      This manuscript couples a 32-parameter model with simulation-based inference (SBI) to identify parameter changes that can compensate for three canonical hyperexcitability perturbations (interneuron loss, recurrent-excitatory sprouting, and intrinsic depolarisation). The study demonstrates a careful implementation of SBI and offers a practical ranking of "compensatory levers" that could, in principle, guide therapeutic strategies for epilepsy and related network disorders.

      Strengths:

      (1) By analysing three mechanistically distinct hyper-excitable regimes within the same modelling and inference framework, the work reveals how different perturbations require different compensatory interventions.

      (2) The authors adopt posterior estimation to systematically rank the efficiency of different mechanisms in balancing hyperexcitability.

      (3) Code and data are available.

      We thank the reviewers for their positive comments on our manuscript.

      Weaknesses:

      (1) A highly dense presentation of the simulated models and undefined symbols makes it hard for readers outside the modelling community to follow the biological message. An illustration of the models, accompanied by some explanations and references to the main equations and parameters discussed in this paper, would make the first section much more straightforward.

      Thank you for this feedback. To clarify our methods, we have added Figure 7, which illustrates the dynamics of the point neurons and their synapses. We have also added explanations and definitions of variables right where they appear. These variables were previously defined only in a table on a different page.

      We also moved the methods section to the back of the paper, as is common in many modern manuscripts. We hope that relegating method details to the end makes the manuscript more accessible.

      (2) This methodology appears to be a brute-force approach, requiring millions of simulations to tune 32 parameters in a network of 500-700 cells. It isn't scalable. Moreover, the authors did not use cross-validation, which, with a relatively low increase in computational cost, would provide a quantitative measure as to how well it generalizes; this combination raises doubts about both scalability and reliability.

      Scalability is indeed a key challenge of SBI methods. Amortized neural posterior estimation (NPE) is a brute-force approach in that it samples solely from the prior distribution, which is extremely wide. Many of these samples are therefore not very informative for the biologically plausible dynamics we are interested in, which is a downside of amortized NPE. However, amortized NPE is extremely scalable because once the estimator is trained, it can estimate the parameter distribution of any given output dynamic. We tried to build an amortized NPE for our simulator, but simulation-based calibration (a method to validate posterior estimates using additional simulations) showed that the estimators were unreliable.

      Sequential NPE is not a brute-force approach because it samples from posterior estimates, which are narrower than the prior. Because the amortized NPE failed, we use sequential NPE to create the two estimators for the baseline and the hyperexcitable condition described in the paper. While millions of prior samples are used to generate the initial posterior estimate, which is then sequentially refined, the sequential refinement requires only 80,000 additional simulations. This requires a significant amount of computational resources, which is why we consider the results worth reporting, but we make the simulator, the simulation results, and the trained estimators available, so other researchers can use or train their own estimators without running millions of simulations. We hope our rewrites make the advantages and disadvantages of the approach clearer.

      Regarding reliability and cross-validation, we agree that our initial submission has fallen short. We presented results from only one density estimator per condition, which we considered sufficient given the large number of samples. In the revised version, we present the key results from two additional density estimators trained on partially new training data (Figure 4).

      (3) Several parameters remain so broadly distributed after fitting that the model cannot say with confidence which specific changes matter. Therefore, presenting them as "compensatory levers" is somewhat questionable.

      It is indeed difficult to determine which changes matter because of the simulator’s complexity. Especially the marginal correlation coefficient (Figure 2 C) are small and the pairwise histograms are broad (Figure 2A), because all other parameters are unconstrained. But the conditional correlation coefficients are larger (Figure 2D) and narrower (Figure 2B). We have added the histograms in Figure 2B in the revised version to highlight the difference. We cannot provide a definitive threshold for correlation coefficients to discriminate between important and unimportant mechanisms. Therefore, compensatory mechanisms discovered with SBI should be validated mechanistically, as we do in Figure 5.

      (4) Every conclusion is drawn from simulated data; without testing the predictions on recordings, we have no evidence that the proposed interventions would work in real neural tissue. Because today we cannot diagnose which of the three modelled pathological regimes is actually present in vivo, the paper's recommendations cannot yet be used to guide therapy.

      This is indeed an unfortunate drawback of our current work. We are working to apply this approach to constrain microcircuit simulators with data from epilepsy patients. But that work is currently ongoing and will not fit into the present manuscript.

      Recommendations for the authors:

      Beyond the issues I wrote above, which are methodological, I would like to raise my concern about the way this manuscript is written:

      We highly appreciate this editorial feedback on clarity and style. Such feedback is rare and we have worked to address each point to improve the manuscript.

      (1) Paragraphs - several paragraphs start with: "To quantify/identify/find specific compensatory mechanisms of hyperexcitability with simulation-based inference". It is a good idea to orient the reader with the specific goal of each section, but it is not helpful to repeat the overall message of the paper in every paragraph. Several paragraphs open with "However," or "Additionally,". Please restructure sentences so that connectors appear after a clear topic sentence.

      We have done major rewrites to improve the readability of our manuscript. We start paragraphs with more specific context sentences, rather than the broad research goal, and also made paragraphs much shorter with clearer main messages.

      (2) Section 1: The way you present NPE, it would seem like it's specific to neuroscience (and it's not). The paragraph starting at line 26 is not clear. Please revise it. Line 29 - missing a "." before the next sentence begins. Avoid phrases like "for the longest time".

      We now stress that NPE, like SBI, is used across scientific domains.

      (3) Section 2 was tough to read. Please present each equation in its own numbered display, followed immediately by a plain-language explanation of every symbol and parameter. Provide an illustrative diagram: a small schematic of the AdEx neuron, synaptic connections, and the three perturbations. Even a simple block figure will orient nonexperts. Keep critical methodological decisions (priors, summary statistics, simulation length) in the main text, but move voluminous tables of parameter bounds, learning rates, and hardware specs to the supplement. Remove mentions of which Python functions you used. Readers care about algorithmic choices, not function names. Please reserve specific code references for the GitHub README.

      We have added schematic panels at the beginning of Figures 3 & 4 and added Figure 7, which illustrates the neuron and synapse models of the simulator. We also made major rewrites to the methods section to remove programmatic implementation details and define variables where they appear.

      In general, I think it would be a good idea to have an editor to polish syntax, verb tense consistency, and punctuation. A thorough language edit will improve the readability and impact of this manuscript.

      We have attempted to improve the points raised by the reviewer. In particular, we have carefully rewritten verb tense and punctuation throughout the revised manuscript.

    1. Author response:

      We are very happy that our work was positively received by the reviewers and editors and we are looking forward sharing our results via eLife. 

      We have added more details on the generation of the Tribolium brainbow-lines and we have submitted the respective plasmids to Addgene and give the respective IDs. Some additional minor changes were done to make the text more clear. 

      Public Reviews: 

      Reviewer #1 (Public review): 

      Summary:

      Pang et al. investigated the expression pattern of the transcription factor foxQ2II in an adult beetle brain. They find nine distinct clusters, with many neurons expressing Glut/ChaT and dopamine. Some of the dopamine neurons resemble cell types described in Drosophila. Several neurons seem to project to prominent higher brain regions such as the MB and CX, and might even connect to both. 

      Strengths: 

      The authors use state-of-the-art labeling techniques for the analysis of individual cell types, such as beetle brainbow, to investigate the until now unknown expression of the transcription factor in the adult beetle brain. 

      We would want to add that this work establishes and introduces the brainbow system for the first time in an arthropod outside Drosophila melanogaster and that we are the first (outside flies) to relate the expression of a neural transcription factor with neural projection and neurotransmitter content.

      Rigorous cell reconstruction and image analysis revealed a better understanding of the anatomy of the labeled cells. 

      Weaknesses: 

      The brainbow labeling seems to include all cells labeled by the enhancer trap line, as well as the ones not expressing foxQ2II. Thus, it is unclear how useful this data is to compare individual cells to other insects. 

      We kindly disagree with the first statement: not all cells of the enhancer trap are labelled but a subset. Therefore, we call it “sparse labelling” in our manuscript while we do not reach “single cell labelling”, which admittedly limits both precision and use.

      The functional relevance of this transcription factor in the adult brain cell is still unknown. It is therefore unclear if the described neurons have any specific function and if they require this transcription factor for normal function. 

      Previously, we published that this gene has an important function in neural development during embryogenesis. Actually, we have done extensive RNAi experiments to test for an e ect during postembryonic development. We found surprisingly small defects when looking at alterations in several imaging lines. However, we found some changes in behavior. Given the extensive data presented in the current paper, we decided to publish these functional data (another 12 figures/suppl. figures) separately. 

      We also note that the identity/function of neurons is determined by a mix of transcription factors. Disentangling the individual role of each of those transcription factors indeed is an exciting question and a major endeavor beyond the scope of this paper.

      Overall, the neural reconstructions are missing single-neuron details; it is difficult to compare the shown cell types to specific cell types in Drosophila based on the presented data, and this finding remains speculative. 

      Indeed, we do not reach single cell resolution, which is below the standards of fly neurobiology. However, compared with all other arthropods we reach a unique level of precision. Specifically, we are the only ones outside fly research that relate the expression of a developmental transcription factor to neural projection and neurotransmitter content. 

      We also think that combining our transgenic line with dopamine-expression was su icient to compare the labelled cells to fly neurons. From what we saw in that analysis, we feel that most homology assessments of single neurons across such large evolutionary distances will remain hypothetical to some degree.

      Reviewer #2 (Public review): 

      Summary:

      The authors provide the first thorough profiling of neurons in Tribolium characterized by the expression of the transcription factor foxQ2, which will be useful for developmental neurobiology. They use state-of-the-art methods convincingly to not only identify the neurons, but also to further characterize them anatomically and neurochemically. 

      Strengths: 

      Thorough and meticulous application of state-of-the-art anatomical methods in a nonstandard laboratory organism. 

      Thank you for this encouraging comment. 

      Weaknesses: 

      No weaknesses were identified by this reviewer. 

      Comments: 

      I don't really have any major suggestions at all. Loved the work. There is only one tiny nitpicking aspect: 

      P21: "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023), which are functions performed by the mushroom bodies and related to the function of the central complex in goal directed navigation, respectively." 

      MBs mainly process olfactory memory. At least in Drosophila, most other kinds of memories are being supported elsewhere. https://pubmed.ncbi.nlm.nih.gov/10454381/ 

      such as, e.g., visual pattern learning in the CX https://pubmed.ncbi.nlm.nih.gov/16452971/

      or motor learning in motor neurons  https://pubmed.ncbi.nlm.nih.gov/38779314/

      or ventral ganglion, antennal lobes, and median bundle for place learning: https://pubmed.ncbi.nlm.nih.gov/10706599/ 

      If the authors focus on MBs, this sentence ought to reflect the fact that the function of the MBs is much narrower than the current sentence appears to suggest. 

      Thanks for this clarification – we have rephrased:

      "Biogenic amines are involved in learning and memory and setting arousal threshholds (Davis, 2023). This relates to the mushroom bodies’ function in olfactory memory, and the function of the central complex in visual pattern learning and goal directed navigation, respectively."

    1. Author response:

      We thank the editors and reviewers for their thoughtful and constructive assessment of Toothy, and for recognizing it as a potentially valuable resource for standardizing dentate spike (DS) analysis across labs. We are especially glad that the reviewers found the manuscript clear and easy to follow, judged the detection and classification algorithms to be appropriate and well-validated, and appreciated the tool's graphical user interface (GUI) based, pip-installable design for lowering the barrier to entry for DS analysis.

      We also understand the concerns raised. Most importantly, we will resolve the data-ingestion and classification errors that reviewers encountered and release an updated version of Toothy that we have verified end-to-end across input formats and datasets. Alongside this, we will provide downloadable demo dataset(s) spanning multiple file formats, probe types, and recording conditions, so that users can confirm a correct installation and see how these cases differ.

      To make the pipeline more transparent, we will add a section describing what Toothy does between user steps, including the rationale for decisions users cannot change, such as detection from a single representative channel. We will also expand the documentation of parameter choices with supporting citations and alternatives, and surface this guidance within Toothy where feasible, consistent with our aim that the tool not function as a black box.

      We will clarify Toothy's scope and current limitations. Recordings with irregular spatial sampling (e.g., tetrodes) are supported for detection but not for CSD-based DS-type classification, which requires a laminar probe spanning approximately the hippocampal fissure to the hilus; we will state this explicitly and evaluate adding an optional waveform-based classification mode (Santiago et al., 2024) to extend type classification to such recordings. We will also add data-quality checks (including sampling rate and inter-electrode spacing) that warn users when a recording may not support reliable results.

      Finally, we will situate Toothy among existing open-source toolboxes, describing how it differs, extends beyond, and interoperates with them, and we will add a comparison of Toothy's outputs to previously published analyses while being explicit about the limits of such comparisons. We will of course also address the remaining technical clarifications and figure edits raised by the reviewers.

      We are confident that addressing these points will make Toothy clearer and more useful to the hippocampal community.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to differentiate between foveal and peripheral attentional mechanisms in visual and frontal brain regions in monkeys engaged in a free-gaze visual search task.

      Strengths:

      The manuscript is clearly written, the question is important, and the behavioral task is interesting.

      Weaknesses:

      I have two major concerns.

      (1) The authors interpret divergence in neural responses to target vs nontarget as attention. But it is not. The subject has to attend to both target and nontarget stimuli to determine the stimulus category and thereby decide on the next action. Thus, divergence between target and nontarget responses could reflect categorical discrimination, but I am not sure this can be interpreted as attentional modulation. While it may be tempting to suggest that finding a stimulus of a specific category is "feature attention", analogous to, e.g., attending to the red stimulus, I don't believe this is correct. For the former, the animals have to attend to a stimulus, and examine the stimulus to determine the stimulus category, unlike a simpler discrimination, which may pop out. Given this, I am unconvinced that the interpretations in this manuscript are valid.

      We thank the reviewer for raising this concern. Selective attention is a process of focusing on goal-relevant stimuli (targets) while ignoring irrelevant distractions. Importantly, attentional selection is not limited to simple visual features (e.g., color, shape, or motion); it can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [1, 2], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [3, 4]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [5-8], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [7].

      Similarly, in our study, monkeys were trained to search for images that matched the category of the cue. The neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors. We also included only neural responses occurring prior to fixations associated with target selection, that is, before the monkeys made a behavioral choice, thereby controlling for potential contributions of target detection or decision-related signals to the observed effects.

      We have clarified and addressed this point in the Discussion as follows:

      “Feature-based attention to simple visual features such as color, shape, or motion has been extensively studied [1, 3, 5, 7-9, 11, 12, 64]. Attention can also operate over more complex features. For example, objects themselves can serve as units of attentional selection [65, 66], and feature-based attentional effects have been observed when searching for images that match the cued images or image patches [6, 67]. In this context, attention can be directed either to overall features of an object or to objects as configurations of multiple non-spatial features. Furthermore, attention to the category of stimuli has been extensively investigated in fMRI experiments in humans [68-71], and it has been shown that attention can warp the representations of semantically related categories when participants search for different categories [70]. In this study, the neural responses to targets versus distractors were compared while constrained to the same stimuli across different trials, ensuring that the observed response divergence was not due to the physical category of the targets and distractors.”

      (2) Regarding the RF classification of foveal and peripheral RFs for IT and PFC, prior work suggests that neurons in IT cortex (especially AIT) and PFC have RFs that largely include the foveal visual field. So, it would be important to include figures that show the RFs of neurons classified as foveal versus peripheral for all three areas.

      We thank the reviewer for raising this important point. We agree with the reviewer that neurons in IT cortex and PFC often have RFs that include the foveal visual field. We did record foveal units with both focal and broad foveal RFs; however, in our analysis we only included neurons with focal foveal RFs to exclude the influence of peripheral stimuli. We defined focal foveal-RF units as those that responded solely to the cue in the foveal region and not to items in the search array presented at least 5° away from the central fixation point, ensuring that their RFs did not extend to these peripheral locations. The items were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their RFs during fixations. By definition, their RFs were restricted to the central point. This is further supported by Fig. S1A-H, which shows no responses to items in the search array at peripheral locations. We have made modifications in the Results and Methods as follows:

      “Notably, the items in the search array were presented at least 5° from the central fixation point and were also separated by at least 5° from each other, excluding the possibility that peripheral stimuli fell within their foveal RFs during fixations.”

      And:

      “In this study, our focus was on units with focal foveal RFs and units with localized peripheral RFs. All further analyses were conducted on these units.”

      We modified Fig. 1 to illustrate the RFs of neurons classified as peripheral, which were also characterized in our previous study using the same dataset [9]. The peripheral population exhibits no responses to the central cue (Fig. S1I–T).

      Reviewer #2 (Public review):

      Summary:

      In natural visual behavior, such as when one is looking for a face in the crowd, the eyes are moved from site to site, seeking possible matching targets. This involves attention both to the current view at the center of vision (the foveal location) as well as to upcoming views via attention to targets in the periphery. While it has been established that attention generally enhances neuronal response (compared to simple visual activation) at the attended spatial location, this study provides solid evidence that attention during active visual search leads to neuronal response enhancement only when the eye moves towards targets that exhibit the desired feature and category. This study thus moves the field towards understanding the neural encoding of active vision.

      This study examines the neuronal basis of feature-selective attention during active, freely behaving visual search. Traditional electrophysiological studies on visual attention in monkeys commonly used an eye fixation with a covert attention paradigm, but have not sufficiently addressed the roles of both foveal and peripheral attention in play during natural looking behavior. Here, the authors present a novel paradigm in which, during eye-movement mediated search, neuronal receptive fields are recorded in multiple cortical areas (sensory V4, temporal, and prefrontal areas). In this manner, as the eye foveates, items in the array fall into foveal or non-foveal recorded sites. Thus, the experimental paradigm is elegant, offering the opportunity to make multiple types of comparisons: target/distractor, towards/away from fovea, and areal. Specifically, following a category cue (face, house, hand, flower), freely initiated saccades are made to locate a categorically matching 'target' in an array of distractors. Feature attention is assessed by comparing eye saccades made to targets vs to distractors. Spatial attention is assessed by comparing saccades made 'towards' vs 'away' from targets. Statistics are rigorous and nicely designed. The detailed association of simultaneously obtained eye movement sequences and neural parameters is well done. These are valuable data that will contribute to our understanding of attentional modulation in visual search.

      Strengths:

      The significance of these findings is fundamental. Decades of attention research in vision have been based on the paradigm of visual fixation and covert peripheral attention. However, increasingly, the field has moved towards understanding how the visual system works during active vision. Here, the authors use an active visual search paradigm and record from multiple areas (V4, IT, PFC). They find enhancement of attention both in the foveal and peripheral locations, and, furthermore, a high degree of feature and categorical specificity. This provides valuable data for the concept of a foveal-peripheral attentional window in natural vision. The controls (comparisons of neuronal response during looks to targets vs distractors, and looks towards and away from the target) and statistical rigor make these findings quite compelling.

      Weaknesses:

      While the study is generally quite strong, there are a few weaknesses to be addressed.

      (1) Little rationale is provided for recording in the selected areas, V4, IT, and PFC. Given the respective roles in sensory, object recognition, and goal-directed behavior, some rationale for this design should be offered, and commonalities/distinctions between these areas should be discussed.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction as follows:

      “V4 and inferotemporal cortex (IT), as the middle and high-level areas of the ventral visual stream, are important for object recognition and categorization [27-34], and their roles have been extensively studied in central vision. At the neuronal level, however, most investigations have largely neglected their functions during active, free-gaze visual search. The prefrontal cortex, including LPFC, has long been implicated as a source of top-down signals that bias the selection of attended features and modulate visual cortical responses [6, 9, 11, 35-40]. Although target-related visual responses have been reported in IT during visual exploration [41], and target-selective responses have been observed in the human medial temporal lobe (MTL) [42] and medial frontal cortex (MFC) [43] during visual search, these studies did not map the receptive fields (RFs) of recorded neurons.”

      We also added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area—consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      (2) Given the reliance of all analyses on saccadic behavior (towards target/distractor, towards/away from target), additional description and summaries of eye movement behavior during single trials and across trials should be provided.

      We thank the reviewer for this helpful suggestion. We have added a description of saccade behavior to the Results as follows:

      “The mean number of saccades monkeys made to find the target after the onset of the search array was 2.25 ± 1.35 (mean ± SD across trials; Table 1) of correct trials, and the mean saccade amplitude was 7.99° ± 3.58° (mean ± SD across saccades; Table 1). Monkeys could fixate on each distractor or the target freely, provided they did not maintain fixation on the target for longer than 800 ms. Across sessions, 42.44% ± 3.6% of saccades were directed to distractors, 57.56% ± 3.6% to targets, and 12.59% ± 3.46% were saccades away from targets (see our previous studies [44-46] for detailed behavioral analyses).”

      We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search task.

      We have also included Table 1, which summarizes eye movement behaviors.

      (3) The dependency of findings on top-down (categorical & feature-specific) task design should be discussed.

      We thank the reviewer for the suggestion and added a discussion as follows:

      “In this task, attention is strongly guided by top-down goals, which bias processing toward behaviorally relevant features and object categories [2, 50, 51]. Top-down attention, including categorical and feature-specific components, has been shown to modulate neural processing across the visual pathway based on task demands and to originate from distributed frontoparietal control networks [11, 35-38, 40]. Our study provides further insight into the mechanisms of goal-directed visual attention, as it is among the first to demonstrate foveal feature attention effects during free-gaze visual search, as well as the distribution of feature and spatial attention across the entire visual field.”

      Reviewer #3 (Public review):

      In this manuscript, the authors investigate the role of attention in foveal processing during a naturalistic task. They record neural activity from extrastriate visual areas V4 and inferotemporal cortex, as well as from the lateral prefrontal cortex, in macaques performing a free-gaze visual search task. In this task, animals searched for a face or house target among multiple complex stimuli, with no constraints on eye movements. Unlike classic studies of visual attention, which often rely on controlled fixation, this work examines neural activity in both foveal and peripheral receptive fields during naturalistic eye movements.

      The main question addressed by the authors is how feature-based attention is distributed and coordinated across foveal and peripheral visual fields during active search, and how this attentional processing influences saccade behavior. The authors show that foveal units in visual areas exhibit feature-based attentional enhancement, with stronger responses when a fixated stimulus is a target compared to when the same stimulus serves as a distractor. Peripheral units in visual and prefrontal areas show both feature-based and spatial attentional modulation, consistent with prior work. Finally, the authors show that attentional modulation depends primarily on stimulus category rather than response magnitude, with neurons showing similar enhancement for all images within the target category regardless of how strongly individual images drive the cell.

      There are several notable strengths of this paper, including:

      (1) Disentangling feature-based and spatial attention during naturalistic vision remains a central challenge. This paper tackles both simultaneously, parsing neural populations by object selectivity (face-selective, house-selective, non-selective) and RF position (foveal vs. peripheral).

      (2) The unconstrained search task (Figure 1A) moves beyond the dominant fixed-gaze, cued-attention designs (Zhou & Desimone, 2011) to study attention as it operates during natural behavior, with sequential fixations and voluntary saccades.

      (3) The scale of the multi-area recordings is a major strength and is well aligned with current trends in primate and human neuroscience toward large-scale, multi-area recordings. Simultaneous recordings from visual and prefrontal areas, comprising over 4,900 foveal units and more than 1,500 peripheral units, enable meaningful cross-area latency comparisons and area-specific analyses of attentional modulation. This study builds on the authors' previous analyses of this dataset by expanding the scope to show that feature-based attention generalizes across neuronal classes and operates on categorical identity rather than response magnitude.

      (4) The combination of simultaneous multi-area recordings and a rich behavioral paradigm provides a dataset that is well-suited for population decoding, cross-area interaction analyses, and trial-by-trial prediction of saccade choices, which could substantially deepen mechanistic understanding beyond the largely univariate comparisons presented here.

      While the data broadly support the paper's main conclusions, several issues limit the strength of the mechanistic interpretation and should be taken into consideration:

      (1) Receptive field size is not explicitly quantified and may confound foveal-peripheral comparisons. Units are classified as foveal or peripheral based on responsiveness to the cue versus the search array (Methods, p. 17), but the manuscript lacks essential information about receptive field sizes, eccentricities, and the number of search stimuli falling within each receptive field and related proper controls. This is critical because receptive fields in visual area V4 at foveal eccentricities are relatively small (Gattass et al., 1988; Desimone & Schein, 1987), whereas receptive fields in inferotemporal cortex can span several degrees to tens of degrees and often include the fovea (Op de Beeck & Vogels, 2000; DiCarlo & Maunsell, 2003; Zoccolan et al., 2007). Given the 2{degree sign} × 2{degree sign} stimulus size, multiple search items could potentially fall simultaneously within peripheral receptive fields. This introduces a potential confound, as attentional modulation is known to be strongest when multiple stimuli appear within a single receptive field (Reynolds et al., 1999). Although the authors acknowledge this issue for visual area V4 (p. 17), it is neither quantified nor controlled for. Without explicit receptive field mapping relative to the search array, comparisons between foveal and peripheral units, as well as between visual areas, are difficult to interpret cleanly.

      We thank the reviewer for the helpful suggestion and apologize for not explicitly providing essential information about the RFs of the units. We added a detailed description of RF properties to the Results as follows:

      “The RFs of these peripheral units were further mapped using a visually guided saccade task and quantified by the number of stimuli that activated each unit (Fig. 1F-K). The eccentricities of the peripheral RFs were 6.22° ± 1.31° (mean ± SD) in V4, 7.04° ± 1.52° in IT, and 6.68° ± 1.56° in LPFC. The sizes of the peripheral RFs were 3.67° ± 1.87° in V4, 6.86° ± 3.11° in IT, and 8.65° ± 3.02° in LPFC. The numbers of items from the search array falling within peripheral RFs were 1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC (also see our previous study [44]).”

      The reviewer is correct that multiple items from the search array did fall within the RFs of peripheral-RF units. However, for focal foveal units, only the fixated stimulus fell within the RF, due to the design of the search array and the definition of these units used in our analyses (see our reply to Reviewer 1, Public Review, Question 2 for details). We agree with the reviewer that attentional modulation is typically stronger when multiple stimuli fall within RFs. In our design, peripheral RFs, on average, contained more stimuli than foveal RFs. Therefore, this difference in RF size would, if anything, be expected to bias toward stronger attentional modulation in peripheral units. This would make our observation conservative, thereby further supporting rather than undermines our main finding of robust feature-based attentional enhancement in foveal units, challenging the prevailing view that such modulation is predominantly peripheral. However, we agree that, when comparing the latency of attentional effects across brain regions in Fig. 3, we cannot rule out the influence of the number of stimuli arising from differences in RF size.

      (2) Attentional modulation is difficult to dissociate from saccade planning and decision-related signals. The free-gaze paradigm enhances ecological validity but introduces a temporal confound: mean distractor fixation durations are approximately 156 ms (p. 9), while attentional effects emerge between 137 and 170 ms after fixation onset (Figure 2). As a result, the reported attentional modulation coincides with the preparation of the subsequent saccade. Neural activity measured in the primary analysis window (150-225 ms; p. 19), therefore, likely reflects a mixture of visual, attentional, motor planning, target recognition, and behavioral relevance signals, all of which are known to modulate responses in visual areas at similar latencies (e.g., Chelazzi et al., 1998). Moreover, target fixations (~257 ms) and distractor fixations (~156 ms) occur on fundamentally different behavioral timescales, which may inflate apparent foveal attentional effects. While the authors suggest that these timing differences support the idea that foveal feature-based attention facilitates prolonged fixation on target stimuli, this interpretation is not fully supported by the current analyses. That said, the saccade-aligned analyses of peripheral units (Figure S3) partially mitigate this concern by demonstrating that featurebased modulation persists through saccade execution.

      We thank the reviewer for raising this important question. We agree that the temporal overlap of visual, motor planning, target recognition, and behavioral relevance signals with attention can result in mixed activity, which needs to be dissociated. Therefore, when calculating feature-based attention, we did implement a series of controls. We added a discussion as follows:

      “A major challenge in interpreting neural activity related to attentional modulation is the inherent temporal overlap of visual processing, motor planning, and target recognition signals in the free-gaze visual search task [73]. To isolate genuine feature-based attention from potential confounds, we applied several stringent analytical constraints, consistent with prior studies [3, 5, 6]. Specifically, by restricting our analysis to fixations where the subsequent saccade was directed away from the RFs, we dissociated attentional modulation from the preparatory motor activity associated with saccade execution. Furthermore, by comparing responses to the same physical stimulus alternating its role as a target or distractor across trials we eliminated any potential bias introduced by stimulus identity or physical category. We restricted our analysis to fixations preceding target selection that is, before the monkeys made a behavioral choice to minimize contributions from target detection or decision-related signals.”

      We thank the reviewer for pointing out the issue of different timescales for target versus distractor fixations. To address this, we conducted a control analysis by computing foveal feature-based attentional modulation using fixations on targets and distractors with matched fixation durations. We obtained similar results. We have updated Fig. S2 to include this control analysis.

      We also clarified this point in the Results as follows:

      “We also obtained similar results when controlling for fixation durations on targets and distractors (i.e., there was no significant difference between fixation durations on targets and distractors; Wilcoxon signed-rank test, P > 0.05; Fig. S2K–P).”

      Lastly, as the reviewer correctly pointed out, the interpretation that foveal feature-based attention facilitates prolonged fixation on the target was not supported. We have revised the Results as follows:

      “On average, target fixations (256.69 ± 197.44 ms [mean ± SD]) were significantly longer than distractor fixations (156.26 ± 45.94 ms; Wilcoxon rank-sum test, P < 0.0001), and during these prolonged target fixation, foveal feature-based attention modulation was consistently observed.”

      (3) The "attention-out" condition for spatial attention lacks directional control. In the spatial attention analyses (Figures 4D-F), the "attention-out" condition appears to include all fixations followed by saccades directed away from the receptive field, regardless of saccade direction. This differs from classic spatial attention designs, which typically use controlled anti-saccades or saccades to fixed locations opposite the receptive field (e.g., Moore & Armstrong, 2003; Gregoriou et al., 2009). Saccades directed toward locations adjacent to, but outside, the receptive field may still partially engage spatial attention mechanisms near the receptive field via broad attentional fields or motor preparation gradients (Bisley & Goldberg, 2010). In addition, the "attention-out" condition likely contains a heterogeneous mixture of trials in which the stimulus in the receptive field is either a target or a distractor, since feature-based attention effects are derived from this same pool of trials. As a result, spatial and feature attention effects are not fully orthogonal, and variance related to feature attention may already be embedded in the spatial attention baseline.

      We thank the reviewer for this important question. We performed a directional control analysis by computing spatial attentional modulation using paired fixations from the attention-in and attention-out conditions. Only saccades directed in nearly opposite directions—defined as having a saccade direction angle ≥ 170° within the 0–180° range—were included. We obtained similar results (Author response image 1). 

      Author response image 1.

      Peripheral spatial attentional modulation in V4, IT, and LPFC. Population response to stimuli followed by saccades directed into their RFs (attention in) versus directed approximately opposite and outside their RFs (attention out), shown for V4 (A), IT (B), and LPFC (C). Shaded area denotes ±SEM across units.

      We did control for feature-based attention when calculating spatial attentional modulation. We apologize for the lack of clarity and have added a description of this control to the Methods as follows:

      “The saccade-target stimulus in the RF during attention-in fixations was matched to a stimulus in the same location during attention-out fixations; in both conditions, this stimulus always served as a distractor for that trial, except in the “Distractor fixations to T” condition (Fig. 5 and Fig. S4), in which it instead served as the target. This design eliminates differences due to feature-based attention between the attention-in and attention-out conditions.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Figure 3C: Unclear how to compare LPFC vs V4 for foveal units since only data from peripheral LPFC is shown?

      We thank the reviewer for pointing out this mistake. In Fig. 3C, we only compared LPFC peripheral units, V4 peripheral units, and V4 foveal units. We have corrected this in the legend of Fig. 3 as follows:

      “Shown are cumulative distributions of feature-attention effect latencies, computed from individual foveal face-, house-, and non-selective units in V4 and IT, and from peripheral non-selective units in V4,

      IT, and LPFC.”

      (2) On page 8, last para: For units with peripheral RFs ... Is this controlled for whether the saccade is to targets or to distractors?

      We thank the reviewer for the question. We indeed addressed this concern by separating fixations based on whether the subsequent saccade was directed to a target or a distractor, and by analyzing attention modulation within each condition. Therefore, attention effects were evaluated while holding the saccade destination constant, effectively controlling for potential confounds related to saccade target selection.

      (3) Page 9: The authors find that target fixations were longer than distractor fixations and conclude that this supports the idea that foveal feature-based attention increases fixation duration, but this interpretation is pure conjecture, and there is no experimental manipulation presented in this paper that helps to establish this interpretation.

      We thank the reviewer for this important comment. We agree that this observation does not, by itself, support our original interpretation, and we have modified it in the Results. Please refer to the last paragraph of our Reply to Question 2 from Reviewer 3 (Public Review).

      (4) Data analysis: receptive field. The authors state that visual response to a cue and the stimulus array was assessed during the 0-200 ms window after stimulus onset. However, after the array onset, the animal could saccade within the 200 ms window. How do the authors ensure uniform stimulation during the 0-200 ms window?

      We thank the reviewer for this question. The activity of units in V4, IT, and LPFC within the 200 ms window after array onset primarily reflected visual stimulation prior to saccades, because typical saccade latencies were approximately 150–200 ms, and the response onset latencies of these units were around 50 ms.

      (5) On page 19, the authors state that to assess feature attention in peripheral RFs, they divided trials into target and distractor fixations. In the former, there was a target in the neuron's RF. This is confusing. I assume target fixations imply fixating on a target, but the authors may mean fixations where a target is in the RF. Please clarify.

      We thank the reviewer for pointing out this confusion. In the original manuscript, we intended to sort fixations by whether a target stimulus was located within the unit’s peripheral RF. To avoid further confusion, we have revised the description in the Methods as follows:

      “we sorted fixations during the search period, following a procedure similar to that in our previous study [5], into two types: “target” – a target stimulus was located within the unit’s peripheral RF; and “distractor” – the same stimulus appeared in the same peripheral RF location but served as a distractor.”

      (6) Figure S1: Are these example units? How many trials? SEM? The sharp rise and no noise are inconsistent; the former suggests minimal smoothing, while the latter suggests lots of smoothing.

      We thank the reviewer for these questions. We showed average responses across all units in Fig. S1. On average, there were 941.79 ± 182.56 trials (mean ± SD across sessions). Shaded areas indicate ±SEM across units. The sharp rise reflects the synchronous response of neurons to the stimulus, while the smooth appearance and low noise result from averaging across a very large number of units and trials.

      Reviewer #2 (Recommendations for the authors):

      Major comments:

      (1) One weakness of this manuscript is the lack of a rationale for choosing V4, IT, and PFC. Specifically, what are the predictions of the roles of these respective areas in the integration of current and peripheral (future foveal) views? There is a significant literature linking the pre-saccadic peripheral stimulus and the post-saccadic foveal stimulus, suggesting that both spatial and temporal integration occur. However, whether such integration occurs at high or low cortical levels is unknown. By recording from mid-tier (V4) and high-order areas (IT, PFC), the authors have an opportunity to address this question. However, there is no mention of this topic, either in the introduction, results, or discussion. I find this omission surprising. At the very least, it should contribute to experimental design rationale and some discussion.

      We thank the reviewer for the suggestion and we modified and added the rationale to the Introduction and a discussion about this integration. Please refer to our Reply to Question 1 from Reviewer 2 (Public Review).

      (2) As both behavior and neural recordings are collected, a figure on saccadic patterns would enhance the reader's understanding. Questions that come to mind are: What does a single search trial look like? How many saccades are there per trial? How often is the target identified after 1, 2, 3, etc saccades? What is the average size of a saccade? Although this is not a study of search strategy per se, a modicum of description of the search sequences would provide context on the behavior. I suggest an illustration of one or more sample trials; a summary of saccade behavior would also be helpful for understanding the data in relation to behavioral performance.

      We thank the reviewer for this helpful suggestion. We have modified Fig. 1A and its legend to illustrate the saccadic patterns of monkeys during the search, providing an example of a single search trial. Additionally, we have added a description of saccade behavior to the Results and included Table 1, which summarizes eye movement behavior. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for further details.

      (3) "Consistently, the probability of making a saccade to a peripheral target was higher following distractor fixations (75.22%) than following target fixations (48.44%, or 63.49% after probability calibration; see Methods), indicating the important role of peripheral feature-based attention in guiding eye movements" It should be noted that this target-oriented visual search is fundamentally a top down task. Once the target is found, the reward is obtained; saccades to distractors are not rewarded, so saccades are more likely. So certainly this task design would increase the post-distractor saccades and decrease the number of post-target saccades. Please clarify the behavioral paradigm: once a reward is obtained, does the task continue, or is a new trial initiated?

      We apologize for the confusion regarding the behavioral paradigm. We would like to clarify that when the target was found and fixated for 800 ms, the reward was delivered and no further saccades occurred. However, if the target was not fixated for 800 ms, the search could continue. It is worth noting that the target fixations in our analyses were restricted to those occurring during ongoing search behavior, excluding target fixations associated with trial termination and reward delivery. Moreover, we compared the probability of making a saccade to the target, rather than the absolute number of saccades, following these fixations. We have modified the Results for clarification, as follows:

      “Two monkeys performed a category-based visual search task, where their objective was to fixate on one of the two search targets that matched the category of the cue (Fig. 1A, B). Specifically, the monkeys were presented with a central fixation point for 400 ms, followed by a cue lasting 500-1300 ms. After a 500 ms delay, a search array appeared with 11 items, including two targets, randomly chosen from 20 possible locations (Fig. 1E). The monkeys had 4000 ms to find one target and maintain fixation on it for 800 ms to earn a juice reward. Fixating on either target completed the trial, and the monkeys did not search for the second target. A new trial began after the reward. It is worth noting that the two target stimuli matched the category of the cue but were different images. The monkeys were required to maintain fixation throughout the cue and delay periods. During search, however, eye movements were unconstrained, and monkeys could revisit each search distractor or target as long as they did not fixate on a target for 800 ms.”

      (4) The fact that there are many more peripheral units in LPFC suggests that this is a region of foveal/periph integration. Combined with the finding that the LPFC leads the attentional effects, this should be a discussion point.

      We thank the reviewer for the suggestion and we added a discussion as follows:

      “Some studies have provided evidence for integration between peripheral and foveal feature information across saccades, including features such as stimulus color [58, 59] and object orientation [60, 61], and visual features have been shown to be predictively remapped prior to saccades [62]. Our finding provides a potential neuronal mechanism that may support this integration process [63]. We found that LPFC’s extensive representation of the visual periphery provides a neural substrate for monitoring the broader search array. Crucially, our finding that LPFC activity temporally precedes attentional effects in the visual area consistent with previous studies [6, 9, 11, 35-40] suggests that it does not merely reflect peripheral sensory input. Instead, LPFC likely acts as a top-down orchestrator, projecting task-relevant templates derived from current foveal goals onto peripheral candidate locations, a possibility that warrants further investigation.”

      Minor comments:

      (1) Figures 2A-D. "These face-selective units also showed slightly enhanced responses to house targets in IT (P < 0.05), but not in V4 (P = 0.89)." It does not appear enhanced.

      We agree with the reviewer that the effect is modest and does not appear strongly enhanced. However, the average response in the 150–225 ms time window to the house target was significantly higher than that to the house distractor in IT face-selective units (Wilcoxon signed-rank test, P = 0.042). We modified the description in the Results as follows: 

      “These face-selective units also showed weakly but significantly enhanced responses to house targets in IT (P < 0.05)”

      (2) Figure 3. For population comparison, a bootstrapped null distribution was used, and a 2-sided permutation test was used to determine the latency difference between the target and distractor; please show these results (described in text) in a figure. Figures 3A-C are described as the latency of individual units. So each of these graphs is the mean of multiple units? So this is also a population analysis? What is the difference between these two comparisons? This is somewhat confusing.

      We apologize for the confusion and thank the reviewer for pointing this out. Each panel in Fig. 3 shows the cumulative distribution of latencies across individual units within each brain region, reflecting the variability of response timing across single neurons. For this analysis, we first calculate the latency of each unit separately. In contrast, population-level latency is measured from the averaged responses of all units within each region (Fig. 2), which captures the overall timing of the population response rather than individual variability. Statistical comparisons at the population level are performed using a two-sided permutation test. We modified Fig. 2 to better illustrate the population-level latency results.

      (3) Did peripheral RFs span more than a single stimulus in the array? If so, how does this impact the interpretation of Figure 5?

      We thank the reviewer for pointing this out. The reviewer is correct that, in peripheral RFs, more than one stimulus from the search array could fall within the receptive field (1.49 ± 0.55 in V4, 2.2 ± 0.72 in IT, and 2.56 ± 0.74 in LPFC). We controlled for this in our analysis of both feature-based and spatial attention effects for peripheral units in Fig. 5. For feature-based attention, we performed the analysis in a stimulus-by-stimulus manner within each category (house and face), such that when a given stimulus served as the target, it was the only target within the RF, and when it served as a distractor, it was the only distractor of its category within the RF. Although additional distractor could still fall within the RF, their identities were random across conditions and thus would be averaged out. A similar approach was applied to spatial attention, where the stimulus-by-stimulus comparison was extended across all four categories, and attention-out stimuli were paired with the corresponding saccade-target stimuli in the attention-in condition, with the effects of other randomly present distractors averaged out. Therefore, the effects shown in Fig. 5 reflect comparisons at the level of individual stimulus, minimizing confounds from other stimuli within the RF.

      (4) Figure 5G: "during "Target fixations to D", there was no significant feature attentional enhancement in response to the peripheral target (Wilcoxon signed-rank test, P > 0.05; Figure 5G-I left panels). It appears that there is some effect of spatial attention during Target Fix to D trials.

      We thank the reviewer for pointing this out and have revised the Results as follows:

      “We further found that spatial attentional enhancements to the saccade target were reduced during target fixations compared to distractor fixations in V4 and IT when activity was aligned to fixation onset (Wilcoxon rank-sum test, P < 0.05; Fig. 5G, H versus Fig. 5A, B), although this effect was not completely abolished.”

      (5) The specific areas of IT and LPFC that were recorded should, as much as possible, be mentioned.

      We thank the reviewer for the helpful suggestions and have added a description of the specific IT and LPFC recording sites to the Methods as follows:

      “Recordings in IT spanned the central IT cortex, encompassing the area between the anterior middle temporal sulcus (AMTS) and the posterior middle temporal sulcus (PMTS), including TE and TEO. Recordings in LPFC were located anterior to the arcuate sulcus (AS) and lateral to the principal sulcus (PS), mainly covering areas 45 and 44.”

      (6) It is often difficult to distinguish the different lines, e.g., red solid vs red dotted, due to their overlap. Would the removal of the error band make this clearer? If so, could put full figure with error bands in the Supplementary Figure.

      We thank the reviewer for this helpful suggestion. To improve visual clarity, we adjusted Fig. 6, Fig. 7, Fig. S2, Fig. S3, Fig. S4, and Fig. S6 by changing the line styles and placing the shaded error bands beneath the traces, allowing the lines to remain clearly visible despite overlap.

      (7) For easy access, the number of saccades to/from targets/distractors should be put into a table.

      We thank the reviewer for the suggestion. We calculated the probability of saccades to and from targets and distractors for each session and report the mean ± SD across sessions in Table 1, as the mean number of saccades per trial was only 2.3. Please refer to our Reply to Question 2 from Reviewer 2 (Public Review) for Table 1.

      Reference

      (1) O'Craven, K.M., P.E. Downing, and N. Kanwisher, fMRI evidence for objects as the units of attentional selection. Nature, 1999. 401(6753): p. 584-7.

      (2) Baldauf, D. and R. Desimone, Neural mechanisms of object-based attention. Science, 2014. 344(6182): p. 424-7.

      (3) Hayden, B.Y. and J.L. Gallant, Combined effects of spatial and feature-based attention on responses of V4 neurons. Vision Res, 2009. 49(10): p. 1182-7.

      (4) Bichot, N.P., et al., A Source for Feature-Based Attention in the Prefrontal Cortex. Neuron, 2015. 88(4): p. 832-844.

      (5) Reddy, L. and N. Kanwisher, Category selectivity in the ventral visual pathway confers robustness to clutter and diverted attention. Curr Biol, 2007. 17(23): p. 2067-72.

      (6) Peelen, M.V., L. Fei-Fei, and S. Kastner, Neural mechanisms of rapid natural scene categorization in human visual cortex. Nature, 2009. 460(7251): p. 94-7.

      (7) Cukur, T., et al., Attention during natural vision warps semantic representation across the human brain. Nat Neurosci, 2013. 16(6): p. 763-70.

      (8) Keller, A.S., et al., Attention enhances category representations across the brain with strengthened residual correlations to ventral temporal cortex. Neuroimage, 2022. 249: p. 118900.

      (9) Zhang, J., et al., Behavioral and Neural Mechanisms of Face-Specific Attention during GoalDirected Visual Search. The Journal of Neuroscience, 2024. 44(46): p. e1299242024.

      (10) Bichot, N.P., A.F. Rossi, and R. Desimone, Parallel and serial neural mechanisms for visual search in macaque area V4. Science, 2005. 308(5721): p. 529-534.

      (11) Bichot, N.P., et al., The role of prefrontal cortex in the control of feature attention in area V4. Nat Commun, 2019. 10(1): p. 5727.

      (10) Cohen, M.R. and J.H. Maunsell, Using neuronal populations to study the mechanisms underlying spatial and feature attention. Neuron, 2011. 70(6): p. 1192-204.

      (11) Maunsell, J.H. and S. Treue, Feature-based attention in visual cortex. Trends Neurosci, 2006. 29(6): p. 317-22.

      (12) McAdams, C.J. and J.H. Maunsell, Attention to both space and feature modulates neuronal responses in macaque area V4. J Neurophysiol, 2000. 83(3): p. 1751-5.

      (13) Motter, B.C., Saccadic momentum and attentive control in V4 neurons during visual search. J Vis, 2018. 18(11): p. 16.

      (14) Sapountzis, P., S. Paneri, and G.G. Gregoriou, Distinct roles of prefrontal and parietal areas in the encoding of attentional priority. Proc Natl Acad Sci U S A, 2018. 115(37): p. E8755-E8764.

      (15) Treue, S. and J.C. Martinez Trujillo, Feature-based attention influences motion processing gain in macaque visual cortex. Nature, 1999. 399(6736): p. 575-9.

      (16) Zhou, H. and R. Desimone, Feature-based attention in the frontal eye field and area V4 during visual search. Neuron, 2011. 70(6): p. 1205-17.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the efforts of the reviewers, which have provided insightful and helpful comments to improve the manuscript. The feedback touches upon a number of topics, focusing on clarification or justification of experimental techniques and on understanding the mechanism by which P. aeruginosa detects HOCl. All reviewers raised the issue of how HOCl activates fro expression, including whether free or protein-bound methionine, cysteine, or other HOCl byproducts induce this expression. For the upcoming revision, we plan to perform experiments that address this issue and will discuss potential mechanistic models in light of the new data. In addition, we plan to perform additional experiments to address a reviewer’s concerns regarding the dependence of the fro response on HOCl production by neutrophils. The revision will correct imprecise statements pointed out by reviewers, and address all remaining issues requiring clarification or further discussion, including the range of HOCl sensitivity, relationship between HOCl and flow sensitivity, and justification for testing the fro response to nitric acid.

      We have completed a number of experiments and responded thoroughly to reviewer comments below. We thank the reviewers again for their details comments, which suggested additional experiments and interpretations that have resulted in significant additional insight into the potential mechanism of HOCl sensing and its relevance with neutrophils.

      Reviewer #1 (Public review):

      Summary:

      Foik et al. report that hypochlorous acid, a reactive chlorine species generated during host defense, activates the transcription of the froABCD in P. aeruginosa. This gene cluster had previously been associated with a potential role during the flow of fluids and appears to be regulated by the sigma factor FroR and its antisigma factor FroI. In the present study, the authors show that froABCD is expressed both in neutrophils and macrophages, which they claim is likely a result of HOCl but not H2O2 production. Fro expression is also induced in a murine model of corneal infection, which is characterized by immune cell invasion. Expression of the fro system can be quenched by several antioxidants, such as methionine, cysteine, and others. FroR-deficient cells that lack froABCD expression during HOCl stress appear more sensitive to the oxidant.

      Strengths:

      The authors provide a number of data supporting their claim that transcription of the froABCD system is induced by reactive chlorine species. This was shown by RNAseq, qRT-PCR, and through microscopy using a transcriptional reporter fusion. Likewise, elevated expression of froABCD was shown in vitro and in vivo, excluding potential in vitro artifacts. The manuscript, while mostly descriptive, is easy to follow, and the data were presented clearly.

      We greatly appreciate the efforts of the reviewer and thank them for their succinct summary of the manuscript.

      Weaknesses:

      (1) Lines 60-62: Some of the authors' conclusions are not supported by the data and thus appear unfounded. One example: "we determine that fro upregulation.....These data suggest a novel mechanism..." Their data do not show that MSR upregulation is a direct effect of FroABCD. Instead, it could be possible that the FroR sigma factor also controls the expression of msr genes, which would be independent of froABCD.

      We thank the reviewer for pointing out this important distinction. We have clarified in lines 63-65 in the clean version of the revision that MSR upregulation depends on FroR rather than FroABCD.

      (2) The authors show increased fro transcription both in neutrophils and macrophages; however, the two types of immune cells differ quite dramatically with respect to myeloperoxidase activation and HOCl production.

      Neither has this been discussed nor considered here.

      We agree that the distinction between the cell types is important and have added a brief description of the differences in respiratory bursts and ROS production between the two cell types and our justification for focusing on neutrophils in lines 102-103. We think it’s very interesting that Fro appears to be activated by macrophages, which are not associated with HOCl production on their own. We think that it would be interesting to identify what is inducing Fro in macrophages in future work.

      (3) With respect to the activation of fro expression upon challenge with conditioned media from stimulated neutrophils, does the conditioned media contain detectable amounts of HOCl? Do chloramines, which are byproducts of HOCl oxidation with amines, also stimulate expression?

      This is an excellent question that addresses which molecules Fro is responding to from neutrophils. We have performed additional experiments that confirm that PMA-stimulated neutrophils produce HOCl (Fig. S2) through the use of a commercial hypochlorite sensor assay, which claims high specificity for detecting HOCl. We further confirmed that this production is inhibited by pretreatment with the MPO-specific inhibitor 4ABAH. These data support the interpretation that the Fro response to stimulated neutrophils requires MPO activity, of which the major product is HOCl. We have described this in lines 117-127.

      Our data does not exclude the possibility that other MPO products could activate Fro expression. We were unable to obtain a reliable source of the major secondary MPO product, taurine chloramine, for our experiments, unfortunately. We believe that understanding the potential for secondary products to activate the response is an important and interesting question that can be explored in a future study. We have added a discussion of this in lines 367-376.

      (4) A better control to prove that this fro expression is indeed induced by HOCl in activated neutrophils would be to conduct the experiments in the presence of a myeloperoxidase inhibitor.

      We thank the reviewer for raising this point. We have performed the suggested set of experiments and found that indeed, the pre-treatment of neutrophils with MPO inhibitor 4-ABAH prior to PMA stimulation suppresses the activation of fro (Fig. 2D and Figure 2-figure supplement 1-2). The results are discussed in lines 117-127.

      (5) The work was conducted with two different P. aeruginosa strains (i.e. AL143 and PAO1F). None of the figure legends provides details on which strain was used. For instance, in line 111, the authors refer to Figure S1B for data that I thought were done with PAO1F, while in 154, data were presented in the context of the infection model, which was conducted with the other strain.

      We thank the reviewer for pointing this issue out. We have ensured that strain names appear in all the revised figure legends. To clarify, only mouse experiments and a related RT-qPCR assay used strain PAO1F due to prior IACUC approval of this strain and its use in previous publications.

      (6) It would be good if immune cell recruitment at 2hrs and 20hrs PI could be quantified.

      We previously quantified neutrophil recruitment at the site of corneal abrasion at 24 hours using the same conditions and strains (Ratitong, B. et al., J. Immun, 2022). While we do not have immune cell recruitment data for the 2 hr and 20 hr time points, the previous data show significant neutrophil recruitment near the latter time point, which is consistent with the interpretation that fro expression is activated by stimulated neutrophils. We have discussed this in lines 212-216.

      (7) The conclusions of Figure 4 are, in my opinion, weak (line 187-188; "It is possible that ....."). These antioxidants likely quench the low amounts of NaOCl directly. This would significantly reduce the NaOCl concentrations to a level that no longer activates expression of fro. There is no direct evidence provided that oxidized methionine induces fro expression. Do the authors postulate that this is free methionine, or could methionine and/or cysteine oxidation in FroR increase the binding affinity of the sigma factor to the promoter? Another possibility is that NaOCl deactivates the anti-sigma factor. None of these scenarios has been considered here.

      We acknowledge that our model of HOCl sensing was unclear and thank the reviewer for their insight. This critique is echoed by reviewer #2 in comment 3 as well. We recently found that the FroI anti-sigma factor has the highest concentration of methionine and cysteine residues of all known P. aeruginosa anti-sigma factors (Appendix 2—Table 1). Given that FroR and FroI form an extracytoplasmic function sigma – anti-sigma pair, which are associated with transducing extracellular signals to the cytoplasm, we have proposed an alternative model in which HOCl or secondary RCS molecules are detected through their oxidation of cysteine and methionine residues in FroI. This is discussed in lines 271-279 and lines 363-367.

      (8) Line 184: The reaction constants of HOCl with Cys and Met are similar.

      We thank the reviewer for pointing out this important clarification. We have revised the sentence to accurately reflect this in lines 243-244.

      (9) Treatment with 16 uM NaOCl caused a growth arrest of ~15 hrs in the WT (Figure 5A), whereas no growth at all was recorded with 7.5 uM in Figure 3A.

      We thank the reviewer for catching this. We have determined that the concentration of NaOCl in the reagent used for this particular experiment was lower than expected, thus requiring a higher concentration to achieve growth inhibition. We have repeated the experiment with new reagent and find that the results (now Figure 6A) are similar to the previous experiment but at a lower concentration of 4 micromolar, consistent with the concentration found to be sub-inhibitory in Figure 3A.

      (10) The concentration range of NaOCl causing fro expression is extremely narrow, while oxidative burst rapidly generates HOCl at much higher concentrations. This should be discussed in more detail.

      We appreciate the reviewer’s comment, which is related to reviewer #2’s comment #9. We have clarified the reported production rates of HOCl, which far surpass the bacterial MIC. After greater consideration, we believe secondary HOCl products including taurine chloramine could have a more significant role in vivo. While this molecule is less potent than HOCl, it is longer-lived, retains bactericidal activity, and retains the ability to oxidize methionine. We have discussed this in lines 377-405.

      Reviewer #1 (Recommendations for the authors):

      (1) Some statements in the text don't match the data shown in the Figures. For instance:

      (a) Figure 2B shows ~65-fold fro expression, but the text states: "...increased expression of fro expression by 30fold..."(line 88).

      The YFP/mCherry value of PMA-stimulated is 67.4 in the Figure (now Figure 2C). The fold-change is computed relative to unstimulated conditioned medium (third column, which has a value of 2.2), which is a 30-fold change. We have added a citation in the main text to the Source Data, which provides these raw values, to help clarify the computation for readers, and added in the legend that the value for unstimulated is greater than 1.

      (b) Figure 3B shows ~35-fold fro expression at 1 uM NaOCl, but the text states: "...NaOCl increased fro expression by up to 74-fold..."(line 110).

      We have clarified that the increase is relative to untreated, added that untreated value is below 1 in the caption, and provided a citation in the main text to the Source Data, which contains the raw values. The change is measured relative to untreated, for which the YFP/mCherry value is 0.48. The value at 1 uM is 35.7, giving a 74-fold change.

      (2) Line 229: While the ∆froR strain was sensitive to HOCl, the strain was tolerant". Please revise.

      We have corrected this typo (now lines 328- 329). We meant to convey that growth was not entirely inhibited in the froR strain.

      Reviewer #2 (Public review):

      Summary:

      Foik et al. studied the regulation of the fro operon in response to HOCl, an oxidant derived from immune cells, especially neutrophils. They use a transcriptional fusion of YFP to the froA promoter in an mCherry-expressing P. aeruginosa strain to determine fro-induction under the microscope. They use this system to study fro expression in medium, in the presence of neutrophils and macrophages, neutrophil-conditioned medium, and several chemical stimuli, including NaCl, HOCl, hydrogen peroxide, nitric acid, hydrochloric acid, and sodium hydroxide. They also use a corneal infection model to demonstrate that froA is upregulated in P. aeruginosa 20 h post-infection and perform transcriptional analyses in WT and a froR mutant in response to HOCl.

      Strengths:

      Their data clearly shows that HOCl is a strong inducer of the fro Operon. The addition of HOClquenching chemicals together with HOCl abrogates the response. They also show that a froR mutant is more susceptible to HOCl than WT. Their transcriptomic data reveal genes under control of the FroR/FroI sigma factor/anti sigma factor system.

      Weaknesses:

      Although the presented evidence is mostly solid, some of their findings need to be evaluated more carefully; explaining the rationale behind some of the experiments might enhance the article, and some of the models proposed by the authors seem far-fetched, as outlined below:

      We greatly appreciate the reviewer’s efforts and thank them for highlighting strengths and areas for improvement.

      (1) In line 76 the authors claim "Relative to P. aeruginosa that were incubated in host cell-free media, P. aeruginosa in close proximity to human neutrophils or that were engulfed in mouse macrophages appeared to increase fro expression (Fig. 1C)". Counting bacterial cells in Figure 1C shows that 1 in 17 bacteria (5.8%) induce the froA-promotor in media in the absence of immune cells, while 4 in 72 bacteria (only 5.5%) do the same in the presence of neutrophils. Contrary to the authors' claims, it appears that P. aeruginosa actually decreases fro-expression in close proximity to neutrophils. There is a slight increase in fro-expression in bacteria co-incubated with macrophages (3 in 21, or 14.3%). A more rigorous statistical analysis might substantiate the authors' claim, but, as is, the claim "neutrophils increase fro expression" is untenable.

      We believe the images alone do not give an adequate representation of the data and have quantified a larger portion of the data, which has been added as Figure 1D. The quantification supports the original claim that fro expression is increased during co-incubation with macrophages and neutrophils. Since there was not sufficient statistical sampling to distinguish engulfed P. aeruginosa from free ones, this part of the claim has been removed from the text (updated in lines 80-84).

      (2) The authors should explain the rationale behind some of the chemicals used. Why did they use nitric acid? Especially at these high concentrations, a strong acid such as nitric acid might have a significant influence on the medium pH. I understand that the medium is phosphate-buffered, but 25 mM nitric acid in an unbuffered medium would shift the pH well below 2. Similar considerations apply to hydrochloric acid and sodium hydroxide.

      We thank the reviewers for pointing out the need for this clarification. We have updated Figure 3D with a lower concentration of NaOH at 1 uM, which is the same concentration as NaOCl that activates fro expression. Due to the high buffering capacity of our medium, a high concentration of 6 mM NaOH was needed to induce a discernible change in pH and this high concentration of NaOH had no obvious effect on growth. Neither 1 uM nor 6 mM NaOH produced a change in fro expression, consistent with our previous findings that the effect is not due to sodium ions or higher pH. These updated findings are described in lines 181-190.

      Since the effect of chloride is already controlled for using NaCl and the concentration of HCl used was not sufficient to cause a significant change in pH in the buffered medium, we have removed the HCl group from the data.

      We have clarified that nitric acid was used because it is a strong oxidizer that is not found in neutrophils and that concentrations used were near the minimal inhibitory concentrations (lines 145-147 and lines 191-198). We acknowledge that the growth inhibition from HNO<sub>3</sub> could be due to pH or oxidation. However, since no change in fro expression was observed at concentrations approaching the inhibitory concentration, we did not address the potential effects of low pH from nitric acid on fro expression.

      (3) In line 187, the authors state that "It is possible that oxidized methionine increases fro expression" and they suggest a model to that effect in Figure 5D. It is unclear why the authors singled out methionine sulfoxide, since a number of other things get oxidized by HOCl. In line 184, the authors state, in the same vein, that "HOCl oxidizes methionine residues 100-fold more rapidly than other cellular components". The authors should state which other cellular compounds they are referring to. Certainly not cysteine and other thiols, which react equally fast and are highly abundant in the cell: P. aeruginosa contains 340 µM GSH, 140 µM CoA-SH (https://doi.org/10.1074/jbc.RA119.009934) plus free cysteine and cysteines in proteins (based on codon usage, 1.34% of amino acids in proteins are cysteine, while methionine is only slightly more present at 2.10%, although a number of starting methionines are removed from mature proteins).

      We acknowledge that our HOCl sensing model had been vague and unclear and thank the reviewer for their insight. This critique is echoed by reviewer #1 in comment 7 as well.

      Our initial suggestion that methionine sulfoxide was sensed was motivated by the observation that methionine sulfoxide reductases are upregulated by HOCl. However, we have revised this based on feedback from reviewers and further consideration of chlorine redox chemistry. Interestingly, we found that the FroI anti-sigma factor has the highest concentration of methionine and cysteine residues of all the P. aeruginosa anti-sigma factors (Appendix 2—table 1). Given that FroR and FroI form an extracytoplasmic function sigma – anti-sigma pair, which are associated with transducing extracellular signals to the cytoplasm, we propose a model in which HOCl or secondary RCS molecules are detected by their oxidation of cysteine and methionine residues in FroI. This is discussed in lines 271-278 and lines 362-366.

      (4) Overall (and this is probably not addressable with the authors' data), some very interesting questions remain unanswered: what is the molecular mechanism of fro-induction? How is the FroR/FroI system modulated by HOCl? Does the system sense free or protein-bound methionine-sulfoxide? Are certain methionine residues in these proteins directly oxidized by HOCl? Many "HOCl-sensing" proteins are also modified at cysteine residues or amino groups; could those play a role? And lastly: what is the connection between shear/fluid flow and HOCl, or are these totally separate mechanisms of fro-induction?

      We thank the reviewer for raising these excellent mechanistic questions. Issues relating to HOCl sensing are addressed in the preceding comment.

      Regarding the connection to shear sensing, Padron et al., 2023 found that the detection of flow in P. aeruginosa can be attributed to chemical transport, in particular to H<sub>2</sub>O<sub>2</sub> that was present in growth media. Based on the same principle, we expect Fro to be upregulated in flow at much lower concentrations than those observed in stationary fluids. The activation of Fro and the effects of HOCl would thus be expected to be flow-sensitive. We have commented on this important factor in the discussion in lines 406-417.

      Reviewer #2 (Recommendations for the authors):

      (1) To address 1, the authors could evaluate the microscopic images in the same manner in which they evaluated the other microscopic images, as, for example, presented in Figure 2B or Figure S1B.

      See response to Weakness point (1).

      (2) To address 2, please explain the choice of the chemicals (why nitric acid?), but also provide the pH of the media with those high concentrations of strong acids and bases, and interpret them in light of the permissible pH range for P. aeruginosa growth. More sensible controls might be a lower NaOH concentration in the range that would be reached through the amount of NaOH in the NaOCl stock at the highest NaOCl concentrations used. As for the acids, I don't see a reason to use these acids at these high concentrations. Please explain.

      See response to Weakness point (2).

      (3) To address 3: The authors could specify their methionine-sulfoxide model a bit more, so that testable hypotheses can be developed. If the authors think the FroR/FroI system senses free methionine sulfoxide, they or others could add methionine sulfoxide to the medium and check induction. If they think specific methionine residues in these proteins are oxidized, they could provide evolutionary evidence of conserved methionine residues. Or, based on a structure or structural prediction, they (or others) could mutate methionine residues, e.g., at the protein's surface or potential protein/protein-interaction sites and assess the effect on HOCl-based activation. Or they could consider other amino acids known to be highly reactive towards HOCl and mutate those in a future study.

      See response to Weakness point (3).

      Further comments:

      (4) The headline of the figure legend of Figure S1 seems incomplete. Please mention the flow experiments shown in Figure S1A.

      We have updated the title of this figure, which now appears in the eLife format as Figure 1 – figure supplement 1.

      (5) What is the difference between the data presented in Figure 4C (bars "UTR" and "NaOCl") and the same bars in Figure S2A? Is this redundant or a re-plot?

      In this revision, Figure 4C has become Figure 5B and Figure S2A is now Figure 5—figure supplement 1. Only the 1 uM NaOCl condition is replotted. We have described this in the legend for Figure 5—figure supplement 1.

      (6) Line 228: hpd is more likely a gene of the aromatic amino acid catabolism.

      We thank the reviewer for pointing this out. We have removed the ‘branched’ descriptor in this sentence, now in lines 325-326.

      (7) Line 229: "While the ΔfroR strain was sensitive to HOCl, the strain was tolerant (Figure 5A)". Please clarify. Which strain was tolerant? WT?

      We have corrected this typo (now line 328-329). We meant to convey that growth was not entirely inhibited in the froR strain.

      (8) Line 247: "Activated neutrophils produce HOCl concentrations as high as 50 µM [24]." This "50 µM" number is often quoted; the citation trail typically leads to Weiss et al. 1982 (https://doi.org/10.1172/JCI110652). However, a more factually correct statement based on that paper would be "2 x 10^6 neutrophils, activated with 30 ng/mL PMA at 37C in 1 mL of Dulbecco's buffer can produce around 50 nmol HOCl per hour". In the particular reference 24, Dybpukt et al used a methodology similar to Weiss et al., and here around 50nmol were produced by the same number of cells in 30 min in response to 100 ng/mL PMA. Please clarify accordingly.

      We thank the reviewer for bringing this to our attention. We have altered the language to indicate that the production rate is for this specific set of parameters. Related to this, Reviewer #1 Comment #10 requested a more detailed discussion of the significance of the Fro response, since it is at much lower HOCl concentration than produced by neutrophils. We have discussed this in lines 377-405.

      (9) Line 275: "which is strain PA14 strain". Please clarify.

      We have fixed this error and entered it into the Key Resource Table.

      (10) Line 311: "Cultures containing densities below 10 P. aeruginosa per frame were concentrated using a syringe filter with 0.2 or 0.8 µm pore sizes (Millipore, Burlington, MA)." Isn't that a bit concerning when testing the induction of an operon that is supposedly activated by shear through fluid flow? Did the authors convince themselves that this procedure does not induce fro?

      We do not expect the filtering procedure to cause changes in gene expression because cells are imaged immediately after filtering. Nonetheless, we performed additional experiments (Fig 3F and 2D) entirely without concentrating cells and found that YFP/mCherry levels were consistent with previous data in Figure 3. This rationale has been added to the Methods section under the “Fluorescence and Phase Contrast Microscopy” section.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This study provides important insights into the neural mechanisms linking sleep and long-term memory consolidation. By combining behavioural, genetic, imaging, and connectomic approaches in Drosophila, it identifies a target neural circuit that will be of broad interest to researchers studying sleep, memory, and neural circuits. The evidence supporting the involvement of the identified circuit in the regulation of sleep and memory is solid and represents a substantial advance in the field. Nevertheless, there is limited evidence to support the mechanistic claim that this circuit directly links sleep and memory consolidation within the available data, and some results should therefore be interpreted with appropriate caution.

      We appreciate the reviewer’s careful evaluation of our manuscript, and we agree that (as with any experimental study) there is still a lot to do to fully understand the mechanisms underlying the linkage between sleep and memory consolidation.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to use state-of-the art behaviour, imaging and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      The investigation follows a logical strategy to collect huge dataset of sleep, appetitive memory and live imaging. The authors identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation dependent manner. The author also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the new data provide TRIC-LUC provided better temporal resolution of neural activity correlates for PAMalpha1-DPM inhibition. Importantly, the author demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP at the memory consolidation period after training.

      Weaknesses:

      Although the revised version carries arguments to satisfy the reviewers' concern, the writing is now less cohesive. Crucially an explanation however remains required for the following experimental contradiction: the central observation of the study indicates that PAM alpha1 activation cause DPM inhibition which disrupt sleep and memory consolidation. Therefore, one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found the opposite is true from now enhanced TRIC-LUC dataset. The authors indicate this data reinforce the inhibitory nature of PAM-alph1-DPM, but it does not explain why such a reduced DPM activity is observed after training.

      We thank the reviewer for their point of view. We have faithfully reported all experimental observations acquired from our enhanced TRIC-LUC dataset as objectively as possible. We note that the reviewer postulates a particular expected activity shift (reduced PAM-α1 activity and elevated DPM activity after memory training), yet it is not clear why. As our data show, this is a highly connected microcircuit and it is not easy to predict how it might change. Our data show that there is change and dismissing our empirically measured results purely based on a theoretical expectation is not justified. The brain operates as an intricately interconnected network; neural activity dynamics cannot always be simply inferred from static circuit polarity. Progress in deciphering neural circuit function relies on iterative rounds of experimental testing. Like most neuroscience investigations, the present study cannot resolve every open question, and we explicitly acknowledge several unresolved directions worthy of future exploration in the Discussion section.

      Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain incompletely understood. The authors examined a specific subset of PAM dopaminergic neurons, PAM-α1, and DPM neurons in Drosophila. These neurons have previously been implicated in memory, and DPM neurons have also been linked to sleep. The study explores whether this circuit provides a mechanistic link between sleep and memory consolidation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-α1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly during the night. The authors further show that perturbation of PAM-α1 and DPM neurons impairs sleep and appetitive memory consolidation under starvation conditions, and that pharmacological sleep induction during the night rescues the LTM defects. Together, these findings suggest that PAM-α1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important observations that advance our understanding of the circuits regulating sleep and memory consolidation.

      Weaknesses:

      Some claims require additional evidence or clarification.

      (1) Previous studies linking impaired memory to reduced sleep have primarily examined conditions involving severe sleep deprivation. In contrast, this manuscript argues that relatively modest decreases in total sleep, accompanied by sleep fragmentation, are sufficient to impair memory consolidation. It remains unclear whether sleep fragmentation of this magnitude is itself critical for LTM consolidation. An independent method for inducing comparably mild sleep loss and fragmentation would be needed to directly test this interpretation.

      We appreciate the reviewer’s suggestion. While alternative assays for inducing sleep loss or sleep fragmentation are indeed available, this line of investigation lies beyond the core scope of the present study. We will certainly take this valuable suggestion into consideration for the future studies.

      Regarding the question of whether sleep fragmentation of this magnitude per se is critical for long-term memory consolidation, we would like to highlight relevant published evidence. Prior work in rodents and human (Bonnet and Arand, 2003; Van Someren et al., 2015; Ramesh et al., 2012; Baud et al., 2014) and our earlier study (Liu et al., 2019) have demonstrated that alterations in sleep architecture, independent of changes in total sleep amount, can affect multiple physiological processes, including memory. Given this existing supporting evidence, we respectfully argue that this does not constitute a weakness of the present manuscript.

      (2) It is unclear why both activation and inactivation of PAM-α1 neurons produce similar effects on sleep and memory. In addition, MB299B-labeled neurons exert stronger effects on memory than MB043B-labeled neurons, whereas MB043B-labeled neurons have stronger effects on sleep. If sleep disruption is the primary driver of impaired memory consolidation, a stronger correspondence between the sleep and memory phenotypes might be expected. The authors speculate that MB043B may affect sleep through non-PAM neurons, but without identifying the relevant neurons, this remains speculative.

      The concern raised by the reviewer that certain interpretations remain speculative represents a common situation in most published research. This interesting direction warrants further investigation in future work, but falls outside the scope of the present study. In the revised manuscript, we have elaborated on the differences observed between these two GAL4 drivers. We have also conducted additional experiments to investigate a well-characterized memory-related recurrent loop of PAM-α1 neurons in sleep regulation. We respectfully note that no single study can comprehensively address all outstanding questions.

      (3) The complex schematic model (Fig. 12), with parallel circuits and unidentified neuronal groups, underscores the difficulty of interpreting the current data. In the "less activity" arm of the model, distinct circuits are proposed to regulate sleep and LTM, respectively, and DPM neurons are not included. This makes it difficult to reconcile the model with the central claim that the PAM-α1-to-DPM microcircuit links sleep and LTM consolidation.

      We appreciate this careful comment on our schematic model in Figure 12. This diagram aims to summarize the key findings obtained in the present study while also explicitly laying out unresolved questions that await future investigation. In our view, including open, outstanding questions in the working model does not undermine the interpretation of our existing experimental results. In the revised manuscript, we have modified the corresponding text to clarify this point and distinguish firmly between conclusions supported by our data and tentative components requiring follow-up validation.

      (4) The TRIC-LUC reporter system is not ideal for resolving dynamic changes in neuronal activity. Activity-dependent Ca<sup>2+</sup> signaling must first reconstitute the TRIC transcriptional system, which then drives luciferase transcription, translation, and accumulation. The original characterization of TRIC indicates that TRIC signals accumulate and decay over several hours. Thus, the kinetics of the TRIC-LUC reporter should be interpreted cautiously, particularly when inferring transient or precisely timed changes in neuronal activity.

      We fully acknowledge the inherent limitations of the TRIC-LUC reporter system, as pointed out by the reviewer. Every experimental tool comes with characteristic strengths and drawbacks. Although TRIC-LUC suffers from temporal delays, our experiment does not aim to capture acute, immediate effects; instead, it examines long-term dynamics of neuronal activity. To date, within Drosophila neurobiology, no superior technique is available for non-invasive long-term monitoring of neuronal activity in freely behaving flies. We share the hope that new tools capable of reporting neuronal activity in real time will be developed and applied, which will facilitate deeper mechanistic understanding of neuronal dynamics.

      (5) Including data from training under fed conditions would provide a more complete understanding of state-dependent neural activity and would help distinguish starvation-specific effects from more general circuit mechanisms.

      We appreciate this suggestion. First, our memory paradigm relies on reward-based associative learning, and starvation is required for flies to express robust memory, so to do this would require a completely new experimental set up. Second, our core findings demonstrate that transient perturbations of this neuronal circuit trigger sleep disturbances and memory deficits specifically under starvation conditions. Therefore, measurements of neural activity under fed conditions are not directly relevant to the central conclusions of the present study. We agree that related experiments on other behaviors under fed states constitute an interesting direction and could be pursued in future investigations.

      Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence how dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strength:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuit interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of the role of dopamine in cognitive functions. Additional experiments investigating the contribution of MBON-α1 to the circuit, connectomic analysis together with a dissection of dopamine receptor function further strengthen the proposed circuit motif and its biological relevance.

      Weaknesses:

      While the study is well designed, presents compelling findings and has been further strengthened by additional experiments, some aspects remain unclear. The role of DPM neurons in memory consolidation seems not yet fully resolved, as different genetic approaches yield variable results. Furthermore, some manipulations impair memory without affecting sleep fragmentation - or vice versa, suggesting that the observed memory deficits cannot be explained solely by impaired sleep-dependent consolidation. It would also have been interesting to discuss potential mechanisms by which dopamine receptor-mediated cAMP signaling could lead to a reduction in Ca<sup>2+</sup> signals. I am confident that these questions can be addressed in future studies.

      We greatly appreciate the reviewer’s positive evaluation of our work and the thoughtful suggestions regarding future directions. As acknowledged in the manuscript, our study centers on identifying a shared circuit that coregulates sleep and memory processes. We have performed preliminary investigations of downstream circuitry, and our results indeed support the idea that sleep and memory can be modulated independently. Importantly, we have avoided drawing definitive conclusions that memory deficits arise purely from impaired sleep-dependent consolidation. Instead, we emphasize the existence of a common circuit mechanism governing both processes.

      We also thank the reviewer for drawing attention to dopamine receptor-mediated cAMP signaling and Ca<sup>2+</sup>dynamics. This observation constitutes an additional finding that requires more extensive mechanistic follow-up. Given the scope of the current work, we have not dedicated a separate discussion section to dissecting this pathway. This promising line of inquiry will be pursued in our future research.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopaminergic circuits interact to regulate memory consolidation in Drosophila and may reveal general principles underlying the neural regulation of memory.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      No further issue, apart from the reverse figure 13 are not found in the main text as indicated.

      We thank the reviewer for this careful check. We have performed a full-text search for Figure 13 throughout the revised manuscript and found no relevant citation. We have also carefully cross-checked all figure numbering and confirm that all figure labels are accurate in the current version.

      Reviewer #2 (Recommendations for the authors):

      In Fig. 8B, some individual GCaMP measurements show values below −100% ΔF/F₀. Under the stated definition, ΔF/F = (Fn - F0) / F0, values below −100% would require Fn to be negative. Since raw fluorescence intensity cannot be negative, values below −100% require further explanation. The authors should clarify whether Fn represents raw fluorescence or processed fluorescence, and whether the plotted traces underwent any normalization, subtraction, detrending, or transformation beyond the stated formula.

      We greatly appreciate the reviewer’s rigorous scrutiny of our data and the valuable question raised. Our fluorescence signals were calculated using the standard formula ΔF/F = (Fn − F0)/F0, where Fn = F_ROI − F_background, and all calculations were implemented accordingly. We have carefully revisited all raw imaging datasets and identified the source of the issue. During initial data processing, we retained all acquired recordings without excluding samples exhibiting focal plane drift. This drift occasionally yielded negative values for Fn (F_ROI − F_background). Beyond the five DPM cell bodies from four brains in Figure 8B highlighted by the reviewer, we further detected six additional DPM cell bodies from four brains in Figure 2A affected by the same artifact. We have now excluded these drifting preparations, regenerated all corresponding plots, and updated the statistical analyses in the revised manuscript. For transparency, we upload both the raw and processed datasets as supplementary materials to clarify this point.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The goal of the study was to address the question of the degree to which social position in a group is a stable trait that persists across conditions. Reinwald et al. use a custom-built cage system with automated tracking and continuous testing for social dominance that does not require intervention by the experimenter. Remixing of individuals from different groups revealed that social position was rather stable and not really predictable from other measures that were taken. The authors conclude that social position is multifaceted but dependent on characteristics like personality traits.

      Strengths:

      (1) Reductionistic, highly controlled setting that allows for the control of many confounding variables.

      (2) Very interesting and important question.

      (3) Confirms the emergence of inter-individual behavior-driven differences in inbred mice in a shared environment.

      (4) Innovative paradigm and experimental setup.

      (5) Fresh perspective on an old question that makes the best use of modern technology.

      (6) Intelligent use of behavioral and cognitive covariables to generate a non-social context.

      (7) Bold and almost provocative conclusion, inviting discussion and further elaboration.

      We thank Reviewer #1 for this constructive and balanced evaluation of our work, and for highlighting both the conceptual importance of the question and the strengths of our automated, highly controlled approach.

      Weaknesses:

      (1) Reductionistic, highly controlled setting that blends out much of the complexity of social behavior in a community.

      Our goal in developing the NoSeMaze was to provide an enriched yet standardized environment that allows animals to interact freely in a complex setting where arena geometry and access contingencies are held constant across groups. As the Reviewer also notes as a strength, this highly controlled setting with minimal experimenter interference enables us to minimize confounds and provides reproducibility across rounds and groups. We now clarify this explicitly and frame the design as a trade-off between ecological complexity and experimental controllability.

      The NoSeMaze is designed as an open system that can incorporate additional configurations. In this first study with the system, we intentionally used a design that focused on single-sex groups without mating, no intruders or external threats, stable environmental conditions, and adult mice. This allowed us to establish social behavior under one defined condition. From here, future studies can add certain levels of ecological and social complexity to progressively understand their impact on specific behaviors and group dynamics.

      Accordingly, we have revised the manuscript to clarify the scope and the role of social and environmental conditions to the here observed phenomena. We also describe how future studies with systematically modified conditions can be used to understand how they change behaviors. We further replaced potentially over-broad terms (e.g., “naturalistic,” “real-world”) with more precise wording throughout.

      We modified the following sections:

      Abstract

      We removed “… within naturalistic mouse groups.” (ll. 52-54) and changed the sentence to “The approach thus enables longitudinal modeling of individuality and social position as key resilience factors.”

      We changed “… in naturalistic groups.” (l. 38) to “… in larger male mouse groups.” and “… from naturalistic tube competitions …” (ll. 41-42) to “… from incidental competitions in the integrated tube tests …”. We also deleted “naturalistic” in l. 50.

      Discussion (ll. 691-701)

      “…The NoSeMaze aims to increase environmental complexity and group dynamics while retaining experimental control, enabling longitudinal high-dimensional phenotyping of individuals. At the same time, it is a controlled laboratory group-housing habitat optimized to capture a subset of the determinants of social complexity present in natural communities. The strength of this design lies in the continuous, observer-independent observation of complex behaviors in defined environmental and social contexts. Accordingly, we interpret our findings as applying to the social contexts and environmental conditions tested here. Building on this, its modular design allows for introducing additional environmental and social factors like stressors, mating behavior, or resource competition in the future to understand their respective impact in modifying social behaviors.”

      Conclusion

      We changed “… complexity of real-world behavior …” (ll. 729-730) to “… complexity of group behavior in semi-naturalistic conditions ...”

      (2) The motivation to enter the test tube is not "trait" (or at least not solely a trait) but the basic need to reach food and water; chasing behavior would be less dependent on this stimulus.

      Tube traversals may reflect different motivations, including routine movement between compartments to access food, the water lickport, the open arena, or the housing area. Nevertheless, we do not interpret tube-entry motivation itself as a ‘trait’. Importantly, our hierarchy readout does not quantify which animals enter the tubes, nor is it confounded by tube-entry frequency itself (cf. ll. 527-532). Rather, it captures the consistent outcomes of incidental dyadic competitions once two animals meet in the tube (push vs. retreat), aggregated across many interactions. Thus, the stable signal we report lies in repeatable competition outcomes, not in traversal propensity. Consistent with this interpretation, the overall number of tube competition events was not associated with social rank, arguing against systematic competition avoidance by low-ranking animals. In contrast, chasing is a voluntarily initiated, asymmetric interaction between an initiator and a recipient and therefore adds a social-action component beyond incidental access-linked encounters, providing a complementary readout of social behavior.

      We also noted that the term ‘trait’ may not be the optimal description in this context and replaced it throughout the ms. with more precise wording, including “internalized social rank” and “propensity to chase”.

      Results

      We added the following paragraph (ll. 524-532):

      “Tube crossings in the NoSeMaze are motivated by the intent to eat, drink, sleep, or socialize. Accordingly, competitions within the tube arise by chance when two animals enter from opposite sides at the same time. Social rank therefore captures the consistent outcomes of repeated incidental competitions (push versus retreat). This measure is not confounded by differences in tube engagement, as social rank was neither associated with participation in tube competitions (ρ = 0.11, p = 0.139; Fig. 6C, Supplementary Fig. S12C) nor with the overall number of tube detections when controlling for chasing (Spearman’s partial correlation between detection count and z-scored David’s score, corrected for the fraction of active chases: ρ = 0.022, p = 0.758).”

      Discussion

      We changed the following paragraphs:

      Line 608-611

      “Importantly, participation frequency and differences in tube entry time did not confound the resulting social ranks, underscoring the robustness of the automated incidental rank assessment.”

      Lines 653-662

      “These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure.”

      (3) Dominance is only one aspect of sociality, social structure is reduced to rank. The information that might lie in the chasing behavior is not optimally used to explain social behavior beyond the rank measure.

      In this manuscript, we focused on the relationship between social rank derived from incidental tube competitions and chasing. We agree that social rank derived from tube competitions captures only one aspect of social structure. We have now clarified this point more explicitly in the revised manuscript. Here, our specific aim was to determine how chasing relates to competition-derived social rank and whether it provides information beyond rank in these mouse groups.

      We also refer to another manuscript dedicated to additional aspects of social structure captured in the video data, including approach and interaction behavior as well as social clique formation. There, these measures are again examined in relation to chasing and social rank. We also plan future studies leveraging chasing in the NoSeMaze as a readout for neurophysiological investigations.

      In the present study, social rank is based on dominance and subordination in incidental competitions in the integrated tube test. We treat chasing as a distinct, volitional social dimension rather than redundant “rank information”. Importantly, our chasing analyses add beyond social rank information in three ways: (1) structural asymmetry (initiator vs recipient roles) not captured by symmetric social rank measures; (2) elite-centric reciprocal dynamics rather than broad top-down enforcement; and (3) context dependence, with social rank–chasing coupling strengthening when group transitivity is lower. We revised the Discussion/Conclusion accordingly and additionally note that other aspects of social organization (e.g., affiliative bonding and higher-order network structure such as clique/rich-club organization; Nelias et al., 2025) require complementary measures (e.g., video-derived interaction networks) and are not the scope of the present study.

      To avoid ambiguity, we also added the following operational definitions in the Introduction and Results:

      Introduction

      Lines 90-94

      “In this study, we use social hierarchy to denote the group-level structure inferred from incidental competitions in the integrated tube tests, and social rank for an individual’s level within that hierarchy. Social position serves as an umbrella term for social rank and chasing behaviors.”

      Results

      Line 306-308

      “Here, social hierarchy refers to the group-level structure reconstructed from incidental competitions in the integrated tube tests, social rank to an individual’s position within that structure, and social position to social rank together with chasing.”

      We additionally clarified the distinction between chasing and social rank in the following sections.

      Discussion (ll. 632-664)

      “In more constrained or despotic conditions, chasing can serve as a unidirectional, dominance-related behavior directed at subordinates [9,43,44]. The larger groups observed in the NoSeMaze reveal a more nuanced role for chasing behavior. Chasing levels are individually stable, but the expression and meaning of chasing are context-sensitive. Chasing was neither broadly distributed nor consistently directed down the social hierarchy. Instead, it was initiated by a small subset of individuals – primarily those occupying high competition-based social ranks – and frequently occurred reciprocally within this group, suggesting intra-elite social dynamics rather than broad dominance enforcement. Rather than solely serving to impose social hierarchy, chasing appeared to function as a means through which individuals with high social rank monitor, negotiate, and maintain their relative standing within the top tier. Notably, the identity of frequent chasers remained stable over time and persisted across changing group compositions, indicating that the propensity to initiate chases reflects a consistent individual-level tendency rather than a purely situational response.

      However, its coupling to social rank depended on the group’s hierarchy structure. In groups with less clearly defined social hierarchies (i.e., lower transitivity), active chases aligned more strongly with social rank, suggesting that mice in less structured groups rely more on proactive signaling to clarify social rank. Indeed, the top-ranked mice in these groups exhibited relatively high levels of active chases. This context-sensitivity highlights chasing’s dual role: it serves as a tool for negotiating social rank among mice at the upper end of the hierarchy, and additionally functions to establish or reinforce hierarchical clarity when social structures are ambiguous. These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure. These findings reveal a novel aspect of social dynamics: chasing is not merely a dominance display but a flexible context-dependent mechanism shaped by both individual disposition and group-level social structure.”

      Conclusion (ll. 718-723)

      “Crucially, social position is not fully described by a single behavioral dimension. We focused here on two separable dimensions: competition-based social rank and proactive chasing. Future work should integrate additional dimensions of social organization, such as affiliative bonding and higher-order network measures [36], to capture additional aspects of social organization. Chasing behavior played a dual role, reflecting a stable individual propensity while also adapting to group-level structure.”

      (4) Focus on rank bears the risk of overgeneralization for readers not familiar with the context.

      As already discussed above (see also point 2 and 3), we have sharpened the framing to reduce the risk of overgeneralization. Throughout the manuscript, we now refer to tube-derived social rank explicitly as a dominance-subordination-related axis of social organization rather than a comprehensive measure of “sociality”, and we also avoid the term “traits”, but prefer the use of “internalized social rank” or “propensity to chase” when discussing stability across rounds and contexts. Also, we removed “personality-like” as the scope of the study is to set specific behaviors in relation and not to enter the field of mouse personality classification.

      We clarified the definition of social rank in this manuscript in the Introduction (ll. 90-94) and the beginning of the Results (ll. 306-308, see also point 3).

      Abstract

      Lines 41-46

      “… Across more than 4,000 mouse-days, hierarchies derived from incidental competitions in the integrated tube tests were non-despotic, transitive, and stable even when group compositions changed. This stability supports an internalized component of competition-based social rank. Chasing was also stable across contexts. Notably, chasing was concentrated among high-ranking individuals, consistent with ongoing negotiation of social rank among individuals at the upper end of the hierarchy.”

      Lines 50-54

      “In summary, high-dimensional tracking with the NoSeMaze reveals that social position in mice is multifaceted and shaped by stable dimensions of individual behavior that persist across changing social contexts. The approach thus enables longitudinal modeling of individuality and social position as key resilience factors.”

      Discussion (ll. 622-626)

      “This temporal and contextual stability supports interpreting social rank as a stable, internalized characteristic of individuals. In this sense, ‘internalized’ refers to stability across repeated rounds and reshuffled groups in this paradigm and to relatively stable tube-competition outcomes.”

      Changes in the Conclusion starting l. 709 as highlighted in answer to the above point 3.

      (5) Conclusion only valid for the reductionistic setting, in which environment, social and non-social changes only within narrow limits, and in which the mouse population does not face challenges

      Our conclusions are bounded by the conditions tested, namely an enriched but stable environment, controlled group composition and remixing, and the absence of explicit ecological challenges such as resource scarcity or predators. As described above in the answer to point 1, we now state explicitly that our conclusions apply to the social-context variation tested here (controlled group reshuffling) under otherwise stable conditions, and we frame resource-competition and environmental-stressor manipulations as future variation to test of how this manipulation affects the here described behaviors.

      We accordingly changed the Discussion as already highlighted in point 1 above and added ll. 693-704.

      (6) Animals are not naive at the beginning of the experiment, but are already several weeks old.

      Animals entered the study as adults and were continuously group-housed (3-5 mice/cage) before entering the NoSeMaze. To reduce effects of initial apparatus novelty, mice underwent two NoSeMaze habituation sessions prior to data collection (each several hours). We now clarify these points explicitly in the Methods and scope the inference accordingly. One of our future steps is to extend this framework to earlier developmental stages (e.g., adolescence/weaning) to capture full lifespan trajectories. This however first required establishing the approach in adult mice under controlled conditions as done in the present study.

      Methods (ll. 739-747)

      “A total of 79 adult male homozygous OXTRfl/fl mice (B6.129(SJL)-Oxtrtm1.1Wsy/J, RRID: IMSR_JAX:008471, Jackson Laboratory) backcrossed > F10 to C57BL/6J background (Charles River, Sulzfeld) were used for the experiments. Animals entered the NoSeMaze as adults (see Supplementary Table S2 for ages at NoSeMaze entry across rounds) and had been continuously group-housed after weaning (3–5 mice/cage), i.e., they were not developmentally or socially naïve at study onset. Of the 79 mice, 26 mice were injected six weeks before the start of the experiment with an AAV expressing Cre recombinase (rAAV1/2-CBA-Cre) into the AON pars centralis to induce bilateral OXTR deletion (OXTRΔAON). The remaining 53 animals received an AAV expressing only dTomato (rAAV1/2-CBA-dTomato). …”

      Discussion (ll. 629-631)

      “… A future direction is to extend this framework to earlier developmental stages to understand which early experiences shape later trajectories of social position.”

      In summary, this is a wonderful study, but not one that is easy to interpret. The bold conclusion is valid only within the constraints of the study, but nevertheless points in an important direction. The paradigm is clever and could be used for many interesting follow-ups.

      To define social position as a personality trait will elicit strong opposition and much debate; the nuances of the paper might be lost on many readers and call for the (re)-consideration of many concepts that are touched. I find this attitude a strength of the paper, but the approach bears the risk of misunderstanding.

      We thank the Reviewer for the helpful comments. As detailed above, we tightened terminology and framing throughout the revised manuscript by defining stability explicitly in terms of repeatability across rounds and reshuffled groups within this paradigm, and described tube-derived social rank as one dimension of social behavior. We hope these edits preserve the conceptual message while reducing the risk of misunderstanding.

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents the "NoSeMaze", a novel automated platform for studying social behavior and cognitive performance in group-housed male mice. The authors report that mice form robust, transitive dominance hierarchies in this environment and that individual social rank remains largely stable across multiple group compositions. They further demonstrate that social dominance and aggressive behaviors, like chasing, are partially dissociable and that dominance traits are independent of non-social cognitive performance. The study includes a genetic manipulation of oxytocin receptor expression in the anterior olfactory nucleus, which showed only transient effects on social rank.

      Strengths:

      (1) Innovative Methodology:

      The NoSeMaze platform is a technically elegant and conceptually well-integrated system that enables fully automated, long-term monitoring of both social and cognitive behaviors in large groups of group-housed mice. It combines tube-test-like dominance contests, voluntary chase-escape interactions, and an embedded operant olfactory discrimination task within a single, ethologically relevant environment. This modular design allows for high-throughput, minimally invasive behavioral assessment without the need for repeated handling or artificial isolation.

      (2) Experimental Scale and Rigor:

      The study includes 79 male mice and over 4,000 mouse-days of observation across multiple group reshufflings. The use of RFID-based identification, automated data logging, and longitudinal design enables robust quantification of individual trait stability and group-level social structure.

      (3) Multidimensional Behavioral Profiling:

      The integration of social (tube dominance, proactive chasing), physical (body weight), and cognitive (olfactory learning task) measures offers a rich, multi-dimensional profile of each individual mouse. The authors' finding that social dominance traits and non-social cognitive performance are largely uncorrelated reinforces emerging models of orthogonal behavioral trait axes or "animal personalities".

      (4) Clarity and Data Analysis:

      The analytical framework is well-suited to the study's complexity, with appropriate use of dominance metrics, mixed-effects models, and permutation tests. The analyses are clearly explained, statistically rigorous, and supported by transparent supplementary materials.

      We thank the Reviewer for their evaluation of our manuscript. We appreciate their recognition of the NoSeMaze as an innovative and conceptually integrated platform, the scale and longitudinal rigor of the dataset, and the clarity and appropriateness of the analytical framework.

      Weaknesses:

      (1) Conceptual Novelty and Prior Work:

      While the study is carefully executed and methodologically innovative, several of its core findings reaffirm concepts already established in the literature. The emergence of stable, transitive social hierarchies, the persistence of individual differences in social behavior, and the presence of non-despotic social structures have all been previously reported in mice, including under semi-naturalistic conditions (e.g., Fan et al., 2019; Forkosh et al., 2019). Although this work extends those findings with greater behavioral resolution and scale, the manuscript would benefit from a clearer articulation of what is genuinely novel at the conceptual level, beyond the technological advance.

      We agree with the Reviewer that transitive dominance hierarchies and stable inter-individual differences have been demonstrated previously in mice, including in semi-naturalistic settings. Our intent was therefore to address two more specific conceptual questions that prior work typically has not tested in an integrated way:

      (1) Whether an individual’s social position generalizes across distinct social contexts created by systematic changes in group composition rather than reflecting stability that is only observable within a fixed group;

      (2) How the two frequently employed social dominance metrics tube competition-based social rank and chasing relate to each other in unperturbed larger groups.

      Specifically, the conceptual advance is enabled by (A) continuous, handling-free estimation of competition-based social rank from incidental tube contests over weeks, (B) a repeated group-reshuffling (“accelerated longitudinal”) design that explicitly tests whether an individual’s social rank generalizes across distinct social contexts, and (C) parallel, continuous measurement of chasing and non-social reinforcement-learning behaviors in the same individuals and environment.

      This integration also lets us dissociate dominance-related social dimensions (competition-based social rank vs chasing) and test their relation to individual styles in non-social reinforcement-learning. Importantly, these questions are addressed in larger intact societies (9–10 mice), where hierarchy structure is shaped by more complex network-level dynamics than in smaller groups.

      We revised the Introduction and Discussion to foreground these conceptual points and to position them more explicitly in the context of prior works.

      Introduction

      We adapted the following section (ll. 76-98) and integrated the suggested references:

      “… In mice, small groups tend to form highly despotic hierarchies [24-26], whereas larger groups exhibit more complex structures [27]. These patterns suggest a strong influence of emergent group-level dynamics on social structure [12,27]. Yet, animals do not enter social groups as blank slates [28-30]; stable latent factors in the individual may also contribute to hierarchy formation [13]. Thus, it remains unclear to what extent an individual’s social position is internalized and persists across different social contexts [31], or instead is primarily an emergent property of group-level dynamics. Disentangling these possibilities requires experimental conditions that allow unperturbed, continuous tracking of all individuals in sufficiently large groups, together with systematic changes of group composition to modulate social context.

      While stable hierarchies and behavioral identity domains have been described previously in semi-naturalistic settings [9-13], many studies quantify these features either within fixed group compositions or in separate assays. Here, we therefore use systematic group reshuffling in 10-member societies to directly test whether individual differences in social rank and chasing persist across distinct social groups. In this study, we use social hierarchy to denote the group-level structure inferred from incidental competitions in the integrated tube tests, and social rank for an individual’s level within that hierarchy. Social position serves as an umbrella term for social rank and chasing behaviors. By continuously measuring social rank, chasing, and reinforcement-learning behavior in the same individuals, we further test how chasing and social rank are related to each other and how they relate to individual styles in non-social reinforcement learning.

      To enable these tests, we developed the Non-invasive Sensor-rich Maze (NoSeMaze).”

      Discussion

      We added a brief framing statement at the beginning of the Discussion (ll. 577-583) to clarify the study’s conceptual novelty:

      “Ecologically enriched, yet experimentally controlled assessments allow us to study behavioral individuality and social structure in group-living animals over extended timescales. Here, we show that individual mice carry stable, individual-specific, and multi-faceted profiles of social position and cognitive styles across changing group contexts. Our approach goes beyond prior work by testing cross-context stability under repeated, systematic group reshuffling in larger mouse societies, while measuring competition outcomes, chasing, and reinforcement-learning behavior in parallel. …”

      (2) Role of OXTR Deletion:

      The inclusion of the OXTR manipulation feels somewhat disconnected from the manuscript's central aims. The effects were minimal and transient, and the authors defer full interpretation to a separate study.

      We appreciate the Reviewer’s point and agree that the OXTR<sup>ΔAON</sup> manipulation can appear secondary to the manuscript’s central aims. We included OXTR<sup>ΔAON</sup> because oxytocin-dependent social recognition memory was hypothesized originally to impact potentially also learning an individual’s position in social hierarchy networks. Even though the effects were small and transient, we nevertheless believe it is relevant to report them, also in relation to the more profound effects of the genetic manipulation reported in a related manuscript (Nelias et al., bioRxiv 2025, 10.1101/2025.08.26.672298). Therefore, we explicitly account for genotype as a covariate in the analyses such that the main conclusions do not depend on this manipulation.

      To improve coherence, we have added one sentence in the Introduction motivating why OXTR<sup>ΔAON</sup> was included, and finally more clearly signposted that deeper mechanistic interpretation is beyond the scope of the present manuscript.

      Introduction (ll. 115-120)

      We added one short sentence introducing the rationale behind the perturbation:

      “… and (3) determine whether social rank, chasing, and non-social reward-seeking behaviors represent stable individual characteristics or dynamic features across time and changing group composition. As a secondary analysis, motivated by oxytocin’s established role in social recognition memory [32,33], we also tested whether OXTR deletion in the anterior olfactory nucleus produces detectable shifts in rank dynamics.”

      Results

      Lines 161-164

      “… This manipulation was included as a secondary biological perturbation. The primary analyses and conclusions focus on the platform and cross-context stability, and genotype is treated as a covariate unless stated otherwise. …”

      Lines 451-454

      “… As a secondary analysis, we tested whether OXTR<sup>ΔAON</sup>, which impairs de novo social recognition memory required for social clique formation in this cohort [36], also affects the measures reported here. Consistent with largely internalized features, OXTR<sup>ΔAON</sup> produced only transient effects. …”

      Discussion (ll. 594-605)

      “This study focused primarily on the relation of social rank and chasing. We however also considered their relation to additional variables including the loss of oxytocin receptors in the olfactory cortex in the adult (OXTR<sup>ΔAON</sup>), involved in de novo social recognition learning. The propensity to chase was largely unaffected by OXTR<sup>ΔAON</sup>. Mice carrying OXTR<sup>ΔAON</sup> displayed a transient reduction in social rank during the first week that normalized thereafter. This transient effect contrasts to the persistent impairment by OXTR<sup>ΔAON</sup> in forming higher-order social bonds that enable membership in stable cliques, as identified by video tracking of self-paced interactions in the same cohort [36]. Together, these findings suggest that OXT-dependent olfactory learning is critical for the formation of social context-dependent higher-order bonds, but plays a limited role in shaping hierarchy-related behaviors.”

      (3) Scope Limitations (Sex and Age):

      The study is limited to male mice, and although this is acknowledged, the title and overall framing imply broader generalizability. This sex-specific focus represents a common but problematic bias. Additionally, results from the older mouse cohort are under-discussed; if age had no effect, this should be explicitly stated.

      We thank the Reviewer for this relevant point. The study is limited to male mice. We therefore revised the title and the abstract and strengthened the limitations to make the sex-specific scope explicit.

      We additionally note ongoing work extending the same framework to female groups, where we find similar hierarchy structure and cross-context stability in the NoSeMaze in preliminary unpublished data.

      Regarding age, our design included two separate cohorts of different adult age ranges (young and older adult animals), but age was not the primary experimental factor. As detailed in our response to Reviewer #3 (point 1), we now quantify age structure explicitly and test age effects using a decomposition that separates between-group age differences from within-group age variation (mean age per group and each animal’s deviation from that mean). We also recomputed stability estimates with age-adjusted ICC models, and the resulting ICCs are highly similar to the original estimates (cf. new Supplementary Table S4), indicating that the reported metrics’ stability is not driven by age differences across groups.

      Title

      We change the title from “Individual differences drive social hierarchies in mouse societies” to “Individual differences drive social hierarchies in male mouse societies”

      Abstract

      We also added male in the abstract (ll. 37-38):

      “The interaction of these behaviors in the shaping of social position in larger male mouse groups remains largely unknown.”

      Discussion (ll. 613-617)

      “… While this study focused on male mice, in which social hierarchies are best established [9], future work is needed to explore sex-specific expressions of social structure and their neurobiological underpinnings in female groups. We therefore restrict our interpretation to male mice. Critically, male social ranks were robustly maintained within the same group over time. …”

      For a more detailed integration of age in the Methods, Results, and Discussion, we kindly refer to the reply to Reviewer #3, point 1.

      (4) Ambiguity of Dominance as a Construct:

      While the study robustly quantifies social rank and hierarchy structure, the broader functional meaning of "dominance" remains unclear. As in prior work (e.g., Varholick et al., 2019), dominance rank here shows only weak associations with physical attributes (e.g., body weight), cognitive strategy, or neuromodulatory manipulation (OXTR deletion). This recurring pattern, where rank metrics are reliably established yet poorly predictive of other behavioral or biological traits, raises important questions about what such measures actually capture. In particular, it challenges the assumption that outcomes in paradigms like the tube test or chase frequency necessarily reflect dominance per se, rather than other constructs.

      We thank the Reviewer for this clarifying point. We agree that stable social rank metrics do not necessarily imply a complete or unitary measure of “dominance.” In the revised manuscript, we therefore clarified that, in this study, social rank is operationalized as consistent competitive outcomes in incidental tube-test encounters in the NoSeMaze (see also Reviewer #1, points 2-4).

      Our data indicate that this competition-based social rank is related to, but not identical with, other social behaviors such as chasing. Likewise, body weight significantly contributes to social rank, but explains only part of the variance. We therefore do not interpret weak or partial associations with other variables as invalidating the social rank measure. Rather, we interpret them as indicating that social position is multidimensional and only partially captured by any single assay.

      In independent subsequent studies that are currently in preparation or revision, we observed that heterogeneity in the neurobiology and response to challenges was best predicted by the competition-based social rank, also compared to the other behaviors assessed here. While these observations are beyond the scope of the present manuscript, they support our view that incidental tube competition and the resulting dominance-subordination structure may provide a biologically informative measure of one important dimension of social position. At the same time, the observation here and in many previous studies that factors such as body weight explain only a limited portion of the variance remains important for our understanding of these constructs.

      We have revised the manuscript accordingly to make this distinction more explicit. In particular, we now define more precisely the terms social hierarchy, social rank, and social position in the context of this manuscript (see also reply to Reviewer #1, point 3; ll. 90-94 in the Introduction and ll. 306-308 in the Results), and we use these terms more consistently throughout. We also revised the text to avoid overstating the meaning of “dominance” where the data support a more specific interpretation.

      Finally, we now emphasize more clearly both the strength and the limitation of the present approach. The NoSeMaze allows these relationships to be assessed continuously in a minimally perturbed group-housing ecology, laying ground for future incorporation of further variables and dimensions to capture how they shape social organization.

      Specifically, we added text in the Discussion and Conclusion to clarify that competition-based social rank and proactive chasing represent separable dimensions related to dominance and subordination, but do not exhaust sociality or individuality, and that future work should integrate additional measures such as affiliative behavior and higher-order network structure.

      Discussion

      Lines 624-6631:

      “… In this sense, ‘internalized’ refers to stability across repeated rounds and reshuffled groups in this paradigm and to relatively stable tube-competition outcomes. The tube-derived social rank describes here a dimension of individual social behavior. The capacity to quantify stable individual differences across changing social contexts highlights the value of the NoSeMaze for lifespan-oriented studies of behavioral individuality. A future direction is to extend this framework to earlier developmental stages to understand which early experiences shape later trajectories of social position.”

      Conclusion

      Lines 718-722:

      “Crucially, social position is not fully described by a single behavioral dimension. We focused here on two separable dimensions: competition-based social rank and proactive chasing. Future work should integrate additional dimensions of social organization, such as affiliative bonding and higher-order network measures [36], to capture additional aspects of social organization.”

      Reviewer #3 (Public review):

      Reinwald et al. present the NoSeMaze, a semi-natural behavioral system designed to track social behaviors alongside reinforcement-learning in large groups of mice. Accumulating more than 4,000 days of behavioral monitoring, the authors demonstrate that social rank (determined by tube competitions) is a stable trait across shuffled cohorts and correlated with active chasing behaviors. The system also provides a solid platform for long-term measurements of reinforcement learning, including flexibility, response adaptation, and impulsiveness. Yet, the authors show that social ranking and chasing are mostly independent of these cognitive traits, and both seem mostly independent of oxytocin signaling in the AON.

      Strengths:

      (1) The neuroethological approach for automated tracking of several mice under semi-natural conditions is still rare in social behavioral research and should be encouraged.

      (2) The assessment of dominance by two independent measures, i.e., spontaneous tube competitions and proactive chasing, is innovative and valuable.

      (3) The integration of a long-term reinforcement-learning module into the semi-natural system provides novel opportunities to combine cognitive traits into social personality assessments.

      (4) The open-source system provides a valuable resource for the scientific community.

      Limitations:

      (1) Apparent ambiguity and inconsistency in age structure and cohort participation across rounds, raising concerns about uncontrolled confounds.

      (2) Chasing behavior appears more stable than tube-test competitions (Figure 4D vs. Figure 3D), which challenges the authors' decision to treat tube competitions as the primary basis for hierarchy determination.

      We thank the Reviewer for the evaluation of our work. We have addressed the limitations raised below with additional analyses, clarifications, and corresponding manuscript revisions.

      Major concerns:

      (1) Unclear and inconsistent handling of age groups and repeated sampling. The manuscript repeatedly refers to "younger" and "older" adults, but it is unclear whether age was ever controlled for or included in models. Some mice completed only one round, others 2-5 rounds, without explanation of the criteria or balancing.

      We thank the Reviewer for this clarifying comment. We now explicitly quantify participation in the different rounds in new Supplementary Table S3. Importantly, our primary stability analyses are implemented using variance-component mixed models (REML) that naturally handle unbalanced repeated-measures data. This is explained in more detail in the Methods section (ll. 1003-1111), where we added a statement that the LMEs are well suited for unbalanced repetitions.

      To further address unbalanced participation, we additionally performed a conservative sensitivity analysis restricted to a balanced subset, including only sessions 1 and 2 and only mice with observations in both sessions. For the key social measures, ICC estimates were highly similar in the full dataset and in the balanced subset (e.g., z-scored competition David’s score, ICC across cohorts: 0.55 without age adjustment vs. 0.56 in the balanced first-two-session subset; active chasing: 0.74 vs. 0.72; being chased: 0.61 vs. 0.60; Supplementary Table S4), indicating that unbalanced participation did not inflate the stability estimates.

      We also addressed age structure explicitly. Although age was not the primary experimental factor, the inclusion of two age cohorts allowed us to assess whether the observed behaviors and their interrelations were robust across most of the adult lifespan (cf. new Figure 1). We therefore decomposed age into a between-group component (mean age per group) and a within-group component (each animal’s deviation from its group mean) and included these terms in the relevant LME models. Tube-based dominance rank (David’s score, z-scored) showed no age effect (p<sub>age, cond.</sub> = 0.63), whereas chasing metrics showed modest age associations (cf. new Supplementary Table S5). This however only indicates that chasing was associated with age to some degree. More importantly, recomputing all stability estimates using age-adjusted ICC models yielded nearly identical ICCs for the core social measures (new Supplementary Table S4), indicating that the reported stability was not affected by age differences across groups.

      In addition, we revised the study design schematic (new Fig. 1; cf. Reviewer #1, Recommendations for the authors) to depict the separate age cohorts and the reshuffling procedure more clearly.

      We adapted the following sections accordingly.

      Methods

      Lines 980-985

      “… Most mice (n = 68) participated in at least two NoSeMaze rounds with reshuffled group members. Supplementary Table S2 summarizes the number of rounds per mouse and missing data due to technical problems. We examined the stability of social and reward-seeking metrics by correlating values from the first and second round (Spearman’s correlation, cf. Fig. 5). These round-1-to-round-2 correlations use one paired observation per mouse and are therefore not inflated by mice contributing >2 rounds.”

      Lines 991-993

      “The ICC treated mouse identity as a random intercept and NoSeMaze group (i.e., round-specific social group) as a random effect, with repetition included as fixed effect (i.e., stability across groups while holding repetition means constant).”

      Lines 1003-1011

      “Variance components for ICC estimation were obtained from LMEs fit by restricted maximum likelihood (REML), which yields less biased variance-component estimates and is well suited for unbalanced repeated-measures designs (i.e., different numbers of rounds per mouse). To assess potential confounding by age structure, age was decomposed into a between-group component (group-mean age) and a within-group component (each animal’s deviation from its group mean) and included as covariates. ICCs were recomputed in age-adjusted models (see Supplementary Table S4). Finally, we performed a conservative sensitivity analysis restricted to the first two sessions per mouse and to mice with observations in both sessions (“balanced first2”), to additionally account for unbalanced participation structure.”

      Results

      Lines 167-173

      “… The study population comprised two age cohorts: younger (16-30 weeks) and older adults (55-97 weeks) (Fig. 1B). The two age cohorts were run as separate experimental series, and group reshuffling was performed within each cohort (Fig. 1C, see Supplementary Table S2). Mice lived in groups of 9-10 for multiple rounds in the NoSeMaze, with different group members in each round (Fig. 1D, see Supplementary Table S1). This allowed us to test which individual behaviors were stable across different group compositions. …”

      Lines 441-450

      “Because the number of rounds in the NoSeMaze was unbalanced between animals (Supplementary Table S3) and age varied across groups (Supplementary Table S1-2), we also performed robustness checks. Stability estimates changed only minimally when recomputed in age-adjusted ICC models (between-group mean age and within-group age deviation; see Methods) and when restricting analyses to a balanced first-two-session subset (sessions 1–2 only; mice with both sessions) (Supplementary Table S4). Mixed models indicated that some chasing and reinforcement-learning measures showed modest age- and/or session-related shifts in absolute levels (Supplementary Table S5), but importantly, these did not affect the observed stability patterns.”

      (2) Stability of chasing appears stronger than the stability of tube competitions. Figure 4D shows highly consistent chasing behavior across weeks, while Figure 3D shows weaker and more variable correlations for tube-based David scores. This is also evident from Figure 5A-B,D. Thus, it appears that chasing, which serves to quantify dominance in similar semi-natural setups, may be a more reliable and behaviorally meaningful measure of dominance than the incidental tube competitions.

      Indeed, active chasing showed higher cross-round correlations than tube-derived social rank (e.g., R1–R2 Spearman ρ = 0.75 vs 0.57; ICC<sub>across cohort</sub> 0.74 vs 0.55, see Fig. 5 and new Supplementary Table S4). Importantly, this does not contradict the central finding that tube-derived social rank is stable across time and across remixed groups. Rather, it may highlight that chasing and social rank capture different aspects of dominance-subordination-related behavior with different statistical properties. Chasing reflects an individual’s propensity to actively initiate interactions (a strongly expressed, asymmetric behavior), which can be highly consistent across contexts. By contrast, social rank is a relational measure inferred from symmetric dyadic win–loss outcomes based on incidental competitions in the integrated tube tests and can vary with the specific set of competitors and interaction opportunities in each reshuffled NoSeMaze group, while still remaining substantially stable overall.

      Accordingly, we do not interpret the higher repeatability of chasing as evidence that it is the “better” hierarchy measure. Tube competitions yield symmetric dyadic outcomes that directly support formal hierarchy reconstruction (David’s score/Elo, transitivity, steepness) and show convergent validity with traditional tube testing. Chasing, in contrast, is asymmetric and volitional and, in our data, is concentrated in the upper social ranks and modulated by group-level social hierarchy structure, consistent with rank negotiation/signaling rather than a mechanism that assigns a full ordering to all individuals. These different properties lead to very different “win-lose” relationships when comparing tube competition events to chasing events that we specifically illustrated in Supplementary Fig. S13. We therefore revised the Results and Discussion to frame tube-derived social rank and chasing as complementary social dimensions.

      Results (ll. 402-413)

      “… Specifically, stability was high for the tube competition-based David’s score (Fig. 5A, ρ = 0.57, p < 0.001), as well as for the fraction of active chases (Fig. 5B, ρ = 0.75, p < 0.001) and of times being chased (Fig. 5C, ρ = 0.54, p < 0.001). Notably, the fraction of active chases showed slightly higher across-round stability than tube-derived David’s score and the fraction of being chased. This is in line with active chases capturing an individual propensity to initiate this behavior, whereas competition-based social rank is a relational measure that is also influenced by the set of competitors and interaction opportunities in each reshuffled group. Nonetheless also social rank and being chased were overall stable. Across the full series, ICCs for these social measures were also in the good–excellent range, indicating high across-round stability (Fig. 5D, ICC = 0.548 to 0.890).”

      Discussion

      Line 653-664 (see also Reviewer #1, point 2)

      “These dynamic aspects of chasing, including its asymmetric initiator–recipient structure and proactive engagement, differ from the nature of tube competitions, which are incidental encounters. Together, tube-derived social rank and chasing describe complementary dimensions of social position, and, alongside other features such as clique formation [36], contribute to describe facets of a broader multidimensional social behavior. Within this complex environment, chasing emerges as a flexible behavioral propensity that is dissociable from formal social rank. Specifically, in the NoSeMaze, chasing contributes dynamically to the maintenance, negotiation, or clarification of social hierarchy structure.”

      (3) Unbalanced participation across rounds compromises stability analyses. Stability analyses (e.g., ICCs, round-to-round correlations) assume comparable sampling across individuals. However, some mice contribute 1 round, others 2, 3, 4, and even 5 rounds. This imbalance may inflate stability estimates or confound group reshuffling effects, and the rationale for variable participation is not explained.

      We thank the Reviewer for raising this point. Indeed, participation was unbalanced across rounds (see the same Reviewer #3, Major concerns 1 and new Supplementary Table S3), mainly due to missing data from occasional technical failures of the RFID detectors or the reinforcement learning water port during acquisition in some groups (for details, see Supplementary Table S2). We added this rationale behind variable participation to our Methods section (for details, see Major concern 1, ll. 981-982, “Supplementary Table S2 summarizes the number of rounds per mouse and missing data due to technical problems.”)

      We now quantify round participation (new Supplementary Table S3) and directly address potential bias from unequal sampling in two ways. First, the round-1-to-round-2 (R1–R2) stability correlations use one paired observation per mouse and therefore are not inflated by mice contributing more than two rounds. Second, ICCs were estimated from REML variance-component mixed models that account for unbalanced repeated-measures by design and use all available observations. To further rule out inflation from unequal sampling or non-random missingness, we additionally report a conservative sensitivity ICC restricted to a balanced subset including only each mouse’s first two observed sessions and only mice with both sessions (“first2-balanced”). Full-sample ICCs and first2-balanced ICCs were highly similar (new Supplementary Table S4), indicating that participation imbalance did not affect the stability estimates. For details on the changes made in the manuscript, see Reviewer #3, Major concern 1.

      Recommendations for the authors: 

      Editor's notes:

      Should you choose to revise your manuscript, if you have not already done so, please include full statistical reporting including exact p-values wherever possible alongside the summary statistics (test statistic and df) and, where appropriate, 95% confidence intervals. These should be reported for all key questions and not only when the p-value is less than 0.05 in the main manuscript.

      Readers would also benefit from noting that the mice were male in the abstract.

      We have revised the manuscript accordingly and now report exact p-values values wherever possible for the key results in the main text. Because full statistical reporting for every analysis in the main text would substantially interrupt readability, we provide the complete statistical details in a Supplementary Excel File (Supplementary Material – Systematic Statistical Reporting), including sample sizes, degrees of freedom, the number and type of permutation tests, exact p-values, and 95% confidence intervals. We also provide Extended Data Sheets for all linear-mixed effects models that account for covariates such as age and round of participation in the NoSeMaze (cf. Reviewer #3, Major Concern (1) for the additional analyses). To guide readers to these resources, we now explicitly refer to the Supplementary Material – Systematic Statistical Reporting at several points in the manuscript:

      Results

      Lines 271-274

      “Detailed statistical reporting for all analyses, including n, degrees of freedom, exact p-values, and 95% confidence intervals, is provided in the Supplementary Material - Systematic Statistical Reporting. Extended Data includes additional linear mixed-effects models controlling for potential confounding variables.”

      Lines 430-431

      “All statistical details are provided in the Supplementary Material – Systematic Statistical Reporting.”

      Additional references to the supplementary statistical reporting were inserted at lines 252-254, 562-563, and 567-568.

      Methods

      Lines 962-965

      “Full details on the statistical tests, including n, degrees of freedom, exact p-values, and 95% confidence intervals, are provided in the Supplementary Material – Systematic Statistical Reporting, as well as in the Extended Data for the LMEs accounting for different covariates.”

      We also revised the abstract and the title to explicitly state that the mice were male.

      Manuscript changes:

      Results, figure legends, and supplementary tables expanded to include full statistical reporting; abstract revised to specify male mice.

      Reviewer #1 (Recommendations for the authors):

      I would recommend being much more explicit about the reductionistic nature of the study and how the limitations are turned here to an advantage, while at the same time acknowledging the challenges of extrapolating beyond these boundaries. Most importantly, dominance should be positioned more clearly and cautiously within a framework of social behavior (and social structure) in general. The authors include cognitive tests, etc., to generate context, but this context is dependent on the same circumstances that possibly contribute to the social structure. The information that lies in the chasing behavior as an additional measured variable might be used better to provide more context.

      We thank the Reviewer for these recommendations. We have better clarified the framing to make explicit that the NoSeMaze is a controlled laboratory group-housing system. The specific conditions are now discussed in more detail in relation the observed social behaviors. We describe the tube-derived social rank as one dimension of social organization, and proactive chasing as another one. This study presents a first necessary step to understand their shared and distinguishing features. We also elaborated the Discussion to better clarify the trade-off between ecological complexity and experimental control, and to highlight that the value of the NoSeMaze for continuous, observer-independent phenotyping under standardized conditions. We now clearly state the importance to vary conditions in the system to see how the social behaviors and also their relation to non-social features changes depending on context conditions. For detailed changes, see our responses to Reviewer #1, points (1), (3), (4), and (5), and Reviewer #2, Weakness (4).

      Manuscript changes:

      Abstract, Introduction, Discussion, and Conclusion revised to clarify scope, construct interpretation, and the complementary roles of competition-based social rank and chasing.

      The precision of the description of the experimental design should be improved. When were the cohorts mixed, or did they stay separate? Which animals were old, which were young? This remained a bit confusing.

      We revised the presentation of the study design accordingly. Specifically, we clarified that the younger and older adult cohorts were run as separate experimental series and that group reshuffling occurred within, but not across, these cohorts. We also revised the study schematic in Fig. 1 and the corresponding description in the manuscript to more explicitly depict cohort structure, timing of NoSeMaze rounds, and between-round reshuffling (details are provided in the Supplementary Tables 1-3). For new age-related analyses and additional robustness checks, see our response to Reviewer #3, Major Concern (1).

      Manuscript changes:

      Figure 1 and its legend were revised to depict cohort structure, NoSeMaze rounds, and between-round reshuffling more explicitly. In addition, the Results subsection “Ecological longitudinal assessment in the NoSeMaze” were updated to clarify that the younger and older adult cohorts were run separately, that reshuffling occurred within but not across cohorts, and which animals belonged to each age-defined cohort. We also now cross-reference Supplementary Table S2 for round-specific age information and cohort composition.

      Results

      Lines 167-170

      “The study population comprised two age cohorts: younger (16-30 weeks) and older adults (55-97 weeks) (Fig. 1B). The two age cohorts were run as separate experimental series, and group reshuffling was performed within each cohort (see Supplementary Table S2).”

      It is also not fully clear across which groups stability measures were obtained: are these across all groups or within the subgroups with a given characteristic?

      We thank the Reviewer for highlighting this point. We now state explicitly that stability measures were computed across all eligible rounds, using mixed-effects models that account for repeated observations of the same mouse and for round-specific group membership. We also added a conservative sensitivity analysis restricted to a balanced first-two-session subset. Full and balanced-subsample stability estimates were highly similar, indicating that the main conclusions are robust to the participation structure. For details, see our response to Reviewer #3, Major Concerns (1) and (3).

      Manuscript changes:

      Methods and Results revised to clarify the level of analysis and model structure, as well as new supplementary robustness tables (Supplementary Tables 3-5) added.

      Minor point: In Figure 3, 18 groups are mentioned, but 19 are shown.

      We thank the Reviewer for this clarification. The apparent discrepancy arose because the group labels in Figure 3 follow the numbering of all experimental groups, whereas only groups with available tube-competition data are shown in this panel. Accordingly, group 16 is absent because no tube data were available for that group (see Supplementary Table S2), and groups 20 and 21 are likewise not included for the same reason. We have revised the figure legend and corresponding text to make this explicit and to avoid the impression of a numbering inconsistency.

      “Fig. 3: Social rank derived from incidental competitions in the integrated tube tests of the NoSeMaze.”

      “C, Box plots of metrics characterizing social hierarchy for 18 groups, including transitivity, steepness, stability, and uncertainty-by-repeatability (for details, see ‘Source Data’). Group labels correspond to original experimental group IDs. Only groups with available tube-competition data are shown. Therefore, numbering is non-consecutive (e.g., groups 16, 20, and 21 are absent; see Supplementary Table S2).”

      Reviewer #2 (Recommendations for the authors):

      (1) To better distinguish this study from previous literature, we recommend incorporating a more focused discussion (or adding to the intro) of how the findings advance our understanding of social hierarchy beyond prior works.

      We have sharpened the conceptual framing in both the Introduction and Discussion. In particular, we now distinguish more explicitly between the well-established observation that hierarchies form and remain stable within fixed semi-naturalistic groups, and the more specific question addressed here: whether an individual’s social position generalizes across changing social contexts created by repeated group reshuffling. We added additional references and also emphasize that the present study integrates continuous measurements of competition-based social rank, chasing, and reinforcement-learning features in the same individuals and environment. For details, see our response to Reviewer #2, Weakness (1).

      Manuscript changes:

      Introduction and Discussion revised to foreground conceptual novelty relative to prior semi-naturalistic work.

      (2) Given the weak associations between dominance rank and other traits such as body weight, cognitive performance, and oxytocin receptor manipulation, we suggest further clarifying what is being captured by these measures.

      We agree and have clarified this point throughout the manuscript. We now define tube-derived social rank explicitly as an operational measure based on repeated competitive outcomes from incidental dyadic tube tests, highlighting it as one dimension of social behavior.

      We also make clearer that proactive chasing is not redundant with rank, but instead captures a distinct, partly dissociable dominance-related interaction mode. For details, see our response to Reviewer #2, Weakness (4).

      We understand the question on the meaning of social rank as the correlations to body weight and cognitive performance are only punctual. We would like to mention here already that the competition-based social rank turns out to be a strong predictor of individual reactivity in a series of challenges. These works are in currently in preparation for publication and will make the relevance of these measures more clear. Related to this, we find in these studies that chasing and competition-based social rank predict different behavioral and neuronal aspects of individual reactivity.

      Manuscript changes:

      Abstract, Results, Discussion, and Conclusion revised to clarify construct interpretation and multidimensionality.

      (3) We recommend modifying the title and abstract to more clearly reflect the male-only design of the study. In addition, please indicate whether any age-related differences were observed. If age had no measurable effect, this should be stated explicitly to justify the combination of age groups.

      We revised the title and abstract to make the male-only design explicit. We also now report age-related analyses directly in the manuscript. Briefly, age was modeled by separating between-group age structure from within-group age variation. Some measures showed modest age associations in mean level, but most importantly, age-adjusted stability estimates for the core social metrics were highly similar to the original estimates, indicating that the reported stability is not driven by age differences across groups. For details, see our response to Reviewer #3, Major Concern (1).

      Manuscript changes:

      Title and abstract revised. Methods, Results, and supplementary robustness analyses expanded to report age effects explicitly.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examined whether infraslow fluctuations in noradrenaline and in heart rate are coupled and how they are affected by sleep transitions. The authors used the fluorescent NA biosensor GRAB-NE2m in the medial prefrontal cortex of mice to record extracellular NA while also recording EEG and EMG during sleep-wake episodes. They also analyzed previously published human data to reproduce relationships they found between sigma power and RR intervals in mice.

      Strengths:

      This is an impressive study with significant strengths, as it involves a rich set of data that includes not only observations of associations between heart rate and noradrenergic dynamics but also optogenetic manipulation of the locus coeruleus. Human data is presented to show parallels in the association between sigma power during sleep and phasic heart-rate bursts.

      We thank Reviewer #1 for their thoughtful, detailed, and constructive evaluation of our manuscript. We appreciate their recognition of the strengths of the study, particularly the integration of noradrenergic recordings, optogenetic manipulation, and cross-species analyses. We are especially grateful for the reviewer’s careful attention to clarity, experimental interpretation, and control comparisons. The comments have helped us sharpen the framing of our hypotheses, clarify causal claims, improve statistical reporting, and better explain our closed-loop approach and heart rate analyses. We have addressed each point in detail below and believe that the revisions substantially strengthen the manuscript.

      Weaknesses:

      (1) Language could be clearer and more precise. As detailed below, in both the introduction and the discussion, the way the hypotheses and study objectives are described could use some revision to be more precise and accurate.

      Thank you for this helpful comment. We have sharpened the description of the study objectives, hypotheses, and interpretation of the findings to better distinguish between what was directly tested, what was inferred, and what remains speculative. We revised the language throughout these sections to improve clarity, accuracy, and overall readability

      (1A) In the introduction on p. 4: The overarching question is framed as "could the peripheral autonomous systems be a read-out of the central LC-NE system and thus be a biomarker of memory consolidation and LC dysfunction?" This gives the impression that the LC function would be the main influence on peripheral autonomous systems. There are, of course, many influences on peripheral autonomous systems, so it would be advisable for the authors to be more specific here about what signal(s) in particular would be predicted to be sensitive markers of LC function.

      Thank you for this important point. We agree that heart rate reflects the integrated output of multiple autonomic mechanisms and should not be interpreted as being exclusively driven by LC activity. Cardiac dynamics arise from the balance between sympathetic and parasympathetic influences, which themselves are regulated by several central and peripheral systems. In addition, recent work shows that the infraslow oscillations observed during NREM sleep are not restricted to norepinephrine alone but also involve other neuromodulatory systems, including acetylcholine and serotonin (e.g., Teng et al., PNAS 2025, Kjaerby et al., iScience, 2026). Our intention was therefore not to imply that the LC is the sole driver of peripheral autonomic dynamics. We have revised our overarching questions to make them more specific: Please see new text below:

      (Introduction, page 4/5). “The sympathetic and parasympathetic autonomic nervous system are involved in HRV, which is conventionally analyzed across three primary frequency bands: high frequency (HF) HRV, low frequency (LF) HRV, and very low frequency (VLF) HRV (Berntson et al., 1997). Due to the frequency overlap with VLF HRV, we wondered if central infraslow NE dynamics could be linked to this poorly understood HRV indicator. Furthermore, are infraslow NE fluctuations directly reflected by HRV under different physiological states or does LC–HR coupling scale differently with LC output? Specifically, if infraslow NE oscillations display faster frequencies - as occurs during sleep fragmentation - will cardiac dynamics exhibit corresponding changes? Conversely, given that stronger infraslow NE dynamics correlate with memory consolidation through their regulation of sleep spindles, could the peripheral autonomic signatures provide an accessible cross-species biomarker of spindle-dependent memory consolidation? Addressing these questions could help bridge mechanistic insights into LC-mediated sleep regulation with established HRV metrics used in human physiology.”

      (1B) In the discussion on p. 12: "In this study, we leveraged real-time measurements of mPFC NE levels and HR measurements from EMG recordings in mice to investigate the causal link between the two variables with high temporal resolution in freely moving sleeping mice, with similar inspection in humans." To test the causal link between mPFC NA levels and HR measures, the study would manipulate NA levels just in the mPFC and not elsewhere in the brain. However, in this study, the manipulation occurred in the LC, and so there would be broad cortical changes in NA levels. Thus, it could be that LC activity causes HR changes via a non-PFC pathway.

      We thank the reviewer for this important comment. Indeed, mPFC NE is merely a readout of LC activations and we expect that NE in other brain regions would show the same patterns. Indeed, mPFC NE is not expected to provide any causal link to heart rate. We have revised added a sentence to the results section and also changed the initial summary part of the discussion to reflect this better.

      (Results, page 6). “mPFC was selected as a representative cortical readout of LC-mediated norepinephrine dynamics, as infraslow NE fluctuations are coordinated across widespread brain regions.”

      (Discussion, page 14/15). “Variability in HR is a non-invasive biomarker of autonomic nervous system function and is frequently disrupted in ageing and Alzheimer’s disease. Here, by combining real-time measurements of mPFC NE dynamics with simultaneous HR recordings in freely sleeping mice, we demonstrate that HR closely tracks the infraslow phasic activity of the LC–NE system.”

      (2) Comparisons with the control condition need further development.

      (2A) While the authors did include a key YFP control condition, in the main text no direct statistical comparison between the closed-loop optogenetic stimulation (ChR2) condition and the YFP control condition was reported. (It was reported in Supplementary Figure 2c-d.) Instead, in the main text, the authors only reported that the effects of stimulation were significant in the closed-loop condition and not in the control. However, that is not the same as demonstrating that the two conditions significantly differed from each other, and it is the direct test that is important for the conclusions, so it seems important to include this result in the main presentation.

      We thank the reviewer for this important point and agree that direct statistical comparisons between ChR2 and YFP conditions are important for interpretation. These comparisons were performed and are shown in Supplementary Figure 2c–d, but we acknowledge that this was not sufficiently emphasized in the main text. We are now more clearly referring to this comparison in Result section:

      (Results, page 9). “The magnitude of pre-stimulation NE descent and post-stimulation NE ascent was reduced as the thresholds increased, indicating less pronounced NE dynamics as LC stimulation became more frequent (Fig. 2e, for direct comparison with YFP control, see Suppl. Fig. 2c-d).”

      Our rationale for prioritizing the within-animal threshold comparisons in the main figure was that the central experimental question concerned how progressive shifts in infraslow NE oscillatory frequency influence the NE–HR relationship. Because variability in viral expression levels (both NE sensor expression and LC opsin expression) introduces substantial between-animal variability, we considered within-animal comparisons across threshold conditions to provide the most informative representation of how changes in LC-driven NE dynamics alter cardiac responses. That said, we agree that highlighting the direct ChR2 versus YFP comparison is important for the overall interpretation. As mentioned, the reference to Suppl. Fig. 2c– d, where the between-group analyses are visualized are now clearly referred to.

      (2B) In addition, the authors should address the issue that the pre-stimulation NE was consistently significantly lower in the YFP condition than in the ChR2 condition (see Supplementary Figure 2c), which is a potential confound.

      We thank the reviewer for bringing up this important point. We agree that differences in stimulation timing between ChR2 and YFP animals could complicate interpretation in a closed-loop design and appreciate the opportunity to clarify this aspect of the experiment.

      In the ChR2 condition, animals were exposed to repeated optogenetic LC activation designed to mimic progressively faster infraslow NE dynamics. Such repeated stimulation is expected to produce a gradual elevation in tonic NE levels across the recording session, which explains the higher pre-stimulation baseline relative to YFP controls. We acknowledge that elevated tonic NE levels could introduce additional physiological effects. For example, higher NE tone would be expected to increase α2mediated autoinhibitory feedback on LC neurons and presynaptic NE release. Within the LC itself, we expect that optogenetic stimulation would largely override such effects due to the strong Na+-mediated depolarization induced by ChR2 activation. However, NE release in downstream regions such as the mPFC may be influenced to some extent by elevated tonic noradrenergic tone. Importantly, such feedback mechanisms would likely also occur under physiological conditions characterized by elevated LC activity, such as stress or sleep fragmentation. Because the goal of our stimulation paradigm was to model progressively faster infraslow NE dynamics under physiologically relevant conditions, we believe this feature of the manipulation may in fact increase the translational relevance of the model. We specifically address this elevation in Figure 3, where we show that very rapid stimulation regimes are accompanied by signs of compensatory cardiovascular regulation, likely reflecting baroreceptor-mediated responses to sustained increases in heart rate.

      We agree that the precise contribution of elevated tonic NE to the overall manipulation cannot be fully disentangled in the present study. We therefore avoid overinterpreting these effects and have instead added text to the manuscript acknowledging this consideration without extensive speculation.

      (Results, page 9) “Since pre-stimulation NE baseline levels progressively became higher in the ChR2 condition compared with YFP controls (Fig. 2c+e, Supplementary Fig. 3a-f), it demonstrates that they arise from stimulation-dependent modulation of noradrenergic tone rather than nonspecific signal drift. As a result, the pre-stimulation state at higher thresholds differed between ChR2 and control conditions, which should be considered when interpreting the immediate effects of laser stimulation across thresholds.”

      (2C) Direct comparison of the strengths of correlations shown in Figure 2h vs. Supplementary Figure 2f should be included. Currently, we see relatively weak correlations in both ChR2 and YFP conditions, and it is not clear if the relationships differ in the control. It seems they are still present in the control condition but weaker which would contradict the apparently broad claim on p. 7 that "No such effects were present in the control condition" (it is not entirely clear whether this claim refers to all effects discussed in the figure or just a subset - this language should be clarified).

      We thank the reviewer for this important comment and agree that the original wording could be interpreted as implying a complete absence of an NE–RR relationship in the YFP condition. To address this concern, we directly compared the strength of the NE– RR relationship between ChR2 and YFP animals. Please see Supplementary Figure 3h.

      Using a linear mixed-effects model that accounted for repeated measurements within animals, we found that the slope of the NE–RR relationship was significantly steeper in ChR2 animals than in YFP controls (all thresholds: slope difference = 1.915, p < 0.0001; thresholds −15, −10, and −5 only: slope difference = 2.367, p < 0.0001). Consistent with this result, comparison of Pearson correlations using Fisher's r-to-z transformation also indicated significantly stronger coupling in ChR2 animals than in YFP controls (see Statistics in the Supplementary File).

      These analyses demonstrate that an inverse relationship between NE and RR is present under physiological conditions in YFP animals, but that optogenetic LC activation substantially strengthens this coupling. We have revised the corresponding text to clarify this:

      (Results, Page 9): “An inverse NE–RR relationship was present in both ChR2 (Fig. 2h) and YFP (Suppl. Fig. 3f) groups but was significantly stronger in ChR2 animals than in YFP controls (Suppl. Fig. 3h).”

      (2D) Did the YFP controls vs. ChR2 animals show any differences in the number of NA states that triggered stimulation in the closed-loop system? With ChR2 animals, stimulation changes NA, which could change future triggering. In YFP animals, nothing changes NA (other than natural fluctuations), so the dynamics of stimulation timing could diverge between groups in a way that complicates interpretation. Specifically, if ChR2 stimulation raises NA and prevents future threshold crossings, ChR2 animals may end up receiving fewer subsequent stimulations than YFP animals (or a different temporal clustering). If the number or pattern of stimulation differed in two groups, it would be important to have a yoked control where matched animals get the same stimulation pattern but not triggered by their own NA.

      We thank the reviewer for this important point. We agree that differences in stimulation timing between ChR2 and YFP animals could complicate interpretation in a closed-loop design.

      Importantly, stimulation triggering was based on relative declines in NE fluorescence calculated against a rolling 2-minute baseline, rather than absolute NE levels. Thus, as tonic NE levels gradually increased in ChR2 animals, the threshold adapted accordingly, reducing the likelihood that elevated baseline NE alone would prevent future triggering. Instead, stimulation continued to occur when NE declined relative to the recent baseline, thereby preserving the infraslow closed-loop structure.

      The number of stimulation events across thresholds is already reported in the manuscript (Methods, p. 30 and corresponding figure legends), but we have now indicated more clearly in the result section where to find the information:

      (Results, Page 9). “Mean traces of NE and RR were aligned to LC stimulation onset (Fig. 2c-d, for number of laser stimulations see Fig. 2 legend or Methods).”

      (Methods, Page 31). “For the LC activation-related analysis, 108 events were found for Threshold -15 (11 of these being YFP), 260 events for Threshold -10 (55 of these YFP), 777 events for Threshold -5 (296 of these YFP), 1,444 events for Threshold 0 (510 of these YFP), and 1,148 events for Threshold 5 (377 of these YFP) across ten animals (four being YFP).”

      While the total number of events was lower in YFP animals, this is expected in part due to the smaller group size (4 YFP vs. 6 ChR2 animals included in this analysis). Furthermore, because ChR2 stimulation increased NE levels by design, more pronounced subsequent declines in NE may have modestly facilitated additional threshold crossings.

      Nevertheless, the overall temporal structure of the stimulation paradigm remained comparable across groups, and YFP animals underwent the same closed-loop stimulation protocol. Our primary comparison was mainly based on shifts in within-animal NE oscillatory frequency and how this change would impact the connection to HR. Thus, we believe the present control condition appropriately addresses the central question of whether optogenetic LC activation are able to conduct a continuum of NE oscillations.

      (3) Some more discussion/explanation of the rationale for the closed-loop approach and how it influences how we should interpret the results could be useful. For instance, currently, it is not clear whether LC stimulation needs to be timed after an NA dip to yield the effects seen.

      We thank the reviewer for pointing this out. The rationale for the closed-loop LC stimulation approach was to test whether the relationship between infraslow LC–NE dynamics and heart rate is maintained only under physiological infraslow conditions or whether it breaks down when the rhythm becomes progressively faster, as occurs during sleep fragmentation and other high-arousal states. Specifically, we asked whether heart rate continues to track LC–NE fluctuations as the infraslow rhythm shifts toward higher frequencies, thereby assessing its utility as a potential biomarker of disrupted restorative sleep.

      To address this while preserving the intrinsic temporal structure of infraslow LC activity, we implemented a closed-loop strategy in which stimulations were triggered following defined declines in the NE signal, using a rolling preceding 2-minute window as baseline. This allowed LC activation to occur during the descending phase of the endogenous infraslow cycle, maintaining its physiological phase structure while systematically increasing its effective frequency. By progressively relaxing the decline threshold, stimulations were triggered earlier in the cycle, thereby compressing the infraslow period in a controlled manner.

      Importantly, the intention was not to test whether LC stimulation specifically needs to occur after an NE dip to elicit the observed effects. Rather, triggering stimulation during the decay phase provided a way to accelerate the infraslow rhythm without disrupting sleep through indiscriminate stimulation. This enabled us to examine whether the coupling between LC–NE dynamics and heart rate remains stable under increasingly rapid infraslow regimes. Our results indicate that this relationship weakens at higher infraslow stimulation frequencies, suggesting that heart rate reliably reflects physiological LC–NE oscillations but becomes less tightly coupled when the rhythm is compressed beyond its normal range.

      We have clarified this rationale in the revised manuscript by adding the below section in the result section.

      (Result, page 8/9). “This approach enabled controlled compression of the infraslow NE cycle by triggering LC activation during the descending phase of the endogenous NE signal, thereby increasing the effective oscillatory frequency while preserving the temporal structure of physiological LC–NE dynamics. This strategy allowed us to test whether heart-rate responses continue to track LC-driven NE fluctuations as the infraslow rhythm becomes progressively faster.”

      (4) The section on heart rate decelerations is hard to follow. In particular, I was not sure how to interpret Figure 3f-j. For Figure 3f, what does the middle line represent? The laser onset or the max RR value after laser onset? What is the baseline that is used to correct the values to obtain amplitudes? If it is the whole period before the maximal RR value or the laser onset, wouldn't baseline values differ significantly across conditions and so potentially account for differences seen between conditions in the reported HR decelerations? Larger HR decelerations may be seen in conditions with higher HR simply as a regression to the mean phenomenon.

      We thank the reviewer for this feedback and agree that additional clarification of Figure 3f–j is warranted.

      For Figure 3f, the central line represents the peak RR value (maximal heart-rate deceleration) identified within the 2–7 s window following laser onset, rather than the laser onset itself. We realize this was not sufficiently clear and have revised the figure and corresponding Results text to clarify this point.

      Regarding baseline correction, RR amplitudes were calculated as the difference between the RR at peak deceleration and the mean RR during the 8–10 s period preceding the RR peak, as described in the manuscript. Thus, the baseline was defined locally for each event and was not based on the entire pre-laser period or stimulation onset. We chose this approach to account for shifts in baseline heart rate across conditions and to capture the relative magnitude of the deceleration response rather than absolute RR values.

      We appreciate the reviewer’s point regarding potential regression-to-the-mean effects, particularly in conditions with higher baseline heart rates. This is an important consideration. However, because the amplitude measure was baseline-corrected on an event-by-event basis, we believe the reported differences are unlikely to be explained solely by higher pre-stimulation heart rate. At the same time, our findings clearly show that elevated baseline heart rate influence the dynamic range of deceleration responses. Our findings show that under physiological conditions with elevated HR, larger heart-rate fluctuations would also contribute to increased HRV, which is often interpreted positively, despite potentially reflecting fragmented or dysregulated sleep states in this context.

      To improve readability, we have revised the figure and associated text.

      (Results, page 10). “To quantify the HR decelerations that happened after LC activation, we took the maximal RR value (so slowest HR) 2-7 s after LC stimulation and baseline corrected the value to the mean RR during the 8–10 s period preceding the RR peak to obtain their amplitude.”

      (5) The findings regarding LC suppression could be further clarified.

      (5A) Page 8: "observed a response in NE decline" - please be more precise. Did NE decline more or less?

      We thank the reviewer for this suggestion and agree that the original wording was imprecise. To clarify the direction and nature of the response, we have revised the text to state:

      (Results, page 11). “...we observed a gradual NE decline sustained throughout the laser period that was not observed in the YFP condition…”

      (5B) It would be helpful to also show the correlation between NE and RR in the control (YFP) condition and whether there were any differences between YFP and Arch conditions (Figure 4e).

      We thank the reviewer for this suggestion. We have now added the corresponding YFP correlation to Figure 4e. In the YFP group, the relationship between NE and RR showed a similar negative trend but did not reach statistical significance (p = 0.053). To directly assess whether the NE–RR relationship differed between Arch and YFP animals, we performed both a linear mixed-effects analysis and a Fisher r-to-z comparison.

      Neither analysis revealed a significant difference between groups. The linear mixed-effects model showed no significant Group × NE interaction (slope difference = 0.994, p = 0.51), indicating that the NE–RR coupling was not altered by LC suppression. Similarly, Fisher's r-to-z comparison found no significant difference between the correlations (p = 0.42).

      We believe this result is consistent with the relatively modest nature of the LC suppression paradigm. While Arch stimulation produced a clear reduction in NE levels, it did not induce a large shift in the overall NE–RR relationship. Instead, the data suggest that heart-rate responses remain coupled to noradrenergic fluctuations under both physiological conditions and during mild LC suppression. We have added the YFP data and clarified this interpretation in the revised manuscript.

      (Results, page 12): “A similar negative relationship was observed in YFP controls (Fig. 4e), and the strength of the NE–RR association did not differ significantly between Arch and YFP animals (Suppl. File, Statistics), suggesting that LC suppression did not substantially alter the underlying coupling between these measures.”

      (5C) This sentence took me multiple readings to understand - it would be helpful to rewrite to make it clearer: "indicating that, while HR generally did not respond strongly to LC suppression, the variability in RR responses was dependent on NE changes to the suppression (Figure 4e)."

      We agree with the reviewer that this phrasing is hard to understand and we have optimized for better clarity. Please see new version below:

      (Results, page 11/12). “Notably, despite the absence of a robust group-level HR effect, NE and RR responses remained negatively correlated across trials (Fig. 4e), indicating that HR dynamics continued to track the magnitude of noradrenergic suppression at the individual-response level”.

      (5D) The two colors in Figure 4 are similar and hard to distinguish.

      We agree that the colors are hard to separate and have altered them to make them easier to separate.

      (5E) The correlations shown in Figure 4j seem to be driven by just two of the cases. Are the effects significant when outliers are removed?

      We thank the reviewer for raising this point. To assess whether the observed correlation in Figure 4j was disproportionately driven by a small number of data points, we performed a formal outlier analysis using the ROUT method (Q = 1%). This analysis did not identify any statistical outliers in the dataset. Therefore, we did not have an objective basis for excluding any observations from the analysis.

      (5F) Page 10: Were there any differences in memory performance between the Arch and YFP conditions?

      We thank the reviewer for this question. The memory experiments were based on a previously published dataset (Kjaerby, Andersen et al., 2022), in which the primary objective was to assess the effect of LC suppression on sleep spindle dynamics and memory consolidation. In the present study, we performed an additional analysis by extracting heart-rate (RR) measures from these recordings to evaluate whether cardiac responses could serve as a biomarker of LC-mediated noradrenergic regulation.

      However, due to technical limitations (EMG recording often suffers from noise) in extracting reliable RR signals from all animals in this dataset, the number of subjects available for this secondary analysis was reduced. All these considerations are described in Methods/Mice. As a result, we were not sufficiently powered to perform a direct statistical comparison of memory performance between Arch and YFP groups based on RR measures alone. Instead, we examined whether RR responses to LC suppression predicted behavioral performance across animals. When pooling Arch and YFP conditions, we observed a correlation between the magnitude of the RR response and subsequent memory performance, suggesting that heart-rate dynamics reflect noradrenergic modulation relevant for memory consolidation.

      To avoid overinterpretation, we have therefore limited our conclusions to reporting this association rather than making direct group-level comparisons between Arch and YFP animals, and we have clarified this point in the revised manuscript.

      (Results, page 13): “Interestingly, across pooled Arch and YFP animals, larger RR increases following LC suppression were associated with better subsequent memory performance. Additionally, RR and NE responses to LC suppression were negatively correlated indicating that animals showing stronger NE reductions also exhibited larger RR changes. Together, these findings suggest that heart-rate dynamics covary with noradrenergic responses during sleep and may reflect physiological processes relevant for sleep-dependent memory consolidation.”

      (5G) Page 10: "We found a correlation between RR responses to LC suppression and sigma power, suggesting that a stronger HR reduction response is linked to higher spindle power." It should be noted in the text that the correlation was not specific to sigma (it was also seen for theta and beta, Figure 4i).

      We agree with reviewer that this should be highlighted. We have changed the sentence:

      (Results, page 12). “Furthermore, we found a correlation between RR responses to LC suppression and sigma power, suggesting that a stronger HR reduction response is linked to higher spindle power; similar correlations were also observed in the theta and beta frequency ranges (Fig. 4h-i).”

      (6) It is not clear which of the sigma power and RR interval findings do/do not exactly line up between the mice and humans. It could be helpful to have a table comparing them. For instance, was the finding in humans that pre-HRB sigma power was positively associated with slowing in heart rate after the HRB also seen in mice? Was there evidence in mice (as seen in the human sample) that sleep-dependent memory improvement was associated with pre-HRB sigma power?

      We thank the reviewer for this thoughtful comment and agree that the cross-species comparisons could be communicated more clearly. Our intention was not to imply exact one-to-one correspondence between all mouse and human findings, but rather to examine whether central–autonomic coupling surrounding phasic heart-rate events shows conserved features across species while acknowledging species-specific physiological differences.

      Importantly, the mouse and human analyses were designed to address related but not identical questions. In mice, we leveraged optogenetic LC suppression to probe a more causal relationship between noradrenergic activity, heart-rate slowing, spindle-related dynamics, and memory consolidation. Previous work using this dataset demonstrated that 2-minute LC suppression robustly enhances spindle density and that spindle enhancement correlates with improved memory performance. In the present study, we therefore asked whether heart-rate slowing covaries with this LC-mediated spindle/memory relationship, supporting HR as a potential biomarker of these restorative processes.

      By contrast, causal manipulation of LC activity is not feasible in humans. Instead, we focused on naturally occurring HRBs as putative downstream signatures of phasic LC– NE activity, motivated by our mouse findings that NE increases precede HR accelerations. We observed conserved coupling between sigma activity and HR dynamics across species, although the temporal profile differed, with sigma activity occurring closer to the HRB in humans than in mice. These temporal differences may reflect species-specific differences in cardiac and sleep physiology.

      Regarding the reviewer’s specific questions, the positive relationship between pre-HRB sigma power and post-HRB heart-rate slowing was tested in humans, where this metric showed the strongest relationship to behavioral outcome. We did not directly test the same measure in mice because, unlike humans, HR recovery following HRBs did not show a pronounced baseline shift (Fig. 5b), limiting the interpretability of this comparison. Similarly, we did not directly correlate pre-HRB sigma power with memory performance in mice, as the more causal LC suppression paradigm already demonstrated a spindle–memory relationship in this species and was the focus of our mechanistic analysis.

      To reduce confusion, we have revised the Results conclusion to make a clearer overview:

      (Results, page 14): “In conclusion, mice and humans displayed evidence of conserved autonomic-central coupling, reflected in coordinated HR and sigma power dynamics surrounding phasic cardiac events.

      However, the temporal relationship between these events differed across species, likely reflecting differences in sleep and cardiovascular physiology. In mice, causal manipulation of LC activity demonstrated that heart-rate dynamics covary with LC-mediated noradrenergic and spindle-related processes linked to memory consolidation. In humans, sigma power preceding HR bursts was associated with both post-HRB heart-rate slowing and sleep-dependent memory improvement, suggesting that autonomic–central coupling surrounding HR events may provide a translational marker of restorative sleep processes (Fig. 5n).”

      (7) Page 18: It is not clear if the sex of mice was balanced across controls and optogenetics groups.

      We thank the reviewer for this important comment and agree that the sex distribution should be reported more clearly. These experiments relied on the availability of animals from the heterozygous TH-Cre transgenic line, and given the relatively small cohort sizes, perfect balancing across sex and experimental groups was not always feasible.

      For the LC activation experiments, the sex distribution was: YFP: 2 male / 2 female; ChR2: 4 male / 2 female. For the LC suppression experiments, the distribution was: YFP: 4 female; Arch: 3 female / 1 male.

      Although the groups were not perfectly sex balanced, we had no strong reason to expect robust sex-dependent differences in the physiological effects of these optogenetic manipulations, particularly given the relatively strong and acute nature of the intervention. At the same time, we acknowledge that the present study was not powered to assess sex as a biological variable, and subtle sex-dependent effects therefore cannot be excluded. To improve transparency, we have now clarified the sex distribution in the Methods/Mice section.

      Reviewer #2 (Public review):

      Summary:

      The major part of this study reproduces previously published findings in both mice and humans and provides incremental analyses on these findings. In essence, the work reaffirms the presence of coordinated infraslow fluctuations in sigma power and heart rate during NREM sleep. It further confirms previous findings that coordination depends on noradrenaline-releasing neurons in the locus coeruleus. Also supporting previously published work in mice and humans, the authors describe a link between the strength of these infraslow fluctuations and memory consolidation in mice and humans.

      Strengths:

      The authors successfully replicate key previously reported phenomena across both mice and humans. Confirmatory studies and demonstrations of reproducibility are essential for progress in neuroscience. To maximize their value, such studies should clearly acknowledge their confirmatory nature and carefully situate what, in their view, are novel results, going beyond existing literature.

      Weaknesses:

      The authors' interpretation of their data needs to be revised. Many of their claims regarding the mechanistic basis of their findings and the predictive value of their correlative datasets are not supported by the available evidence.

      In the present manuscript, several citations of literature on the work they reproduce lack precision or completeness, which reduces transparency and obscures how the reported findings relate to previously established results.

      We thank Reviewer 2 for the thoughtful comment regarding positioning of our findings relative to the literature, and caution in mechanistic interpretation. In response, we have revised the Introduction, Results, and Discussion to more clearly acknowledge foundational studies in this area and to better clarify how the present work extends beyond them.

      We agree that prior work has demonstrated infraslow coupling between sigma activity, norepinephrine (NE) dynamics, and heart rate (HR), and has established a role for the locus coeruleus (LC) in coordinating these oscillations. However, cardiac measures in these studies were typically treated as secondary observations rather than as primary experimental targets. A central goal of the present study was therefore to provide a systematic and mechanistically grounded characterization of NE-mediated HR dynamics during sleep across multiple timescales, including infraslow oscillations, sleep–wake transitions, and causal manipulations of LC activity.

      Importantly, we also aimed to relate infraslow HR fluctuations to the very-low-frequency (VLF) component of heart rate variability (HRV), which remains comparatively under-characterized and mechanistically unresolved in the clinical HRV literature. By linking LC activity, NE dynamics, and HR fluctuations across behavioral states, our findings provide a biologically grounded framework that may help explain this component of HRV.

      A second major objective of the study was translational. Because direct LC recordings are not feasible in humans, we asked whether cardiac dynamics alone could reflect the infraslow, memory-consolidating potential of sleep and thus serve as a noninvasive biomarker. By directly manipulating LC activity and demonstrating corresponding changes in HR dynamics, our results strengthen the mechanistic rationale for using HRV—particularly its VLF component—as an accessible proxy of LC-dependent sleep physiology.

      We therefore respectfully disagree with the suggestion that the present study does not provide novel insight. Rather, the revised manuscript now more clearly emphasizes that our contribution lies in (i) systematically characterizing NE-dependent HR dynamics across sleep states, (ii) linking these dynamics to the poorly understood VLF component of HRV, and (iii) establishing a causal and translational framework for using cardiac measures as markers of LC-mediated sleep processes.

      We hope the reviewer finds that the revised Introduction and Discussion better highlight both the existing literature and the specific advances provided by the present work.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I have been convinced by discussions about the replicability crisis that it should be a standard practice to share data upon publication in a publicly accessible online repository such as Open Science Framework or OpenNEURO. Making data available can increase the impact of the research study. Simply stating "data available upon request" as done in the current draft is not sufficient as, unfortunately, when data are not shared upon publication in a public repository, it can be impossible to gain access by request from the researchers - the majority of requests from other researchers to obtain data are not complied with (e.g., Vanpaemel, Vermorgen, Deriemaecker, & Storms, 2015; Wicherts, Bakker, & Molenaar, 2011).

      We fully agree with the reviewer about the importance of data sharing for transparency, reproducibility, and maximizing the impact of research. Consistent with these principles, we will make the human dataset publicly available on the Open Science Framework upon publication and have added the link to this repository in the manuscript under Data Availability (https://osf.io/g6emj/). With respect to the mouse dataset, we respectfully note that this dataset is currently the subject of multiple planned and ongoing analyses that extend beyond the scope of the present manuscript. Releasing these data publicly at this stage could compromise these efforts and lead to potential misinterpretation prior to completion of the full analytic pipeline. For this reason, we believe that it would not be appropriate to share the mouse data in a public repository at this time. However, we remain committed to transparency and will make the mouse data available upon reasonable request during this period, with the intention of publicly releasing the dataset once the planned analyses are complete.

      (1) p. 4: There is some orphan text at the top of the page ("marker of Alzheimer's disease. Furthermore, maintaining LC neural density prevents neurodegeneration (13). Given the reported reduction in HRV in aging and Alzheimer's disease, suppressed").

      Thank you. We have removed the orphan text.

      (2) p. 22: What does NFR refer to?

      Novel-to-familiar ratio. We have removed the abbreviation from the main text. It is now only used in the figure.

      (3) I might have missed this information, but it was not clear to me how the epochs to be analyzed were selected, and for the mice, how much wake vs. sleep they included.

      We thank the reviewer for this comment and apologize that the epoch selection criteria were not sufficiently clear. They were mentioned under Methods. For the optogenetic analyses in mice, epochs were selected based on NREM sleep including microarousals to ensure that physiological responses were evaluated within stable sleep conditions while allowing for natural brief interruptions.

      For the LC activation (ChR2) experiments, stimulation epochs were included only if NREMinclMA began at least 30 s before laser onset and continued for at least 30 s after laser onset. For the LC suppression (Arch) experiments, the criterion was similarly ≥30 s of NREMinclMA prior to laser onset, but extending ≥60 s following laser onset to accommodate the longer suppression response profile.

      Thus, analyses were intentionally restricted to sleep periods, and wakefulness was not included except where it emerged naturally as an outcome of the manipulation or transition under investigation. We have now clarified these inclusion criteria in the Methods/ Event marker selection section to improve transparency.

      (4) In the Supplement, Figure 3 appears before Figure 2.

      Thanks. We have corrected it.

      Signed, Mara Mather

      Reviewer #2 (Recommendations for the authors):

      The authors' interpretation of their data needs to be revised. Many of their claims regarding the mechanistic basis of their findings and the predictive value of their correlative datasets are not supported by the available evidence.

      There are three major directions in which this study would need to be revised.

      Part 1 - Literature citations of both mouse and human literature need to be revised, and citations placed in a manner that accurately reflects what has been previously done. In detail:

      (1.1) The statement on p. 4 regarding "the extent to which phasic infraslow NE fluctuations ...is not well understood" disregards a previous publication in which closed-loop optogenetic stimulation of LC was already shown to regulate HR variations on the infraslow time scale (10.1016/j.cub.2021.09.041).

      We thank the reviewer for this important point. The study by Osorio-Forero et al. (2021) was already cited as ref. 4 in our original manuscript and in the discussion, we specifically highlighted this study: ‘These findings build on Osorio-Forero et al. (35) as well as other studies’; however, we agree that our wording did not sufficiently emphasize its key finding that closed-loop optogenetic manipulation of LC activity can coordinate infraslow heart-rate fluctuations and spindle clustering during NREM sleep.

      We have revised the relevant paragraph in the introduction to explicitly acknowledge that ref. 4 demonstrated causal coordination between LC activity, sleep spindle dynamics, and heart-rate fluctuations. We have also refined our statement of the knowledge gap to clarify which novel questions our study addresses. We believe these revisions more accurately position our work within the existing literature and clearly distinguish our contributions from prior studies.

      We have updated the introduction in several places to address these comments:

      (Introduction, page 4). “It has previously been reported that HR fluctuates at similar infraslow frequencies as NE fluctuations and sigma power in mice (Lecci et al., 2017; Osorio-Forero et al., 2021) and that phasic HR fluctuations correlate with infraslow changes in pupil diameter (Carro-Domínguez et al., 2025), a proxy for changes in NE levels (Murphy et al., 2014; Reimer et al., 2016). Importantly, optogenetic manipulation of the LC has demonstrated that infraslow LC activity coordinates sleep spindle clustering and heart rate fluctuations during NREM sleep (Osorio-Forero et al., 2021). Together, these findings support a functional coupling between the central LC–NE system and peripheral cardiac dynamics, further supported by findings that HR increases accompany MAs during NREM sleep (Carro-Domínguez et al., 2025;

      Osorio-Forero et al., 2025).”

      (Introduction, page 4/5). “Due to the frequency overlap with VLF HRV, we wondered if central infraslow NE dynamics could be linked to this poorly understood HRV indicator. Furthermore, are infraslow NE fluctuations directly reflected by heart-rate variability under different physiological states or does LC–HR coupling scales differently with LC output? Specifically, if infraslow NE oscillations display faster frequencies - as occurs during sleep fragmentation - will cardiac dynamics exhibit corresponding changes? Conversely, given that stronger infraslow NE dynamics correlate with memory consolidation through their regulation of sleep spindles, could the peripheral autonomic signatures provide an accessible cross-species biomarker of spindle-dependent memory consolidation? Addressing these questions could help bridge mechanistic insights into LC-mediated sleep regulation with established HRV metrics used in human physiology.”

      (1.2) In this same published paper, optogenetic stimulation of LC was already used to "determine the causal relationship..." (p.6). However, the authors do not cite these data.

      We thank the reviewer for this comment. As noted in our response to Comment 1.1, ref. 4 was already cited in the Introduction, and we have now revised that section to more explicitly emphasize this paper. In addition, we have modified the wording in the Results section.

      (Results, page 8). “After finding the inverse correlation between NE and RR in natural sleep transitions, we next sought to further characterize the causal influence of LC activity on HR dynamics, using a closed-loop optogenetic approach to modulate NE oscillatory frequency during NREM sleep.”

      (1.3) The authors' speculation about LC-induced sympathetic and parasympathetic actions is premature: this study does not provide pharmacological experiments in this direction. However, two published studies implied a parasympathetic mechanism linked to infraslow fluctuations of LC activity (10.1016/j.cub.2021.09.041, 10.1016/j.cub.2017.12.049). These findings should be appropriately cited. While HR decelerations may reflect compensatory autonomic responses, there is no direct evidence in the present study that these effects are sympathetically mediated. An alternative, and equally plausible, interpretation is enhanced parasympathetic activity. This distinction is particularly important given that the observed increases in mean HR during LC stimulation cannot distinguish between reduced parasympathetic activity and increased sympathetic drive.

      We thank the reviewer for raising this important point and agree that the current study does not provide direct mechanistic evidence to disentangle sympathetic versus parasympathetic contributions to LC-mediated heart rate regulation. This was not the intention of our study; rather, the relevant discussion section was meant to provide mechanistic interpretations and hypotheses based on the observed physiology. To avoid overstating our conclusions, we have revised the wording to more clearly emphasize the speculative nature of these interpretations. Furthermore, we have added the suggested references demonstrating that muscarinic blockade reduces heart rate and pupil fluctuations during sleep, which support the possibility of a parasympathetic contribution to infraslow LC-related dynamics.

      (Discussion, page 16/17). “Prior findings demonstrate that LC is linked to the autonomic nervous system. Stimulation of LC projections decrease parasympathetic cardiac vagal activity (Wang et al., 2014) and also influences sympathetic output through direct projections to the preganglionic cells in the sympathetic nervous system (Karemaker, 2017; Nygren and Olson, 1977; Samuels and Szabadi, 2008). While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014), there is a clear parasympathetic component (Taylor et al., 1998). This combined with the ability of pharmacological blockage of the parasympathetic system to block infraslow oscillations of HR (Osorio-Forero et al., 2021) and pupil diameter during sleep (Yüzgeç et al., 2018), led us to expect a slowing of HR during LC suppression due to parasympathetic disinhibition.”

      (1.4) It is not clear why prefrontal NE signals were associated with HR fluctuations. Literature evidence indicates other brain areas that are functionally more directly linked to autonomous fluctuations.

      We thank the reviewer for raising this point. We do not speculate that mPFC NE activity is causally linked to heart rate fluctuations or that the mPFC directly mediates the observed autonomic dynamics. Instead, mPFC NE signaling was used as an experimentally accessible readout of LC activity. Importantly, accumulating evidence suggests that infraslow NE oscillations are globally coordinated phenomena that are expressed across multiple brain regions during sleep, making mPFC NE a valid proxy for LC-driven neuromodulatory state dynamics. We have clarified this in the result section:

      (Results, page 6). “mPFC was selected as a cortical readout of LC-mediated norepinephrine dynamics, as infraslow NE fluctuations are coordinated across widespread brain regions.”

      (1.5.a) The lack of effect of Arch-inhibition of LC on infraslow NE signals is concerning. Prior work showed that bilateral LC inhibition does affect NE signals and also infraslow sigma power fluctuations (10.1016/j.cub.2021.09.041, 10.1038/s41593-024-01822-0). This discrepancy should be explicitly acknowledged and discussed on p.9.

      We thank the reviewer for highlighting this important point. To clarify, we did observe a robust effect of LC inhibition on noradrenergic signaling, with clear suppression of the NE signal following Arch-mediated LC inhibition (Fig. 4c), aligned to laser onset. In addition, LC suppression increased neuronal synchronization, including enhanced sigma power relative to YFP controls (Fig. 4h), consistent with previous reports showing that reduced LC activity promotes synchronized sleep-related oscillations. Thus, we do not interpret our findings as indicating an absence of LC suppression effects.

      The apparent discrepancy relates specifically to the absence of a statistically significant group-level shift in infraslow NE or HRV power during NREM sleep, rather than the efficacy of the manipulation itself. We note that our inhibition paradigm was intentionally mild, consisting of repeated 2-minute suppression periods separated by 4-minute intervals, and was designed to introduce subtle shifts within physiological ranges rather than globally reorganize infraslow sleep structure. Accordingly, we consider the immediate NE and heart-rate responses to LC inhibition to be the most sensitive physiological readouts of LC-mediated regulation in this context.

      We also note that the studies cited by the reviewer used different suppression paradigms, including more frequent manipulations. Thus, while our suppression scheme did not result in detectable infraslow reorganization, we do not believe they reflect an ineffective LC suppression as we demonstrated a clear NE reduction in response to time-locked LC suppression.

      We have added a sentence to the result section to explain the lack of effect infraslow power:

      (Results, page 12). “This likely reflects the relatively mild and intermittent LC suppression paradigm, which was designed to remain within physiological ranges and therefore did not globally reorganize infraslow sleep dynamics.”

      (1.5.b) Additionally, the authors state in the Discussion that they "find no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection is more strongly associated with sympathetic activity rather than parasympathetic inhibition". However, the results primarily demonstrate an absence of modulation in mean HR, while preserving a significant relationship between RR intervals and stimulation. This indicates a modulation of heart rate variability, even in the absence of changes in average HR. Notably, such variability-related effects may fall outside the VLF range and could instead involve higher-frequency components.

      We thank the reviewer for this important clarification. We agree that our original wording may have conflated the absence of modulation in mean HR with the absence of autonomic modulation more generally. To address this point, we revised the Discussion to emphasize that LC activity may influence VLF HRV through sympathetic activation and/or indirect modulation of cardiac vagal activity, even in the absence of robust changes in average HR. We have softened our previous interpretation that the LC-HR relationship is primarily sympathetic in nature and instead discuss a more nuanced interaction between sympathetic and parasympathetic influences on HRV dynamics.

      (Discussion, page 16/17). “Prior findings demonstrate that LC is linked to the autonomic nervous system. Stimulation of LC projections decrease parasympathetic cardiac vagal activity (Wang et al., 2014) direct projections to the preganglionic cells in the sympathetic nervous system (Karemaker, 2017; Nygren and Olson, 1977; Samuels and Szabadi, 2008). While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014) there is a clear parasympathetic component (Taylor et al., 1998). This combined with the ability of pharmacological blockage of the parasympathetic system to block infraslow oscillations of HR (Osorio-Forero et al., 2021) and pupil diameter during sleep (Yüzgeç et al., 2018) led us to expect a slowing of HR during LC suppression due to parasympathetic disinhibition. Interestingly, we found no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection may also somehow be driven by sympathetic outflow. Previous research had indicated that LF power may also represent sympathetic activity, but this interpretation has been challenged due to the mixed contribution of both autonomic branches (Houle and Billman, 1999; Japundzic et al., 1990; Reyes et al., 2013). LF and HF ratio (LF/HF) were traditionally thought to reflect balance between sympathetic and parasympathetic activities (i.e., the sympatho-vagal balance), though currently considered as an oversimplification of non-linear integration of autonomic signals (Billman, 2013; Pagani et al., 1986). Our findings implicate the VLF may offer a precise marker for central arousal states, given its overlaps with infraslow phasic fluctuations of LC-NE levels. Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014). Within this framework, infraslow LC–NE dynamics may represent one central contributor to these slow cardiac fluctuations during sleep either directly through sympathetic activation or indirectly by inhibition of cardiac vagal activity.”

      (1.6) The relationship between sigma power fluctuations and HR is different in humans than in mice. This has been shown before (10.1126/sciadv.1602026, 10.1038/s41593025-02159-y). This work should be mentioned on p. 11.

      We already acknowledge prior studies demonstrating species differences in the relationship between sigma power fluctuations and heart rate in the discussion section. However, to accommodate the reviewer’s comment and improve clarity for the reader, we have now also added the suggested references to the Results section, where the relationship between sigma power fluctuations and HR is first discussed.

      (Results, page 14). “These temporal differences may reflect species-specific physiology differences in cardiac timescales, which has also been previously reported (Bergel et al., 2025; Carro-Domínguez et al., 2025; Lecci et al., 2017).”

      (1.7) Correlations between the strength of infraslow sigma power fluctuations and memory consolidation have been published and should be discussed (10.1126/sciadv.1602026). It is surprising to see that correlations with learning in humans are done using pre-HRB sigma peaks rather than heart rate. This is a measure that is very close to the one used by Lecci et al.; this similarity should be clearly acknowledged. The way the data are currently presented limits this manuscript's novelty, also in its translational aspect.

      We agree that the work by Lecci et al. (2017) established an important relationship between the association of infraslow sigma power fluctuations and memory consolidation, which is highly relevant to our findings.

      Importantly, the underlying infraslow fluctuations in neuromodulatory tone are increasingly recognized as key regulators of sleep spindle dynamics (sigma power), including from our own previous work demonstrating that direct manipulation of locus coeruleus–norepinephrine infraslow rhythms alters spindle organization and sleep continuity. Thus, our findings are conceptually aligned with prior studies linking sigma fluctuations to memory consolidation.

      However, we would like to clarify an important distinction in our translational approach. While Figure 4 demonstrates that heart rate dynamics during sleep can predict memory performance in mice, Figure 5 was designed to address the translational potential of these findings in humans, where direct neuromodulatory readouts are not readily accessible. Here, we deliberately focused on heart rate bursts and their associated sleep dynamics as a clinically tractable physiological measure.

      We acknowledge that the pre-HRB sigma increase may appear conceptually similar to the measure used by Lecci et al.; however, our approach is not equivalent. Rather than selecting spindle or sigma peaks themselves, we aligned analyses to heart rate accelerations and examined the robust upregulation of sigma activity preceding these events. In this framework, sigma activity serves as a physiological readout linked to autonomic dynamics, rather than being the primary anchor of analysis. We chose this measure because heart rate bursts are influenced by multiple physiological factors, and the associated sigma dynamics provided the clearest and most robust relationship with memory outcomes in the human dataset.

      We have revised the Discussion to more clearly acknowledge the similarity to prior work. We believe our findings extend prior observations by providing evidence that sleep-related heart rate fluctuations may serve as a non-invasive readout of the memory-preserving function of sleep.

      (Discussion, page 18). “Previous work in humans demonstrated that the strength of infraslow sigma oscillations correlates with sleep-dependent memory consolidation in humans (Lecci et al., 2017).”

      (1.8) Conclusions as to whether sigma fluctuations might be slightly slower and less powerful in mice are not justified. More work is required to determine which infraslow manifestations are most useful for cross-species comparisons. Moreover, little is currently known about LC activity in human sleep. A careful look into how pupil diameter correlates with sigma power should provide clues for further discussion (see Carro-Dominguez et al). A detailed study of infraslow fluctuations in human sleep should also be discussed https://doi.org/10.1101/2024.11.06.620875.

      We thank the reviewer for this thoughtful comment. We agree that our original phrasing suggesting that sigma dynamics in mice may be “slower and less powerful” than in humans was overly interpretive. We have removed this sentence from the Results section. In addition, we have expanded the Discussion to more thoroughly integrate recent human literature on infraslow sleep dynamics.

      (Discussion, page 19). “Importantly, infraslow fluctuations of sigma power in human sleep have received growing attention. Recent work demonstrates that the infraslow fluctuation of sigma power segments N2 sleep into functional phases associated with arousal and memory-related sleep markers (Dimitriades et al., 2024). Complementary findings using pupillometry show that pupil diameter fluctuates on similar infraslow timescales during NREM sleep and is inversely related to spindle clustering, providing indirect evidence that arousal-related noradrenergic dynamics shape human sleep microstructure (Carro-Domínguez et al., 2025). However, LC activity during human sleep remains inferred rather than directly measured, and systematic perturbation studies linking LC output to spindle–autonomic coupling in humans are currently lacking. Together, these observations underscore both the promise and the current limitations of cross-species comparisons of infraslow sleep dynamics.”

      (1.9) Citation of literature should be as explicit as possible. Referring to "many studies rely on plasma levels..." while including some that actually did real-time fiber photometric measures is misleading.

      We thank the reviewer for this suggestion and agree with the reviewer about the importance of accurately representing prior studies. We have now updated the discussion to clarify this.

      (Discussion, page 15). “Many studies have linked HR to NE (Fawaz and Simaan, 1963; Sundaram et al., 1991; Tanoue et al., 2022; Watson et al., 1979). Many rely on plasma levels of NE, which, while linked to central NE (Gurguis and Uhde, 1998), has low temporal resolution, making causal interpretations harder. In recent years, the use of biosensors and fibre photometry allows for very reliable estimate of the temporal dynamics of NE changes making association to HR more precise (Osorio-Forero et al., 2021).”

      (1.10) Regarding the discussion on the baroreflex: The emphasis is placed predominantly on sympathetically mediated effects. However, the description of the baroreflex loop is incomplete, as it overlooks the substantial contribution of parasympathetic modulation. In particular, heart rate adjustments within the baroreflex are primarily mediated by parasympathetic mechanisms.

      We agree with reviewer that this important notion should be added. We have rephrased the discussion as below:

      (Discussion, page 20). “These neurons suppress the activity of the rostral ventrolateral medulla, ultimately resulting in reflex parasympathetic activation with sympathetic inhibition lowering the HR (Aicher et al., 2000; Lanfranchi and Somers, 2002).”

      (1.11) Regarding the interpretation of HRV analysis, the discussion places disproportionate weight on sympathetic modulation in the interpretation of HRV metrics. This framing is inconsistent with recent conceptual clarifications, including a recent Nature Reviews Cardiology article by Menuet et al. (10.1038/s41569-02501160-z), which cautions against simplistic low-frequency/high-frequency (LF/HF) interpretations of autonomic balance.

      We thank the reviewer for this thoughtful comment. We agree that HRV frequency bands should not be interpreted as exclusive markers of specific autonomic branches. In the original manuscript, we cited Billman (2013), which challenges the validity of LF/HF as a measure of sympatho-vagal balance, to acknowledge these conceptual limitations. However, we recognize that some of our phrasing, particularly in the section discussing compensatory HR decelerations, may have implied branch-specific dominance.

      We have updated the Discussion section as follows:

      (Discussion, page 17). “Interestingly, we found no consistent modulation in HR during LC suppression, suggesting that the LC-HR connection may also somehow be driven by sympathetic outflow. Previous research had indicated that LF power may also represent sympathetic activity, but this interpretation has been challenged due to the mixed contribution of both autonomic branches (Houle and Billman, 1999; Japundzic et al., 1990; Reyes et al., 2013). LF and HF ratio (LF/HF) were traditionally thought to reflect balance between sympathetic and parasympathetic activities (i.e., the sympatho-vagal balance), though currently considered as an oversimplification of nonlinear integration of autonomic signals (Billman, 2013; Pagani et al., 1986). Our findings implicate the VLF may offer a precise marker for central arousal states, given its overlaps with infraslow phasic fluctuations of LC-NE levels. Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014). Within this framework, infraslow LC–NE dynamics may represent one central contributor to these slow cardiac fluctuations during sleep either directly through sympathetic activation or indirectly by inhibition of cardiac vagal activity.”

      We have also replaced the title of the Discussion section title “Locus-coeruleus-mediated sympathetic control drives compensatory heart rate decelerations” with “Locus-coeruleus activation drives compensatory heart rate decelerations via autonomic feedback mechanisms”

      (1.12) In particular, the manuscript attributes VLF power primarily to sympathetic activity, despite evidence that very-low-frequency RR-interval oscillations are strongly dependent on parasympathetic integrity. Notably, parasympathetic blockade has been shown to nearly abolish VLF oscillations in humans (Taylor et al., Circulation, 1998; doi:10.1161/01.CIR.98.6.547).

      We thank the reviewer for highlighting this important point. It was not our intention to imply that VLF HRV is driven primarily by sympathetic activity or to disregard the important role of parasympathetic integrity in shaping VLF oscillations. Our intention was to present the multifactorial and incompletely resolved nature of the VLF component; however, if our wording can be interpreted otherwise, we agree that clarification is warranted.

      In response, we have revised the manuscript to better reflect the current understanding of VLF physiology. Specifically, we now make clearer that, while VLF HRV remains less mechanistically defined than HF and LF HRV, substantial evidence supports a strong parasympathetic contribution to the expression of VLF oscillations. We have incorporated the reviewer-suggested references within this comment and others and clarified this point throughout the revised manuscript.

      (Discussion, page 16). “While the mechanisms generating VLF HRV are not well defined (Armour, 2003; Shaffer et al., 2014; Wang et al., 2014) there is a clear parasympathetic component (Taylor et al., 1998).”

      (Discussion, page 17). “Although parasympathetic activity appears important for the expression of VLF oscillations, growing evidence suggests that VLF dynamics reflect broader interactions between the heart and autonomic nervous system rather than simple sympathetic or parasympathetic control alone (Armour, 2003; Shaffer et al., 2014).”

      (1.13) Moreover, based on work by Armour (2003) and Kember et al. (2000, 2001), the VLF rhythm is thought to emerge from stimulation of afferent sensory neurons within the heart, further arguing against a purely sympathetic interpretation. Together, these findings indicate that the discussion overemphasizes sympathetic mechanisms and underrepresents the contribution of parasympathetic and afferent cardiac pathways to HRV, particularly in the VLF range.

      We thank the reviewer for this important point. Our response to this comment is largely aligned with our response to comment (1.12). In the revised Discussion, we now more clearly acknowledge that VLF oscillations likely arise from more complex cardioautonomic interactions than simple sympathetic and parasympathetic innervation alone. At the same time, we have intentionally avoided an extensive discussion of these mechanisms, as our experimental design does not directly address these pathways.

      Part 2 - There are a number of conceptual issues that need more careful elaboration:

      (2.1) A highly problematic point throughout this study is the choice of AUCs, notably of NE signals and RR intervals, rather than the signal amplitudes. It confuses the correlations to events of different durations, such as MAs and wakefulness. It is not possible to draw conclusions of the kind "cardiac rhythm are tightly coupled with the infraslow phasic NE ..." because such statements ignore that AUCs conflate amplitude and duration of a signal.

      We thank the reviewer for this comment and agree that the distinction between amplitude- and duration-related signal features is important. We deliberately chose AUC measures because our intention was to capture the overall physiological response over time, including both the magnitude and temporal evolution of the signal, rather than relying solely on a single peak value. In this context, AUC provides information about the shape and sustained nature of NE and RR changes within a defined time window, which we considered particularly relevant for temporally dynamic responses.

      At the same time, we appreciate the reviewer’s concern that AUC may conflate response amplitude and duration, particularly when comparing events of different lengths such as microarousals and wakefulness. To minimize this issue, the AUC windows were intentionally kept relatively short and fixed, thereby limiting the influence of prolonged wake episodes or differences in transition duration on the measure. Thus, our intention was not to quantify the total duration of awakenings, but rather the immediate physiological response profile surrounding the event.

      To address this concern more directly, we compared amplitude- and AUC-based measures across all mice. The results are now shown in Suppl. Figure 1g. We found a strong correspondence between amplitude and AUC measurements for NE signals, indicating that the observed relationships are not dependent on the choice of metric. A similar, albeit weaker, relationship was observed for RR responses. This likely reflects physiological constraints on heart-rate dynamics, where the initial heart-rate acceleration is relatively similar across vigilance-state transitions (Figure 1e), while the duration of the response differs substantially. As a result, amplitude measures may underestimate differences between transitions, whereas AUC better captures the extent of the cardiac response. Importantly, direct comparison of NE and RR amplitudes still revealed a significant relationship, supporting the overall conclusion that cardiac and noradrenergic responses are coupled. However, this relationship was weaker than that observed using AUC measures, suggesting that incorporating temporal aspects of the response captures additional biologically relevant information.

      In addition to the figure, we have revised the Results section to include these considerations:

      (Results, page 7): “These comparisons were performed using area under the curve (AUC) estimates of NE and R-R responses. Importantly, comparing AUC and peak amplitude measures showed a similar overall relationship, and the coupling between NE and RR remained significant when only peak amplitudes were considered (Suppl. Fig. 1g), indicating that the findings are not solely driven by the response duration aspect of AUC. The somewhat stronger relationship observed with AUC-based measures may reflect rapid saturation of heart-rate responses across vigilance-state transitions, making response persistence an informative component of the physiological signal.”

      (2.2) A next problematic aspect of the study is the analysis of noradrenergic signals and HR at transitions (e.g., from NREM sleep to sleep wakefulness or to microarousals). The abstract does not mention these data, leaving open how they fit into the paper's message. There are also two problems with it: a) Noradrenaline levels increase with wakefulness, as do many other neuromodulators. This is not novel. Furthermore, why use an AUC measure for 0-25 s when MAs last only 5 or 15 s? b) biosensor signal comparisons are difficult to make for state transitions, because blood flow changes and modifies the fluorescent signal.

      We thank the reviewer for these comments and appreciate the opportunity to clarify the rationale and interpretation of these analyses.

      Regarding the inclusion of vigilance-state transitions, our intention was not to claim novelty in the observation that NE levels increase during wakefulness. Rather, these analyses were included to provide an additional physiological context in which to examine the coupling between NE and heart-rate dynamics. Specifically, the transition analyses allowed us to determine whether graded changes in NE across sleep-to-wake transitions were mirrored by corresponding RR changes (Fig. 1d–e) and whether these responses covaried (Fig. 1f), thereby strengthening the overall conclusion that cardiac dynamics track noradrenergic signaling across naturally occurring sleep-state fluctuations.

      Regarding the use of AUC measures, we deliberately chose this metric because it captures the overall physiological response over a defined time window, including both magnitude and temporal evolution, rather than relying solely on a peak value. In this context, AUC was intended to reflect differences in the overall response profile, for example that NE and RR responses during microarousals may return more rapidly toward baseline than during sustained wakefulness. To minimize the influence of differing event durations, the analysis window was kept fixed (0–25 s) across all transition types. Importantly, we directly compared AUC- and amplitude-based measures and found that they produced largely similar relationships (Suppl. Fig. 1g). The coupling between NE and RR responses remained significant when only peak amplitudes were considered, indicating that the observed relationship is not solely driven by response duration. However, the relationship was somewhat stronger with AUC-based measures, likely because heart-rate responses rapidly saturate across vigilance-state transitions, making response persistence an informative component of the physiological signal. As mentioned in the previous comment, we have added a Figure and new result text to highlight this.

      Regarding the concern about blood-flow related artifacts in fluorescent biosensor signals, we agree that hemodynamic contamination is an important consideration for neuromodulator recordings and applies broadly to biosensor-based measurements, including analyses of infraslow fluctuations. To minimize this issue, ΔF/F calculations were performed using the isosbestic control channel, which serves to correct for movement- and hemodynamic-related signal fluctuations. Furthermore, hemodynamic artifacts typically occur rapidly at state transitions, whereas GRAB-NE signals display slower dynamics. If uncorrected hemodynamic contamination strongly influenced the signal, we would expect abrupt signal distortions tightly aligned to arousal onset, which was not evident in our recordings. While we cannot completely exclude residual hemodynamic influences, we do not believe they account for the graded NE responses observed across vigilance-state transitions or the corresponding relationship with RR dynamics.

      (2.3) Wordings such as 'extent of arousal' to compare MAs and wakefulness are problematic. Microarousals and wakefulness are qualitatively different behaviorally, physiologically, and in terms of neuromodulatory conditions

      We thank the reviewer for the helpful comment. Our intention was not to imply that microarousals and wakefulness are qualitatively identical states differing only in magnitude. Rather, we used arousal in the broader neurophysiological sense, referring to the degree of activation of central arousal systems, with wakefulness representing part of this continuum. However, we recognize that in the sleep field, arousal is often used more specifically to describe brief EEG desynchronization events and that we furthermore use micro-arousals as a broad term for these sleep arousals, which may make our wording confusing. To avoid ambiguity, we have revised the manuscript to replace this terminology with vigilance state, sleep–wake state, or state transitions, depending on the context.

      (2.4) To interpret linear correlations between datasets, even for the ones with Rsquare values < 0.5, as 'predictive' represents an overextended interpretation that is not supported by available evidence. This concern is aggravated due to the use of AUCs that confound amplitudes and time courses. This is in particular the case for Figure 5d, 5k, or 5m…

      We thank the reviewer for this important comment. We agree that the term predictive may overstate the interpretation of these correlations, particularly given the modest R² values in some analyses. Our intention was to highlight an association between heartrate and NE-related measures rather than imply strong predictive performance or causality. We have therefore revised the wording throughout the manuscript to avoid predictive language and instead refer to these relationships as associations or correlations.

      Regarding the use of AUC, we selected summary metrics based on the physiological characteristics of the signal of interest. In cases where responses were characterized primarily by rapid shifts to a new level, amplitude measures were used. In contrast, AUC was chosen when both the magnitude and duration of the response were considered physiologically relevant.

      Part 3. A substantial number of experimental and analytical points require clarification. Here is a list of a few examples; many observations noted here apply equally to other figure panels.

      (3.1) Many figure panels leave it open about whether averages or representative data are shown, how many animals are included, and what kind of measures are plotted.

      We thank the reviewer for this comment and agree that clarity in figure presentation is important. In response, we have carefully revised the figure legends throughout the manuscript to more explicitly state whether data shown are representative examples or group averages, clarify the number of animals included in each analysis, and specify the measures being plotted. We have also added “data are shown as mean ± SEM” where this information was previously not explicitly stated and clarified when n refers to the number of animals. We believe the revised figure legends now provide clearer guidance for interpretation, and further statistical details are available in the accompanying statistics table. It should also be noted that number of animals used for all experiments can be found in the Method section ‘Mice’.

      (3.2) In case experiments were done in a paired manner (e.g., the LC stimulations), individual data points should be shown connected for the different conditions.

      We thank the reviewer for this suggestion. While we agree that connecting individual data points is valuable for paired experimental designs, this visualization is not appropriate for the LC stimulation analyses presented here. Specifically, each condition reflects pooled stimulation events selected across animals rather than a single summary value per animal that can be directly matched across columns. As such, individual points in one condition do not map one-to-one onto points in the next condition, making connected visualizations potentially misleading.

      Importantly, although the data are presented as pooled event-level measures, the paired structure of the experiment was accounted for in the statistical analyses, such that repeated measurements within animals and the paired nature of the design were included in the relevant comparisons.

      (3.3) Numerous analyses involve heart rate measures from the neck EMG during wakefulness. However, it is not specified how, in this case, RR peaks could be detected within the high-activity EMG.

      We thank the reviewer for this question. The procedure for heart-rate detection during wakefulness is described in detail in the Heart rate detection Methods section. Briefly, RR intervals were extracted from preprocessed neck EMG recordings using a previously validated approach for mouse sleep studies. To minimize contamination from movement-related EMG activity during wakefulness, R-peaks were not detected within periods extending 100 ms before and 250 ms after detected movement, as these segments were considered too noisy for reliable peak detection. Heart-rate estimates during these excluded periods were subsequently interpolated using surrounding valid RR intervals to preserve temporal continuity and enable analysis of HR dynamics before and after movement episodes.

      (3.4) Figure 1b: PSD for RR intervals. The supplementary figure says that 11-minutelong NREMS or 5-minute-long NREMS periods were used. These are very rare events in mice, for which the average bout duration is around 2 min and the cycle length is 10 minutes. The methods do not explain how these bouts were chosen, how many of them were included, and why shorter bouts were not analyzed. Single cases or means?

      We thank the reviewer for this important point and agree that additional clarification was warranted. The selection of long NREM (including microarousals) periods was motivated by methodological considerations related to spectral analysis of very slow oscillations rather than by an assumption that these bout lengths are representative of average NREM duration in mice. Because our analysis focused on very-low-frequency dynamics (~1 cycle every 50 s), sufficiently long continuous recordings are required to reliably estimate power at these frequencies and avoid fragmentation-related edge effects that disproportionately affect shorter bouts.

      For this reason, shorter NREM episodes were excluded from the PSD analysis, as they do not provide sufficient duration to robustly capture slow-frequency components. We initially compared PSD estimates using NREM periods of at least 11 min (allowing ~10 VLF cycles) and 5 min duration to assess whether the shorter 5 min recordings introduced bias (Supplementary Fig. 1f). While longer periods resulted in higher overall power estimates, the frequency distribution remained highly similar between conditions. We therefore selected 300 s (5 min) as the inclusion criterion for the remainder of the study, as periods exceeding 10 min are uncommon in mice and would substantially limit analyses across experimental paradigms.

      To further reduce bias related to bout duration, PSD estimates were weighted by NREM episode length, as longer bouts showed systematic effects on power estimates. For Figure 1, the analysis included 144 NREM-with-MA bouts across 7 animals, and data shown represent group means rather than single examples.

      This description is also included in the Methods section/Data Analysis.

      (3.5) Figure panel 1c,d: In Panel c, what is plotted?

      As stated in the figure legend, it is the cross-correlation between NE and R-R. We have added more description in the figure legend.

      A cross-correlation between two signals? If yes, how were these signals chosen per transition? Looks rather like they plot some time course across a transition. What is time point 0?

      We thank the reviewer for pointing out that this analysis was insufficiently explained. The analysis shown represents a cross-correlation between the NE and RR signals, performed to assess their temporal relationship across different vigilance-state transitions. Specifically, for each transition type, NE and RR signals were extracted within the corresponding time windows and cross-correlated to determine the strength and timing of their interaction.

      In this context, time point 0 (lag = 0) represents perfect temporal alignment between the two signals. Positive or negative lags indicate whether changes in one signal systematically precede or follow changes in the other. Due to methodological differences between the rapid electrical heart signal and the slower fluorescent NE signal, we intentionally avoided overinterpreting fine temporal lead–lag relationships. Rather, the aim of this analysis was to characterize the overall interaction between the two signals across transitions.

      Because the relationship between NE and RR was predominantly inverse, negative cross-correlation values indicate that increases in NE are associated with decreases in RR (i.e., faster heart rate), and vice versa. Thus, this analysis served primarily to confirm and extend our other findings by quantifying the interaction between NE and heart-rate dynamics across vigilance-state transitions.

      We have revised the descriptive sentence in the Results section to improve clarity and have added a more detailed description of the signals included directly in the figure panel.

      (Results, page 7). “Here, cross-correlation analysis revealed a predominantly negative relationship between NE and RR across vigilance-state transitions, indicating that increases in NE were associated with reductions in RR (i.e., faster heart rate; Fig. 1c).”

      (Figure 1 legend). “Cross correlation (how strongly and at what temporal offset the two signals covary) between NE and RR during transitions”.

      In Panel d, what is measured here? Is the time point of the dotted line a NA trough or the moment of a transition? Show the data with connected lines. Heart rate calculation during wakefulness?

      We thank the reviewer for these questions and apologize that this was not sufficiently clear. In panel d, the dotted line indicates the NE trough, not the moment of a vigilance state transition, as specified in both the figure and figure legend. The analysis is aligned to detected NE troughs during sleep and examines the subsequent physiological dynamics, including transitions into wakefulness.

      We have updated the sentence in the result section to make it a bit more clear:

      (Results, page 7): “To explore the NE-RR relationship across sleep-wake transitions, we examined four progressive sleep-to-wake transitions using the preceding NE trough as the time stamp (time 0):…”

      Regarding the suggestion to connect data points, as noted in an earlier response, these analyses are based on event detection, where multiple events contribute from each animal. Thus, each condition represents pooled events across animals rather than a single matched value per animal, making connected-line visualizations inappropriate and potentially misleading. Importantly, the paired structure of the experimental design was accounted for in the statistical analyses.

      Heart-rate detection during wakefulness was usually not possible due to movement artefacts in the EMG. In the detection, we excluded periods with movement. Our EMG-based detection approach is described in the Heart rate detection Methods section. Briefly, movement-contaminated periods were excluded from R-peak detection (100 ms before and 250 ms after detected movement), and RR intervals were subsequently interpolated using surrounding valid values to preserve temporal continuity. Because analyses were aligned to NE troughs occurring during sleep, we limited the temporal window to avoid excessive contamination from movement-related noise associated with subsequent wakefulness. We have slightly revised the text to make these points clearer.

      Panel C is a cross correlation between NE and RR from the same traces included in the mean traces. x=0 is the NE through like the other figures.

      Panel D is the mean traces of NE during transitions with the dotted line at x=0 being the NE though

      Yes, but x = 0 for the cross-correlation is not the NE trough. It says something about how aligned NE and R-R are in time (see response further up in (3.5)).

      (3.6) The sigma band should ideally be chosen between 10-15 Hz for better consistency with the literature.

      We thank the reviewer for this suggestion and agree that consistency in the definition of frequency bands is important for comparison across studies. The sigma range used in the present manuscript was selected to match our previous publications and analyses, thereby allowing direct comparison with our earlier findings on LC-mediated regulation of sleep spindles and infraslow sleep dynamics. Maintaining the same band definition also ensured consistency across the datasets analyzed in this study.

      We acknowledge that having a shared definition of sigma power would improve alignment and facilitate comparisons across laboratories. At present, there remains some variability in the exact frequency boundaries used for spindle and sigma analyses across studies and species, although there are ongoing efforts within the sleep field to improve standardization. Importantly, we do not expect that modest adjustments of the sigma-band boundaries would materially affect the conclusions of the present study, as the spindle-related activity of interest lies well within the selected frequency range and the observed effects are broad rather than restricted to a narrow frequency bin.

      Moving forward, we aim to follow emerging consensus recommendations where appropriate. To clarify this point for readers, we have added a statement in the Methods section explaining that the sigma band was chosen to maintain consistency with our previous publications.

      (Methods: EEG power, page 29). “The sigma band was defined as 8 - 15 Hz to maintain consistency with our previous publications. Modest differences in sigma-band boundaries are not expected to affect the main conclusions.”

      (3.7) What are 'extreme LC stimulation frequencies'. The only information available is that stimulations were done for 2s at 20 Hz.

      We thank the reviewer for pointing out that this wording was unclear. By “extreme LC stimulation frequencies”, we did not refer to the within-stimulation pulse frequency (which remained constant at 20 Hz for 2 s across all conditions). Rather, we referred to the effective frequency of LC activation at the infraslow timescale, which was progressively increased through the closed-loop stimulation paradigm.

      Specifically, stimulations were triggered when NE levels crossed increasingly permissive thresholds during the descending phase of the endogenous NE signal. As thresholds increased over time (from −15 ΔF/F (%) to +5 ΔF/F (%)), stimulations occurred progressively earlier in the infraslow cycle, thereby compressing the oscillatory period and increasing the effective frequency of LC recruitment while preserving the endogenous temporal structure of NE dynamics.

      Thus, “extreme stimulation frequencies” refers to the highest rate of repeated LC activations achieved through the closed-loop paradigm, where stimulations became increasingly frequent at the infraslow level rather than changes in the 20 Hz pulse train itself. To avoid confusion, we have revised the wording throughout the manuscript to refer more explicitly to faster infraslow LC activation frequencies or increased infraslow stimulation frequency.

      We have added more information in the result section to highlight this better.

      (Results, page 8). “We employed a closed-loop paradigm, where LC stimulations (2 s 20 Hz (10 ms) blue laser pulses with a light intensity of 5 mW) were triggered when NE levels fell below increasing thresholds (-15, -10, -5, 0 and 5 ΔF/F (%), Fig. 2a-b, Methods). This approach enabled controlled compression of the infraslow NE cycle by triggering LC activation during the descending phase of the endogenous NE signal, thereby increasing the effective oscillatory frequency while preserving the temporal structure of physiological LC–NE dynamics. This strategy allowed us to test whether heart-rate responses continue to track LC-driven NE fluctuations as the infraslow rhythm becomes progressively faster.”

      (3.8) Could the discrepancy between panels 4c, left and right, be due to limited sample size? The whole figure lacks indications of sample numbers, making interpretation difficult.

      We thank the reviewer for this comment. We assume the reviewer is referring to the apparent discrepancy between the NE response and heart-rate response following LC suppression (Fig. 4c–d), where LC inhibition induced a clear reduction in NE levels, whereas mean heart-rate responses were less pronounced.

      The figure legends state sample sizes and event numbers. Specifically, for these analyses, n = 8 animals (4 Arch, 4 YFP) were included, comprising 48 Arch events and 38 YFP events, our interpretation is that this discrepancy reflects a biological observation rather than a failed manipulation. Specifically, LC suppression robustly reduced NE levels, confirming the effectiveness of the optogenetic intervention, whereas HR did not exhibit a similarly consistent group-level response. However, as highlighted by the correlation analyses, variability in RR responses remained associated with the magnitude of NE suppression, suggesting that heart-rate dynamics still reflected noradrenergic modulation at the individual-response level despite the absence of a strong mean effect.

      (3.9) It would be great if Figure 4 j could be more explicitly illustrated. For example, behavioral traces that lead to higher NFR and corresponding changes in RR AUC should be shown for animals with large and small effect sizes. The sample size seems excessively low. Can this explain the difference in slopes compared to Figure 4e, right panel?

      We thank the reviewer for this suggestion. To improve the interpretation of Figure 4j, we have now added representative examples illustrating animals with high and low memory performance and their respective RR responses following LC suppression.

      We agree that the sample size for this analysis is limited. As noted in the Methods, Figure 4j is based on a secondary analysis of a previously published dataset (Kjaerby, Andersen et al., 2022), where heart-rate measures were retrospectively extracted from EMG recordings. Due to noise-related limitations in RR detection, reliable cardiac measures could not be obtained from all animals, reducing the number of subjects available for this analysis. For this reason, we have deliberately avoided direct statistical comparisons between Arch and YFP animals and instead limited our conclusions to the observed association between RR responses and memory performance across animals.

      Regarding the difference in slope compared with Figure 4e, the two analyses are based on different levels of aggregation and address different questions. Figure 4e examines the relationship between NE and RR responses across individual LC suppression events, resulting in multiple observations per animal. In contrast, Figure 4j uses a single mean RR response and a single behavioral outcome per animal. Furthermore, Arch and YFP animals were pooled in Figure 4j to maximize statistical power and because the dataset was not sufficiently powered for direct group comparisons. Consequently, the slopes are not expected to be directly comparable between the two figures.

      (3.10) It would be important to show anatomical validation of viral expression in THcre animals and optic fiber positioning.

      We thank the reviewer for this comment. Anatomical validation of viral expression in TH-Cre animals and optic fibre placement was performed for these experiments and has been reported previously in Kjaerby et al. Nature Neuroscience paper, from which this dataset was derived. Specifically, viral targeting and fibre positioning were histologically verified as part of the original experimental validation. This is mentioned in the Method section ‘Surgery’: Viral expression and injection sites were validated through immunostaining of perfused brain slices from the experimental animals (see Kjaerby et al. (3) for more information).

    1. Author Response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study reveals distinct representations of task-related information in the dendrites and somata of cortical neurons during sensorimotor learning and behavioral adaptation. The evidence is compelling, combining simultaneous imaging of dendritic and somatic activity during behavior to demonstrate compartment-specific encoding of sensory cues, motor actions, and corrective signals. The work will be of broad interest to neuroscientists studying dendritic computation, motor learning, and the cellular mechanisms underlying adaptive behavior.

      Thank you for this excellent summary. We recommend one change: removing the word “simultaneous”. It could perhaps be replaced with “concurrent” or simply omitted. Tuft dendrites and somata were imaged on alternating days, and most readers will probably interpret “simultaneous” as implying a faster, interleaved sampling rate.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Scheib et al. identify distinct calcium dynamics in the somata and tuft dendrites of layer 5 pyramidal cells in mice performing a licking task. Animals are trained to lick water ports on the left or right following an acoustic cue, and can adjust their targeting when the ports are displaced. For tongue premotor cortical neurons projecting to the ventromedial thalamus, calcium transients in tuft dendrites are tightly locked to the direction-instructive cue, while somatic calcium signals are more broadly dispersed and more frequently synchronized with tongue motion and port contact. Finally, when the targets are shifted, tufts exhibit a sparse but large corrective signal on an improperly-targeted first lick, and the changes in population activity in the tufts and somata differ after adaptation to the new port locations.

      Strengths:

      (1.1) In my opinion, this is a very strong manuscript which reports several novel and significant observations, contains high-quality data and (for the most part) reasonable analyses, and is clear and well-written. Most prior studies of cortical sensorimotor processing have measured the output of neurons using extracellular recording - an approach which obscures potentially important signaling differences between neuronal compartments. This study leverages cutting-edge imaging techniques in mice to document large, time-dependent differences between calcium signals at cortical somata and tuft dendrites. This phenomenon could have major implications at the cellular level for synaptic plasticity, and at the systems and behavioral levels for motor adaptation. As described below, I have only one major technical concern (which should be addressable with additional analysis), along with several relatively minor suggestions for improving the manuscript.

      We thank the reviewer for their insightful summary of the significance of the differences that we identified in the task-related activity of tuft dendrites and somata.

      Weaknesses:

      (1.2) At a conceptual level, the authors may wish to elaborate a bit on what sensorimotor computation they think the circuit is implementing, and how their results help explain this implementation. Several possibilities are raised: tuft activation could "prime" the pyramidal cells in advance of movement initiation (line 319ff), or could track errors to engage plasticity (line 351ff) and solve the credit assignment problem (line 362ff). It might be helpful to make one of these proposals more concrete with a computational model, but this is not strictly necessary.

      We thank the reviewer for this feedback. We absolutely agree that detailed computational models of each proposed computation will be very valuable and constitute an important follow-up to this work. We hope to collaborate with theorists to take that next step. Each possible computation noted by the reviewer reflects distinct differences that we observed in the task-related activity of tuft dendrites and somata. They are not mutually exclusive hypotheses to explain the same phenomenon. As such, we think they are best addressed independently in future modeling work. By making the data and a concise description of the main findings available immediately, we hope to allow computational experts in each of these areas to take advantage of the results of this study without delay.

      (1.3) My only major technical concern relates to the analyses in Figures 4F-H, 5G-I, and 6H-K (c.f. equations 2-5). Typically, one identifies population-level factors by projecting neural activity onto fixed dimensions of interest; this makes it possible to see how activity evolves over time along interpretable coordinates. Here, however, the coding directions are redefined at each time point, so the "choice" activity at time t is actually a different signal from the "choice" activity at t+1. This procedure is a bit like comparing the activity of one neuron at one time point with the activity of a different neuron at a later time point. It also makes the physiological interpretation more complicated: if the dimensions are fixed, one can see how a downstream neuron could "read out" the signal by computing a weighted sum of the activity of upstream neurons, but it is harder to see how this could happen if the weights are always rotating.

      We thank the reviewer for raising this point. We agree that our use of projections along coding directions (CDs) defined at each time point is a less conventional use of coding directions, although nearly identical calculations have been previously used to assess population-level selectivity and code stability in this task (Chen et al., 2017; Yang et al., 2022). As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task-dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task-variable. Finally, because the primary goal of the study was to identify any differences between tuft dendrite and somatic encoding, we think that calculating the population selectivity at each timepoint gives readers a less biased view of the selectivity of the two compartments, whereas calculating a CD over a single arbitrary time window could conflate differences in dynamics with differences in selectivity.

      We agree that calculating the CD at each timepoint makes it hard to see where the code is stable and where it is rotating, and thus how a downstream neuron might “read out” the signal. To provide this information, we have added new panels to the supplement showing the correlation of CDs across time (Figure 4 - figure supplement 1B,D). We also now provide this information for CR-CA in Figure 5—figure supplement 2A (the plots previously presented in 2A were the correlations of CR with CA, rather than CR-CA; an error that has been fixed). The following changes were also made to the Results section to clarify this issue:

      “From the linear model, we calculated coding directions (CDs) at each timepoint that maximally separated Stimulus, Choice, and Outcome activity (Figure 4F; Figure 4—figure supplement 1A) and estimated the direction and selectivity along each dimension across time (Figure 4—figure supplement 1B-E; see Methods). Allowing CDs to rotate in time (see Figure 4—figure supplement 1B,D), although unconventional, ensured that comparisons of population selectivity across the two compartments were not biased by the selection of an arbitrary CD time window.”

      We also identified a mistake in the description of CD orthogonalization in the Methods, which has been corrected as follows:

      “For each timepoint, each selectivity CD was then orthogonalized with respect to the other two selectivity CDs by a QR decomposition in which that selectivity CD was last in the order.”

      (1.4) A few comments on the behavioral task and results. After the port shift, the error rate is quite high, and doesn't diminish much between the early and late epochs (approximately 42% and 38% error rate, respectively; Figure 1I). That is, mice do not seem to fully master the task. Clearly, animals do alter their aim, but even this does not seem to change much between early and late periods (Figure 1J). I recommend that the authors show the behavioral data at a finer level of granularity (e.g., by plotting the change in exit trajectory on all individual trials across sessions, with a loess fit) to allow an assessment of the adaptation rate and when adaptation saturates. It would also be more conventional to refer to the behavioral changes as "motor adaptation," instead of "skill learning." (The latter would be appropriate if the port offset were randomized across trials, and animals received two separate cues for direction and offset, but I suspect this task would be too difficult for mice to learn.)

      We agree with the reviewer that by the end of the late period, performance on the right side (Figure 1I) has still not returned to pre-shift levels. This may reflect mice not fully mastering the task, as the reviewer suggests, or it may reflect that after the shift, the right port is substantially more difficult to reach than the left port. Unfortunately, because of high animal-to-animal and lick-to-lick variability, plotting the post-shift lick angle at a finer level of granularity is not statistically informative.

      With regard to the nature of the learning in our task, we selected “skill learning” as the best description of the motor learning task based on distinctions between adaptation and the learning of motor skills by Krakauer et al., 2019 and Heald et al., 2021. Conceptually, the difference is whether an existing motor controller memory is simply updated with new parameters, or whether the motor context has changed sufficiently that a distinct motor controller memory (which can still use parts of previous memories) is formed. In our task, after the port shift the left port forms an obstacle to reaching the right port. This obstacle was simply not present before the shift. Before the shift, ports were approximately equidistant from the mouth and easily avoided given the port separation and tongue width. Thus, avoiding an obstacle would presumably not be part of the initial motor controller memory and a distinct memory would need to be constructed.

      We agree with the reviewer, however, that given that we do not have fine-timescale dynamics of behavioral changes in response to the shift, and did not conduct other experiments (such as returning the ports to their original location) that would typically be conducted to identify “adaptation-like” or “skill learning-like” dynamics, we cannot empirically distinguish between the two. We now clarify in “Study limitations” that we call the studied behavior “skill learning” based on the nature of the task, but that our behavioral analysis cannot distinguish between adaptation and skill learning:

      “We refer to the behavioral paradigm as motor “skill learning” strictly based on the nature of the task. After the shift, mice must avoid a new obstacle close to the mouth (i.e., the left port), which we assume requires the formation of a distinct motor controller memory and therefore would be considered skill learning (Krakauer et al., 2019). However, we did not confirm that the mice exhibited specific behavioral characteristics of skill learning and it is possible that other kinds of motor learning (e.g., motor adaptation) were dominant.”

      (1.5) This is perhaps a semantic point, but it might not be entirely accurate to refer to the activity evoked by the directional cue as "sensory." Typically, a "sensory" response should encode some feature of a stimulus - in this case, the frequency of a tone. Here, it seems likely that the cue-aligned activity reflects the instructed lick direction, rather than the auditory information per se. (Presumably, these premotor neurons do not have well-behaved auditory tuning curves.) By comparison, in macaques performing center-out reach tasks, activity in dorsal premotor cortex rapidly ramps up following a visual cue instructing the direction of an upcoming reach, but one usually wouldn't refer to this activity as "visual" or "sensory" (though this is sometimes done). I suggest the authors either use "Instruction" or similar (e.g., in Figure 4F), or clarify in the text whether they think the activity is a genuine auditory response or something else.

      We understand how this could cause confusion. “Sensory” was meant to denote the nature of the differences in external events between the trial types used to calculate selectivity, not to imply that the activity was necessarily selective for detailed features of the cues outside the context of the task. Previous work in ALM cortex has labeled this selectivity direction as “stimulus” (Yang et al., 2022; Chen et al., 2024) to better emphasize that it is simply defined by the external cue. Where appropriate, we have revised the article to use “stimulus” or “instructional cues” in place of “sensory” for clarity and to better conform with convention.

      Reviewer #2 (Public review):

      Summary:

      The authors set out to compare functional encoding in the tuft dendrites and somata of a specific cortical cell type during motor planning and learning.

      Strengths:

      (2.1) The investigation of a specific projection type (L5 ET) is a strength that aids reproducibility and interpretation. The elegant approach to increasing the depth of field of dendritic imaging is another strength. The data analyses are largely clear in their methods, scope, and interpretation. The writing is extremely clear and appropriately referenced, with an excellent Introduction, in particular.

      We thank the reviewer for their appreciation of the study design, imaging methods, and scholarship of the article.

      Weaknesses:

      (2.2) It is not obvious whether the selected labeling strategy avoids labeling Layer 6 CT neurons, which would contaminate dendritic recordings. The images provided suggest enrichment in L5, but a discussion of this important potential caveat is warranted, especially since within-cell comparisons of apical dendrites to somata were not performed.

      We thank the reviewer for emphasizing the need to discuss this potential issue. For the following reasons, it is likely that the vast majority of dendrites we imaged in layer 1 originated from layer 5 ET neurons. First, as the reviewer notes, the provided images suggest enrichment in layer 5. This enrichment likely reflects the fact that most L6 CT neurons in motor and premotor cortex send denser projections to other thalamic nuclei than to VM thalamus (Winnebust et al., 2019, Cell), where we targeted our retrograde-Cre injections. Second, L6 CT neurons are predominantly untufted (Ledergerber and Larkum, 2010, J. Neurosci.), including in motor and premotor cortex (Peng et al., 2021, Nature; Ichikawa, 2025, Front. Neuroanat.). A recently identified subclass of L6 CT neurons in secondary motor cortex has dense projections to VM thalamus, but this class also appears to extend minimal dendrites into L1 (Li et al., 2024, bioRxiv). Nonetheless, we did not label post-hoc tissue collected from imaged mice with markers of precise laminar boundaries, and thus cannot definitively rule out the possibility that dendrites from a subclass of L6 CT neurons with tuft dendrites were also imaged. We have added the following paragraph to the “Study limitations” section to make readers aware of these issues:

      “L5 ET neurons in premotor cortex elaborate extensive tuft dendrites in L1, whereas Layer 6 (L6) corticothalamic (CT) neurons are predominantly untufted (Jiang et al., 2020; Peng et al., 2021). Thus, although we cannot rule out the possibility that dendrites from a subclass of L6 CT neurons were also sampled, it is likely that the vast majority of dendrites we recorded in L1 originated from L5 ET neurons.”

      (2.3) The application of DeepInterpolation to dendritic data appears to be novel, and little detail or vetting is provided. The reader is left guessing: Was the model retrained or fine-tuned on dendritic data? How does the denoising affect the resulting segmentation and activity traces? Is denoising necessary for this workflow?

      We thank the reviewer for requesting this useful additional information.

      In all cases, the model was retrained for each dendritic or somatic imaging session. Denoising improved segmentation consistency, as measured by comparing segmentations of individual sessions from the same animal. This is now specified in the Methods as follows:

      “The DeepInterpolation model was trained on each imaging session prior to denoising of that session. Denoising prior to NMF-based segmentation resulted in more robust and consistent dendrite segmentation than NMF-based segmentation without prior denoising (0.79 +/- 0.01 ⍴ vs. 0.46 +/- 0.01 ⍴; mean of the max Spearman correlation of components across sessions; random subsample of N = 3 mice, 15 sessions, 400 components).”

      With regard to how denoising impacts activity traces, examples were shown in Figure 2I, K. To provide more quantitative information to the reader, we calculated estimates of the power and reliability of the spectral content of dendrite activity traces extracted with or without denoising. These data are now shown in the new panel, Figure 2 - figure supplement 2G. The power spectral density of the denoised activity and the estimated reliable power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C). Some frequencies beyond this point have been suppressed beyond what would be expected due to photon shot noise (as estimated by the replicate coherence-weighted PSD, or “recoverable” PSD). Further characterization of the precise nature of the suppressed high-frequency information – which could be suppressed artifacts (e.g., fast brain motion) or lost signal detail (i.e., GCaMP8m rise kinetics) – is beyond the scope of this paper.

      Details of the PSD calculations have been added to the Methods, and the following statement has been added to the Results: “Power spectral density of the denoised traces and the coherence-weighted power spectral density of the raw traces match up to approximately 2.6 Hz (Figure 2 - figure supplement 2G; Methods), which is not far from the bandwidth of GCaMP8m, given its estimated combined rise and decay (Figure 2 - figure supplement 3B, C).”

      (2.4) The activity patterns of the recorded cells appear to lack the characteristic ramping during the delay epoch previously reported in both calcium imaging and electrophysiology studies. Given that a major contribution to the significance of the work is to constrain models of ALM function, a discussion of how the data aligns with previous measurements in the same circuit would improve the work.

      Preparatory selectivity and ramping activity can be seen in Figure 3H, Figure 6I, and Figure 5 – figure supplement 1B. We note that in ALM cortex, the ramping mode explains a minority of the total variance (~17%, Yang et al., 2022), but it can appear particularly prominent in projections along certain fixed CDs.

      (2.5) It would be very informative to compare differences in signals between dendrites and somata of the same cells. Consistently tracing dendrites to their respective somata would assuage worries of potential contamination from dendrites of deeper cells and enable more direct comparisons of signal transformations between dendrites and somata. It would be good to understand the relationship between dendritic calcium signals and backpropagating action potentials in this task. The authors detect less frequent calcium events in tufts versus somata; is this due to selective backpropagation of action potentials? The dynamics of this process were recently investigated by Adam Cohen's group in vivo and in vitro, and measurements in the present settings could be compared to such work.

      We agree with the reviewer that being able to compare differences in signals between the dendrites and somata of the same cells would be very valuable. However, reliable tracing of tuft dendrites to somata from in vivo 2P anatomical imaging requires extremely sparse labeling, such that very few neurons are recorded per animal (Kerlin et al., 2019, eLife; Otor et al., 2022, Science). As stated in the “Study limitations” section of the Discussion, we suspected (correctly) that some task-related selectivity (i.e., selectivity for corrective action) would be sparsely represented in the dendrites, and thus adopted a labeling and image processing strategy that allowed us to record from many dendrites per animal. This strategy necessarily comes at the expense of generating a labeling density that precludes reliable tracing of tuft dendrites to their respective somata based on 2P morphology alone. As discussed in our response to reviewer comment 2.2 and a new paragraph of “Study limitations,” substantial contamination of the dendrite recordings by dendrites of L6 CT neurons is highly unlikely. Future studies could use simultaneous functional imaging across large volumes combined with activity-based segmentation or post-hoc high-resolution imaging of tissue sections registered to in vivo 2P imaging to accomplish both high-throughput dendritic imaging and reliable tracing.

      We thank the reviewer for pointing out that we could discuss selective backpropagation as a potential mechanism more explicitly. Our results are consistent with previous studies of L5 tufts in vivo (Francioni et al., 2019, eLife), including in ALM cortex (Maristany de las Casas et al., 2026, Science), that reported that rates of multi-branch calcium transients in the tuft dendrites of L5 neurons are lower than somatic spike rates. As discussed in “Study limitations,” there is not a clear approach in our data to determine the precise nature of the events underlying the calcium transients we measured in the tuft dendrites. Selective backpropagation of action potentials is certainly one possibility and we agree that recent research from Dr. Adam Cohen’s group should be discussed. We have added the following to the Discussion:

      “Based on previous calcium imaging of L5 tufts in ALM cortex of mice engaged in similar tasks (Kerlin et al., 2019; Maristany De Las Casas et al., 2026), we suspect that most of the activity we measured was coincident with global tuft or hemi-tree events, as well as somatic spiking. Recent in vivo voltage imaging in the hippocampus has also indicated that most spikes in distal dendrites start as bAPs that have been selectively amplified (Wu et al., 2026; Lee et al., 2026).”

      (2.6) The Coding Direction analyses presented in this work, while consistent with previous literature on population codes in ALM, are at odds with the nature of the measurements here. The changes in representation that occur between the dendrites and soma of an individual cell are probably best thought of in terms of the dynamics of signals themselves within individual neurons, rather than in the information encoded across a population.

      We thank the reviewer for giving us the opportunity to clarify this issue. As noted in the article, given low numbers of error trials and high trial-to-trial variability, we found that estimating the selectivity of individual ROIs for these task dimensions was not robust and was subject to overfitting. Cross-validated projections at each timepoint provided a far more robust measure of population selectivity. Furthermore, we were able to orthogonalize stimulus, choice and outcome CDs to better identify distinct encoding of each task variable in the population activity. Thus, the analyses are not at odds with the nature of the measurements in the study.

      Nevertheless, it is true that by recalculating the CD at each timepoint, our selectivity projections do not provide the same information as conventional projections along a fixed CD, which can indicate where the selectivity code is stable and where it is changing. To provide this information we have added new panels to the supplement showing the correlation of selectivity CDs across time (Figure 4 - figure supplement 1B, D).

      (2.7) This work is largely observational, describing signals that might reflect computational transformations and/or instruct plasticity, but those possibilities have not yet been deeply investigated. The manuscript does a good job of laying out these as future directions.

      We agree with the reviewer. As noted by the reviewer in comment (2.1), we combined a number of approaches in an innovative manner to explore how tuft dendrite activity differs from somatic activity at the population level during motor learning. These measurements provide the necessary foundation for future mechanistic studies and we think it is appropriate to share them at this stage of investigation and in the format of this article.

      Reviewer #3 (Public review):

      Summary:

      This article by Scheib et al. investigates how layer 5 extratelencephalic (ET) neurons in the frontal cortex encode sensorimotor information during motor learning, focusing on differences between their apical tuft dendrites and somas. The authors alternated recordings among these ET neuronal compartments in the mouse anterior lateral motor cortex (ALM) during a cued directional licking task with a target port shift. They found that while tuft dendrites predominantly encode sensory cues, with a subset selectively active during corrective actions, somatic activity was more strongly associated with action timing. Additionally, learning induced divergent plasticity: tuft dendrites increased their selectivity but decreased response gain, maintaining stable net selectivity, whereas somas showed increased net selectivity early in learning. Together, these findings reveal distinct sensorimotor representations and learning-related plasticity in dendritic and somatic compartments, providing insight into how compartment-specific activity in the frontal cortex may contribute to motor skill acquisition.

      Strengths:

      The authors developed an innovative imaging approach and a comprehensive data analysis pipeline to address a knowledge gap in the literature. By alternating imaging of dendritic tufts and somas in the same animals, they compare compartment-specific activity during motor learning and identify distinct encoding of task variables and learning-related plasticity across these compartments. Interestingly, a subset of dendritic tufts shows activity associated with corrective actions. The findings are discussed in the context of current theories of dendritic computation, credit assignment, and motor learning, providing a useful foundation for future mechanistic studies.

      We thank the reviewer for highlighting interesting findings in the paper and their assessment that it provides a “useful foundation for future mechanistic studies”.

      Weaknesses:

      No major weaknesses were identified.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A very minor suggestion: it would be useful to mention the model organism in the abstract or title.

      (1.6) Thank you for catching this. We have added the model organism to the abstract as follows:

      “Using longitudinal two-photon calcium imaging, we investigated sensorimotor encoding in the apical tuft dendrites and somata of L5 extratelencephalic (ET) neurons in the frontal cortex of mice during learning of a discrete change to a cued dexterous action.”

      Reviewer #3 (Recommendations for the authors):

      Major:

      (3.1) Lines 197-199: It is unclear why the authors conclude that somas have stronger representations of choice and task outcome. In Figure 4G, there is no significant difference between dendrites and somas for Choice or Outcome coding selectivity. The differences in the Sensory/Choice and Sensory/Outcome ratios shown in Figure 4H,I are likely explained by stronger Sensory selectivity in dendrites (Fig 4g), rather than by stronger Choice or Outcome encoding in somas.

      We agree with the reviewer’s interpretation of the data. The statement at 197 - 199 was meant to reflect relative selectivity, but it was imprecise. We have replaced that sentence with the following, more precise sentence:

      “Somatic activity also encoded these features, but the representation of the stimulus was weaker – and the representation of action timing was stronger – than in the tuft dendrites.”

      (3.2) Figure 4B: The authors realign FL-associated IRFs to GO-cue timing using the mean FL latency for each trial type and animal. Because FL timing is jittered across trials and may differ between CL and CR trials, this could smear the realigned traces and complicate the interpretation of contact-associated activity. The authors should consider using trial-by-trial FL timing for realignment or quantify the impact of FL-timing variability on the resulting traces.

      We aligned average GO- and contact-IRFs in Figure 4B so that comparisons of their magnitudes could be drawn from the same time window.

      With regard to jitter across trials, we think the reviewer may have misinterpreted how the mean IRFs in Figure 4B are calculated. The contact-IRF, by its nature, is calculated once per animal and trial type with respect to FL timing and shifted once based on mean FL latency. There is no smearing due to trial-to-trial FL timing.

      With regard to systematic differences in FL timing across animals and CR vs. CL, the reviewer is correct that this could – in theory – smear the realigned mean contact-IRF shown in Figure 4B. However, differences in mean FL latency across animals and trial-types are small compared with the long-timescale contact-IRFs. Thus, the non-realigned (i.e., always FL-aligned) mean contact-IRF looks nearly identical to Figure 4B just globally offset in time, as shown in Author response image 1:

      Author response image 1.

      Since this is nearly identical to data already presented in Figure 4B, we do not think it is necessary to include it in the revised article. However, we have added the following to the Methods:

      “Population averages of contact-IRFs that were not shifted prior to averaging were nearly identical (excluding the overall temporal shift; data not shown), indicating that pooling of mean IRFs across animals and trial types produces minimal smearing of the final population IRF.”

      (3.3) Figure 5A: Are CA trials specific to motor learning, or do they reflect a corrective lick toward the alternative port after an unrewarded lick? An analysis of the second lick on left-error trials or pre-shift right-error trials could help distinguish whether correction licking reflects a general decision change after failed reward, or a motor-command correction specific to post-shift motor learning. The authors should also report the prevalence of CA versus AP trials and clarify whether these trial types are behaviorally distinct.

      We thank the reviewer for highlighting the need to emphasize that CA trials reflect a distinct behavior related to reaching the displaced port.

      By definition, CA trials started as Motor Error trials and thus reflected a corrective lick toward the same port after an unrewarded lick. Almost all first contact licks on Motor Error trials were well outside the distribution of correct left licks both pre- and post-shift (Figure 1 - figure supplement 1B,D), consistent with the interpretation of this first lick as directed toward the right port. Thus, we see no evidence suggesting that CA trials involve a decision change. CA trials are exceedingly rare pre-shift, because Motor Error trials are rare pre-shift (Figure 1I, only ~5% of all right trials).

      With regard to other error types before the shift, most expert-trained mice did not immediately sample the other port with a second lick after an unrewarded lick. They usually either stopped licking immediately or licked the unrewarded port multiple times before switching ports. When port switches occurred pre-shift, timing was highly variable across mice and trials. Even on rewarded trials, some mice would “check” the unrewarded port after consuming the reward, as can be seen in Figure 3I, J. All of these behaviors are clearly distinct from the stereotyped second lick that occurred on CA trials after the shift. We agree that the prevalence of CA and AP trials, as well as the prevalence of immediate port alternation, should be reported, and we have added that information to the article as follows:

      “On Correction Attempted (CA) trials, the first lick made contact with the incorrect port, and the mouse chose to direct a second lick toward the correct port (Figure 5A; prevalence: 54% of motor error trials). We interpreted these licks as a corrective action, because the tongue exit angle shifted further toward the correct target (Figure 5B). Abandoned Port (AP) trials were the same as CA trials, except the mouse either did not make a second attempt or the second lick was directed toward the incorrect port (Figure 5A; prevalence: 46% of motor error trials).”

      (3.4) The classification of pre-shift errors into motor and decision errors is not clear. If error-trial exit angles follow a unimodal distribution (Figure 1- Figure Supplement 1C), then the distinction between motor and decision errors may not be behaviorally well separated. The authors should explain how these categories are validated and whether conclusions depending on this classification are robust to alternative definitions.

      We do not conclude that motor errors and decision errors are distinguishable pre-shift. Pre-shift licks were classified into motor error and decision error categories only to demonstrate that the boundary we established for classifying post-shift licks classifies extremely few (~5%, Figure 1I) pre-shift licks as motor errors. No conclusions were drawn from comparisons between pre-shift licks classified as decision errors and those classified as motor errors. The categorization is defined by the distribution of exit angles pre-shift and validated by the bimodal distribution of exit angles on error trials post-shift. To improve clarity regarding our classification of pre-shift errors, we have added the following to the Results:

      “Exit angles after the shift exhibited a bimodal distribution across error trials (Figure 1G,H; Figure 1—figure supplement 1C,D), supporting this distinction in error type. The frequency of licks classified as motor errors on right-cued trials increased significantly after the shift (median pre-shift 0.06, median post-shift 0.42, p < 0.001; Figure 1I; Figure 1—figure supplement 1C,D), reflecting the new challenge of avoiding the left lickport. In contrast to after the shift, exit angles on error trials before the shift were unimodal (Figure 1—figure supplement 1C). These errors were classified based on the fixed CB in order to demonstrate that very few pre-shift licks qualify as motor errors (Figure 1I), and not to suggest that tongue trajectories before the shift are behaviorally well-separated.”

      (3.5) Figure 1- Figure Supplement 1D, post-shift decision errors: Are these truly decision errors? The lick angles appear similar to those observed before the shift, suggesting that these trials may reflect execution of a "default" or "uncertain" lick trajectory rather than an incorrect choice under the new contingency.

      The post-shift exit angles on right-cued decision error trials (Figure1 - figure supplement 1D, grey) are similar to the lick angles on correct left-cued trials pre-shift (Figure 1 - figure supplement 1A, red) and clearly different from the correct right-cued trials pre-shift (Figure 1 figure supplement 1A, blue). Thus, to the extent that the animal’s intention can be measured from lick trajectory, it was targeting the incorrect (left) port. It is also true that it may still target the previous location of the left port (a “default” left trajectory), but because the decision error makes precise targeting irrelevant to the task outcome (it is easy to reach the left port after the shift), we do not designate it as a joint decision error and motor error. As to whether the deliberative process leading to this action is somehow cognitively distinct from other behaviors typically labeled as decision errors or incorrect choices, we cannot say.

      Minor:

      (3.6) Vocabulary consistency: soma vs somata.

      When data are shown for, or derived from, multiple somata, we use “somata”. When data are shown for an individual soma (such as in a panel with data from a single example soma), we use “soma.” We could not find any use of “somas,” which would indeed be inconsistent.

      (3.7) Figure 1C: I am not sure why the lick trajectories do not depict the tongue exiting the mouse. What time window is shown? Why does it look like the trajectories are shifted to the left?

      We thank the reviewer for identifying this issue. The definition of the location labeled “mouth” was accidentally omitted. The lick trajectories in Figure 1C do depict the tongue tip once it became visible to the cameras. Jaw opening and shifting partly determined the location where the tongue became visible in the videography. These movements varied from mouse to mouse and trial to trial, so exit angle was measured from the approximate midpoint between the temporomandibular joints, which is the grey point in 1C. We have fixed the captions and Methods to precisely define this location. With regard to the appearance of a slight leftward shift in the trajectories, this reflects how the tongue exits the mouth and how the tongue tip curves downward as the tongue approaches the port.

      (3.8) Figure 1- Figure supplement 1: it could ease the comparisons to report population statistics, such as median, from panel A to panel B and D, population statistics from B to D.

      Thank you. We have added these statistics to the Figure 1 - figure supplement 1 caption.

      (3.9) Choice boundary (CB) should be defined in line 100, not 110.

      Thank you. We have fixed this.

      (3.10) Line 109: claim not supported by referenced figure (Figure 1 - Figure Supplement 1). Lick angle histogram to the right port, pre-shift does not overlap substantially with lick angle to the left port, post-shift.

      We thank the reviewer for the opportunity to clarify this. We agree that Figure 1 - figure supplement 1 is not sufficient to support the claim. First, we want to make clear that Figure 1 - figure supplement 1 does not contradict the claim. The new location of the left port can obstruct the tongue during right-cued licks, regardless of the distributions of left licks pre- or post-shift. Second, to confirm that the new location of the left port would obstruct a substantial fraction of pre-shift right-cued lick trajectories, we measured the minimum distance between tongue trajectories and the post-shift location of the left port. Of pre-shift right-cued exit trajectories, 30 +/- 5% came within 1.25 mm – half of the combined tongue width (1.5 mm) and port width (1 mm) – of the port center.

      To make this claim more precise, we have changed the statement as follows:

      “Thus, on right-cued trials, mice continuing to follow the pre-shift motor plan would be biased to more frequently contact the new left port location (Figure 1E,F; 30 +/- 5% of pre-shift trajectories came within a tongue-width of the new location) and receive punishment (i.e., timeout).

      (3.11) Line 113: claim not supported by referenced figure. Figure 1G does not display error trials.

      We have changed the line to refer to “both correct and error trials”, such that reference to Figure 1G is also appropriate.

      (3.12) Figure 2 - Figure Supplementary 3 & method: how is noise estimated?

      Thank you. The following has been added to the Methods:

      “For Figure 2 - figure supplement 3, noise was estimated as the square-root of the geometric mean of the Welch power spectrum in a high-frequency band (0.25–0.5 times the frame rate; Giovannucci et al., 2019).”

      (3.13) Figure 2D: Was imaging during the shift epoch always performed in dendrites? If so, could the imaging schedule bias comparisons between dendritic and somatic activity during learning, especially given that mice show behavioral learning between early and late post-shift sessions (Figure 1J)?

      No, imaging during the shift was not always performed in the dendrites. The following has been added to the Methods to make clear that the post-shift data reflect dendritic and somatic imaging conducted on the day of the shift with roughly similar frequency:

      “For Figure 5 and Figure 6, which make comparisons between dendritic and somatic activity during the post-shift period, 67% of animals providing somatic data (4 of 6 mice) underwent somatic imaging on the day of the shift and 80% of mice providing dendritic data (8 of 10 mice) underwent dendritic imaging on the day of the shift.”

      (3.14) Lines 163-164, "we observed that the onset of tuft activity was consistently time-locked to the GO cue (vertical green line; Figure 3B). This was in contrast to somatic activity, which had more variable timing (Figure 3E)." The authors cite panels B and E in support of this point, but these appear to be example ROIs. It would be helpful to clarify how representative these examples are, since the corresponding population summaries in panels G and H do not make the effect immediately apparent.

      These examples are representative, as supported by the population summary of activity time-locked to the GO-cue versus port contact in Figure 4B.

      (3.15) Figure 4B: It could be useful to add the lick traces here as well. To allow the reader to have an idea of contact timing with respect to the Go cue and compare the sustain response with the licking pattern.

      We understand how this could be helpful. However, since these exact traces are already present in Figure 3I,J, we think that adding them to Figure 4 is unnecessary and would add complexity to an already very busy figure.

      (3.16) Figure 5E: Why are the imaging sessions labeled 0 and +1 rather than 0 and +2? Are the dendritic and somatic imaging not alternated?

      Yes, imaging was not alternated for 3 of the 22 mice. We have clarified this in the Methods, as follows:

      “Somatic and dendritic imaging sessions alternated every other day (19 of 22 mice), except for 3 mice in which only one compartment was imaged daily (dendrite-only: 2 mice, soma-only: 1 mouse). The exceptions were due to brain curvature or the angle of the coverslip with respect to the brain, such that only one compartment could be imaged and the other compartment was underneath skull regrowth or dural thickening that made high-quality imaging impossible.”

      (3.17) Figure 6C, legend: I suppose the authors meant "remapping", not "Post-shit SI distribution" for the description of the right column.

      Thank you. We have fixed this label.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study presents a useful finding on the effects of arginine vasopressin (AVP) on islet cells in pancreatic tissue slices, using technically sophisticated spatio-temporal calcium recordings to confirm that AVP influences α and β cells differently depending on glucose concentrations. While the study's methods - particularly the calcium imaging techniques and peptide ligand design targeting V1b receptors - are strong, the reviewers were concerned about several aspects of the experimental design. However, the results on βcell responses are incomplete and insufficient to support the manuscript's claims, especially due to the high variability of islet responses and lack of mechanistic and functional (hormone release) data. There are also concerns about the possibility of off-target effects and incomplete receptor specificity, noting that the study would have been significantly strengthened by inclusion of signaling pathway interrogation, hormone output assays, genetic validation (e.g., β cell-specific deletion of V1br), and receptor localization, although the work will still be of interest to researchers studying islet physiology in the context of health and diabetes.

      We sincerely thank the reviewers and editors for their thorough evaluation of our manuscript and their recognition of its technical strengths, including the advanced spatio-temporal calcium imaging and the rational design of selective V1b receptor ligands. We appreciate their acknowledgement of the study’s relevance for understanding AVP effects in a physiologically intact islet context and their positive assessment of our methodological rigour and innovation. The reviewers’ constructive feedback has helped us clarify the boundaries and intent of our study, which focuses on the glucose- and context-dependent modulation of α- and β-cell activity, rather than exhaustive molecular dissection.

      While the reviewers rightly emphasize the importance of receptor specificity and downstream signaling validation, we respectfully suggest that some of their concerns may reflect a lingering bias toward reductionist frameworks. Our interpretation is rooted in the emerging understanding that β-cell behaviour is largely defined by dynamic intercellular interactions within the islet collective, rather than by static gene expression or receptor localization alone (Jin et al., 2025; Korošak et al., 2021; Rutter et al., 2024). Recent studies have demonstrated that roles such as “leader” or “hub” β cells are transient and emergent, governed more by timing, environment, and local network structure than by fixed molecular identity (Postic et al., 2023; Gosak et al., 2018).

      This has profound implications for how we interpret cell responsiveness to agents like AVP: what appears as biological variability may in fact reflect context-sensitive transitions within a non-linear, self-organizing system (Stožer et al., 2021). Hence, we chose to focus on functional collective dynamics using intact pancreatic slices, rather than isolated cell models which fail to preserve the essential network architecture of islets. Although the addition of genetic models or isolated receptor measurements would strengthen receptor-specific conclusions, we argue that such approaches alone cannot resolve the physiological complexity of a system where function arises from cell–cell communication and spatiotemporal context.

      Indeed, the lack of direct correlation between receptor transcript abundance and functional outcomes has been noted in prior studies, reinforcing the view that function cannot be strictly predicted by molecular presence (Rutter et al., 2024). As articulated in our manuscript, the islet behaves as a sensory collective (Fancher & Mugler, 2017), where emergent patterns— not static cell identity—determine behaviour. This perspective aligns with broader shifts in biology away from strict genetic determinism toward causal emergence and collective agency (Ball, 2023; Levin, 2021).

      We therefore believe our study contributes not only new pharmacological insights but also a conceptual reframing of how AVP responses should be interpreted in a complex organ like the pancreas. We have added new data addressing reviewer suggestions—such as glucagon secretion assays, clarifications on the role of forskolin, and an analysis of event timing—that further support our conclusions. We also expanded the discussion on how islet variability is functionally meaningful, not just noise, and explained why β-cell responses to AVP must be interpreted within this probabilistic framework.

      We agree that future work should include receptor-specific knockouts and more direct signaling pathway assays, but these would need to be designed with careful consideration of the islet’s dynamic topology and the emergent nature of β-cell roles. In this light, we see our study not as the final word, but as a necessary systems-level foundation for more targeted interventions. We thank the reviewers again for their careful critiques and hope that our response clarifies both the rationale and scope of our work. Our revisions aim to enhance the paper’s clarity while maintaining its commitment to an integrative, physiology-rooted approach.

      We thank the reviewers and editors for their thoughtful and constructive assessment of our work. We are especially grateful for their recognition of the study’s technical strengths, including the use of spatio-temporal calcium imaging in intact pancreatic tissue and the strategic development of receptor-selective peptide ligands. We also appreciate their acknowledgement that our study contributes to the understanding of glucose-dependent AVP effects in islet physiology. The reviewers’ concerns regarding variability, receptor specificity, and functional validation helped us further clarify the scope and context of our study.

      We respectfully submit that some reservations stem from a reductionist framing that may not fully account for the collective behaviour of islets. As we and others have shown, β-cell function arises from emergent, self-organizing network dynamics, not just from static gene expression or receptor abundance (Jin et al., 2025; Korošak et al., 2021; Postic et al., 2023). In this view, pharmacological heterogeneity across islets is not simply noise or experimental inconsistency, but a signature of dynamic attractor states within the islet network (Stožer et al., 2021). Because an islet functions as a coupled system, most response variability originates from its emergent collective behavior, which eclipses variability in receptor expression or metabolic state.

      For this reason, even single-islet receptor quantification or ATP measurements would provide limited explanatory power: it is the state of the network—not absolute receptor levels—that determines whether a perturbation elicits activation or inhibition. As we illustrate in our graphical abstract, a single islet tested repeatedly under identical glucose conditions can yield divergent responses, simply because it occupies different dynamic states. These findings are in line with systems biology and network science approaches, which have revealed that cell function, especially in the β-cell collective, cannot be fully understood through reductionist parameters alone (Gosak et al., 2018; Ball, 2023).

      We have included glucagon secretion assays and new analyses to address key reviewer suggestions. Still, we chose not to pursue extensive knockouts or cAMP imaging, as these would require a different experimental scope and could risk disrupting the very dynamics we aim to understand. Likewise, while direct measurements of V1bR or IP3R expression would add molecular detail, they are not definitive without network context. The bell-shaped AVP dose-response curve and its explanation through IP3R inactivation are supported by prior studies; we invoke this mechanism not speculatively, but because it provides the most parsimonious explanation for the glucose-dependent shift in β-cell responsiveness.

      We also clarify that our study does not aim to resolve every mechanistic detail, but rather to offer a systems-level insight into how AVP modulates islet dynamics across varying glucose and cAMP contexts. The implications extend beyond AVP pharmacology, suggesting that perturbations to β-cell function must be understood within a probabilistic, state-dependent framework (Fancher & Mugler, 2017). This resonates with emerging concepts in cell physiology that emphasize causal emergence and local agency over static molecular determinism (Levin, 2021; Rutter et al., 2024).

      In summary, we see our work as part of a necessary shift in perspective—from linear receptor-function models to context-sensitive dynamic systems. We are grateful for the opportunity to revise our manuscript in response to insightful feedback and hope our clarifications and new data will strengthen its impact for the islet research community.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors confirmed earlier findings that AVP influences α and β cells differently, depending on glucose concentrations. At substimulatory glucose levels, AVP combined with forskolin - an activator of cAMP -did not significantly stimulate β cells, although it did activate α cells. Once glucose was raised to stimulatory levels, β cells became active, and α cell activity declined, indicating glucose's suppressive effect on α cells and permissive effect on β cells. Under physiological glucose levels (8-9 mM), forskolin enhanced β-cell calcium oscillations, and AVP further modulated this activity. However, AVP's effect on β cells was variable across islets and did not significantly alter AUC measurements (a combined indicator of oscillation frequency and duration). In α cells, forskolin and AVP led to increased activity even at high glucose levels, suggesting that α cells remain responsive despite expected suppression by insulin and glucose.

      Experiments with physiological concentrations of epinephrine suggest that AVP does not operate via Gs-coupled V2 receptors in β cells, as AVP could not counteract epinephrine's inhibitory effects. Instead, epinephrine reduced β cell activity while increasing α cell activity through different G-proteincoupled mechanisms. These results emphasize that AVP can potentiate αcell activation and has a nuanced, context-dependent effect on β cells.

      The most robust activation of both α and β cells by AVP occurred within its physiological osmo-regulatory range (~10-100 pM), confirming that AVP exerts bell-shaped concentration-dependent effects on β cells. At low concentrations, AVP increased β cell calcium oscillation frequency and reduced "halfwidths"; high concentrations eventually suppressed β cell activity, mimicking the muscarinic signaling. In α cells, higher AVP concentrations were required for peak activation, which was not blunted by receptor inactivation within physiological ranges.

      Attempting to further dissect the role of specific AVP receptors, the authors designed and tested peptide ligands selective for V1b receptors. These included a selective V1b agonist; a V1b agonist with antagonist properties at V1a and oxytocin receptors; and a selective V1a antagonist. In pancreatic slices, these peptides seem to replicate AVP's effects on Ca<sup>2+</sup> signaling, although responses were highly variable, with some islets showing increased activity and others no change or suppression. The variability was partly attributed to islet-specific baseline activity, and the authors conclude that AVP and V1b receptor agonists can modulate β cell activity in a statedependent manner, stimulating insulin secretion in quiescent cells and inhibiting it in already active cells.

      We applaud the reviewer to capture the essence of work in their introduction.

      Strengths:

      Overall, the study is technically advanced and provides useful pharmacological tools. However, the conclusions are limited by a lack of direct mechanistic and functional data. Addressing these gaps through a combination of signaling pathway interrogation, functional hormone output, genetic validation, and receptor localization would strengthen the conclusions and reduce the current (interpretive) ambiguity.

      Thank you!

      Weaknesses:

      (1) The study is entirely based on pharmacological tools. Without genetic models, off-target effects or incomplete specificity of the peptides cannot be fully ruled out.

      We partially agree with this comment and acknowledge that genetic models would provide a valuable complementary approach to address possible off-target effects or incomplete peptide specificity. However, genetic models also have important limitations, particularly when the aim is to resolve subtle, population-level physiological differences in beta cell activity. We therefore used pharmacological tools at different concentrations to test whether the observed effects were concentration-dependent and consistent with the expected receptor-mediated actions. An advantage of the pancreatic slice preparation is that it preserves much of the native tissue environment and allows pharmacological manipulation within concentration ranges closer to in vivo efficacy, thereby reducing the likelihood of nonspecific effects. To compensate for the lack of genetic models, we now emphasize the collective activity analysis as an additional strength of the study and have clarified this limitation in the revised manuscript.

      (2) Despite multiple claims about β cell activation or inhibition, the functional output - insulin secretion - is weakly assessed, and only in limited conditions. This aspect makes it very hard to correlate calcium dynamics with physiological outcomes.

      We agree that the functional output needed stronger support and have therefore expanded the hormone secretion experiments. While the effects of AVP and its analogues were tested during a stable plateau phase in the Ca<sup>2+</sup> imaging experiments, this phase provides only a narrow dynamic range for insulin release measurements in mouse slices. We therefore added a sequence of stimulations on the same slices, using 8 mM glucose and 500 nM forskolin, with glucose lowered to a non-stimulatory range between different AVP concentrations. These new experiments better define how AVP-dependent changes in Ca<sup>2+</sup> dynamics translate into insulin secretion under conditions with a broader secretory dynamic range. The new insulin and glucagon secretion data have now been added to the manuscript as Figure 5, and the text has been revised accordingly.

      (3) Insulin and glucagon secretion assays should be provided; the authors should measure hormone release in parallel with Ca2+ imaging, using perifusion assays, especially during AVP ramp and peptide ligand applications.

      We added insulin and glucagon secretion assays for AVP ramp to Figure 5.

      Additionally, there is no standardization of the metabolic state of islets. The authors should consider measuring islet NAD(P)H autofluorescence or mitochondrial potential (e.g., using TMRE) to control for metabolic variability that may affect responsiveness.

      We agree that standardization of the metabolic state of the islets would further strengthen the interpretation of the responsiveness data. We attempted to address this experimentally, but the results were inconclusive and therefore not included in the manuscript. Based on our previous unpublished observations, NAD(P)H levels appear to be significantly higher and less variable in islets within tissue slices than in isolated islets, suggesting that the slice preparation may better preserve the native metabolic state. However, we acknowledge that this remains an important limitation and we now indicate that in the manuscript. Additional experiments will be required to establish a robust and standardized approach, for example by combining NAD(P)H autofluorescence and/or mitochondrial potential measurements with Ca<sup>2+</sup> imaging.

      (4) There is a high degree of variability in response to AVP and V1b agonists across islets (activation, no effect, inhibition). Surprisingly, the authors do not fully explore the cause of this heterogeneity (whether it is due to receptor expression differences, metabolic state, experimental variability, or other conditions).

      This is a well-taken point and has indeed been one of the major bottlenecks in interpreting the results of this study. We agree that the variability in responses to AVP and V1b agonists may reflect several factors, including receptor expression, metabolic state, experimental conditions, and differences in the functional state of individual islets. However, our data also suggest that the beta cell population within an islet should be considered as a dynamic, non-linear system, in which even small differences in initial conditions or collective state can result in qualitatively different outcomes, including activation, no apparent effect, or inhibition. In this framework, the response to AVP is not determined by receptor expression alone, but by the current physiological context of the islet network. This is also why we believe that pharmacological tests are most informative when interpreted within a defined functional state rather than as isolated receptor-specific readouts. As indicated in the graphical abstract, apparently similar islets may occupy different dynamic states and therefore respond differently to the same Gq/PLC/IP3R stimulus. We have now expanded the discussion to make this interpretation more explicit and to acknowledge that receptor expression, metabolic variability, and experimental factors remain possible contributors that will require further targeted studies.

      The following text has been added to expand the discussion:

      “The heterogeneous responses observed across different islets, where some showed increased activity while others showed no detectable change or inhibition, could intuitively be attributed to variability in V1b receptor expression or signaling capacity among β cells. Such an explanation would be consistent with differences in receptor density, coupling efficiency to Gq proteins, or downstream signaling components such as PLC or IP<sub>3</sub> receptors. However, our data suggest that receptor-level variability alone is unlikely to fully explain the observed response spectrum, and that the current functional state of the islet collective must also be considered. The islet behaves as a non-linear dynamic system in which the same molecular perturbation can produce different functional outcomes depending on the current state of the β-cell collective. In such systems, cells or cell populations do not occupy a single deterministic activity state, but rather move within a landscape of possible states, with perturbations shifting the probability distribution of transitions between them. This concept is well established in dynamical systems approaches to biological cell-state transitions, where attractor landscapes, noise, and signaling inputs determine the probability of moving between alternative functional states rather than enforcing a single fixed output.

      In this framework, AVP and V1b receptor-selective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that βcell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone. Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (5) There is no validation of V1b receptor expression at the protein or mRNA level in α or β cells using in situ hybridization, immunohistochemistry, or spatial transcriptomics.

      We agree with the reviewer that spatial validation of V1b receptor expression is important for interpreting the cellular targets of AVP signaling in the islet. We have therefore added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      (6) AVP effects are described in terms of permissive or antagonistic effects on cAMP (especially in relation to epinephrine), but direct measurements of cAMP in α and β cells are not shown, weakening these conclusions. The authors should use Epac-based cAMP FRET sensors in α and β cells to monitor the interaction between AVP, forskolin, and epinephrine more conclusively.

      We agree that direct measurements of cAMP dynamics in alpha and beta cells would provide a more conclusive assessment of the interaction between AVP, forskolin, and epinephrine signaling. We attempted to address this experimentally; however, within the time domain of the Ca<sup>2+</sup> oscillations analyzed here, the temporal resolution and robustness of currently available cAMP readouts were not sufficient to resolve these interactions reliably. Even at slower time scales, cAMP sensor signals can be difficult to interpret quantitatively and may be overinterpreted if not tightly linked to the functional readout. We have therefore moderated the wording of the manuscript and now describe the proposed permissive or antagonistic interaction between AVP/V1b and cAMP-dependent signaling as an interpretation supported by the pharmacological Ca<sup>2+</sup> response patterns, rather than as a directly demonstrated cAMP mechanism. We now explicitly acknowledge in the limitations that most experiments were performed under cAMP-permissive conditions, which increases sensitivity for detecting AVP-dependent modulation but complicates the separation of direct beta cell effects from intra-islet interactions. Future studies using optimized cell-type-specific Epac-based sensors will be required to resolve this interaction.

      (7) Single-islet transcriptomics or proteomics (also to clarify variability) should be provided to analyze receptor expression variability across islets to correlate with response phenotypes (activation vs inhibition). Alternatively, the authors could perform calcium imaging with simultaneous insulin granule tracking or ATP levels to assess islet functional states.

      We agree that single-islet transcriptomics, proteomics, or simultaneous metabolic readouts could provide useful complementary information, particularly for describing molecular variability across islets. However, we do not think that differences in receptor expression or ATP levels alone are sufficient to explain the diversity of response phenotypes observed here. Our interpretation is that the beta cell population behaves as a collective dynamic system, in which the same input can lead to different outcomes depending on the current state of the network and its local physiological context. In such a system, AVP/V1b signaling does not necessarily impose a single deterministic response, but changes the probability distribution of accessible states, including activation, inhibition, or no detectable response. Theoretically and partially confirmed by the preliminary data, even the same islet exposed repeatedly under apparently identical conditions could be expected to display different responses if it occupies a different position within this dynamic state space at the time of stimulation. This concept is summarized in the graphical abstract and is central to our interpretation of the pharmacological data. We have therefore clarified in the Discussion that receptor expression, ATP levels, and other molecular parameters may modulate the response landscape, but are unlikely to fully define the observed functional phenotype without considering the collective dynamics of the islet.

      Added to Discussion section: “In this framework, AVP and V1b receptorselective agonists may reshape the probability landscape of β-cell activity. Depending on the initial metabolic, electrical, and Ca<sup>2+</sup>-handling state of the islet, the same stimulus may increase oscillation frequency, produce little detectable effect, or shift the system toward reduced activity or functional inactivation. This interpretation is also consistent with studies of pancreatic islet dynamics showing that β cell Ca<sup>2+</sup> activity emerges from coupled electrical, metabolic, and network interactions rather than from the properties of individual cells alone (63). Thus, molecular variability in V1b receptor expression or signaling capacity may contribute to the heterogeneous responses, but it is unlikely to determine them without considering the collective dynamic state of the islet.”

      (8) While the study implies AVP acts through V1b receptors on β cells, the signaling downstream (e.g., PLC activation, IP3R isoforms involved) is simply inferred but not directly shown.

      We agree that downstream signaling was not directly resolved at the level of PLC activation or specific IP3R isoforms. However, we did not infer Gq/PLC/IP3R involvement solely from AVP pharmacology, but used ACh as an independent Gq-coupled receptor reference stimulus in the same pancreatic slice preparation. With ACh concentration ramps, we could reproduce both activation and inactivation patterns observed with AVP/V1b stimulation, supporting the interpretation that these responses arise from modulation of the Gq-dependent Ca<sup>2+</sup> signaling axis.

      In addition, in prelilminary expriments we could observe that inhibition of Gq activity with YM254890, as well as interference with IP3R-dependent signaling using Xestospongin C, diminished the response, although not completely. This incomplete suppression is important, because it suggests that beta cell Ca<sup>2+</sup> homeostasis and collective islet activity are not controlled by a single linear pathway, but by partially redundant and context-dependent mechanisms. We have therefore revised the manuscript to state more cautiously that our data support the involvement of Gq/PLC/IP3R-dependent signaling, while acknowledging that direct measurements of PLC activity and IP3R isoform-specific contributions remain outside the scope of the present study.

      (9) The interpretation that IP3R inactivation (mentioned in the title!) underlies the bell-shaped AVP effect is just hypothetical, without direct measurements. Assays in β (and/or α)-cell-specific V1b KO mice and IP3R KO mice must be provided to support these speculations.

      We agree that the involvement of IP3R-dependent signaling should be stated with appropriate caution. However, the concept of IP3R inactivation as a mechanism contributing to bell-shaped Gq-dependent Ca<sup>2+</sup> responses is not purely hypothetical, since IP3R inactivation has been directly demonstrated in previous studies and provides a parsimonious explanation for the shift from activation to suppression at higher AVP concentrations. In the present study, this interpretation is further supported by the glucose dependence of the AVP concentration-response relationship, where different stimulatory glucose conditions shift the apparent efficacy peak.

      We also agree that cell-specific V1b receptor and IP3R knockout experiments would be valuable future approaches. In this respect, we have obtained preliminary results from a small sample of IP3R triple-knockout mice, which cannot yet be fully included because they are part of an ongoing collaboration. In these experiments, supraphysiological AVP concentrations did not produce the IP3R-like beta-cell response pattern observed in controls, namely reduced halfwidth and increased frequency, whereas alpha cell stimulation was preserved similarly to WT slices.

      At the same time, we believe that definitive knockout experiments must be carefully designed, because the beta cell population behaves as a dynamic collective system in which the response to AVP depends on the current functional state of the islet, glucose context, and intercellular coupling. We therefore now present IP3R inactivation as a strongly supported mechanistic interpretation rather than as a directly proven mechanism in this study, and we explicitly acknowledge that cell-specific V1b and IP3R genetic models will be required to fully resolve this pathway.

      Reviewer #2 (Public review):

      Summary:

      In this paper, Drs. Kercmar, Murko, and Bombek make a series of observations related to the role of AVP in pancreatic islets. They use the pancreatic slice preparation that their group is well known for. The observations on the slide physiology are technically impressive. However, I am not convinced by the conclusions of this manuscript for a number of reasons. At the core of my concern is perhaps that this manuscript appears to be motivated to resolve 'controversies' surrounding the actions of AVP on insulin and glucagon secretion. This manuscript adds more observations, but these do not move the field forward in improving or solidifying our mechanistic understanding of AVP actions on islets. A major claim in this manuscript is the beta cell expression of the V1b Receptor for AVP, but the evidence presented in this paper falls short of supporting this claim.

      Observations on the activation of calcium in alpha cells via V1b receptor align with prior observations of this effect.

      I have focused my main concerns below. I hope the authors will consider these suggestions carefully - please be assured that they were made with the intent to support the authors and increase the impact of this work.

      We thank the reviewer for their detailed input and support to increase the impact of our work and our understanding of important cellular processes overall. We have considered their suggestions carefully to further expand the strenghts of our approach and analysis.

      Strengths:

      The main strength of this paper is the technical sophistication of the approach and the analysis and representation of the calcium traces from alpha and beta cells.

      Thank you!

      Weaknesses:

      (1) The introduction is long and summarizes a substantive body of literature on AVP actions on insulin secretion in vivo. There are a number of possible explanations for these observations that do not directly target islet cells. If the goal is to resolve the mechanistic basis of AVP action on alpha and beta cells, the more limited number of papers that describe direct islet effects is more helpful. There are excellent data that indicate that the actions of AVP are mediated via V1bR on alpha cells and that V1bR is a) not expressed by beta cells and b) does not activate beta cell calcium at all at 10 nM - which is the same concentration used in this paper (Figure 4G) for peak alpha cell Ca2+ activation (see https://doi.org/10.1016/j.cmet.2017.03.017; cited as ref 30 in the current manuscript).

      We thank the reviewer for this important comment and agree that the literature on AVP actions in vivo is complex, with several possible sites of action outside the islet. We have therefore revised the Introduction to make the rationale more focused and to better separate systemic effects of AVP from studies addressing direct actions on pancreatic islet cells. At the same time, we chose not to restrict the Introduction only to the alpha cell V1bR literature, because one of the aims of the manuscript is precisely to address why AVP effects on insulin secretion have remained difficult to interpret across experimental contexts.

      Our results fully confirm a central aspect of the study cited by the reviewer, namely that V1bR activation robustly stimulates alpha cell Ca<sup>2+</sup> activity under non-stimulatory glucose conditions, and that 10 nM AVP does not produce a uniform activation of beta cell Ca<sup>2+</sup> activity. In fact, in a substantial fraction of beta cell populations, 10 nM AVP failed to activate oscillations, consistent with the view that alpha cells are the more sensitive and more direct cellular target of AVP/V1bR signaling. However, we do not think that the available transcriptomic evidence is sufficient to categorically exclude V1bR expression or functional relevance in beta cells. Re-analysis of the published dataset, together with more recent datasets and our newly added RNAscope data, supports a higher relative expression of V1bR transcripts in alpha than in beta cells, but does not justify treating beta (or non-alpha) cell expression as absent.

      We have therefore revised the manuscript to avoid overstating beta cell V1bR expression as a major isolated claim. Instead, we now present the data as evidence that AVP/V1bR signaling acts most prominently through alpha cells, while beta cell responses emerge in a concentration-, glucose-, and statedependent manner within the intact islet. This interpretation is consistent with the reviewer’s concern that 10 nM AVP preferentially activates alpha cells, but it also accommodates our observation that beta cell collective activity can be modulated under defined pharmacological and metabolic conditions. We believe that this is an important distinction, because the absence of a uniform beta cell Ca<sup>2+</sup> activation at one AVP concentration does not exclude beta cell modulation by AVP/V1bR signaling within the intact islet network. The Introduction and Discussion have been revised accordingly to clarify that our study does not simply challenge the alpha cell V1bR model, but expands it by examining how AVP-dependent alpha cell activation, possibly lower beta-cell receptor expression, and collective beta cell dynamics interact in the native pancreatic slice preparation.

      (2) We know from bulk RNAseq data on purified alpha, beta, and delta cells from both the Huising and Gribble groups that there is no expression of V2a. I will point you to the data from the Huising lab website published almost a decade ago (http://dx.doi.org/10.1016/j.molmet.2016.04.007) - which is publicly available and can be used to generate figures (https://huisinglab.com/dataghrelin-ucsc/index.html). They indicate the absence of expression of not only AVP2 receptors anywhere in the islet, but also the lack of expression of V1bra, V1brb, and Oxtr in beta cells. Instead of the detailed list of expression of these 4 receptors elsewhere in the body, it would be more directly relevant to set up their pancreatic slice experiments to summarize the known expression in pancreatic islets that is publicly available. It would also have helped ground the efforts that involved the generation of the V1aR agonist and V2R antagonist, which confirm these known AVP/OXT receptor expression patterns.

      We thank the reviewer for pointing us more directly to the publicly available islet expression datasets. We agree that the expression of AVP/OXT receptors in purified alpha, beta, and delta cells provides an important reference frame for interpreting our pharmacological data, and we have revised the manuscript to summarize these islet-specific datasets more directly rather than emphasizing receptor expression in other organs. These data support the absence or very low expression of V2 receptors in islet endocrine cells and confirm that V1b receptor expression is substantially enriched in alpha cells compared with beta cells.

      At the same time, as outlined in our response above, we do not think that the currently available transcriptomic datasets are sufficient to categorically exclude low-level V1bR transcript expression or functional relevance in beta cells within the intact islet. For this reason, we added independent RNAscope validation to assess V1bR transcripts in the pancreatic slice preparation. These data confirm stronger V1bR expression in glucagon-positive alpha cells, while also showing a broader expression pattern within the islet and pancreas.

      We have also revised the rationale for the pharmacological experiments using V1aR- and V2R-directed tools. We now present these experiments not as evidence for unexpected receptor expression, but as functional controls that are consistent with the known AVP/OXT receptor expression patterns in pancreatic islets. This better aligns the manuscript with the existing transcriptomic literature while preserving the main physiological question of the study: how AVP/V1bR-dependent signaling reshapes alpha-cell activity and beta-cell collective dynamics in intact pancreatic tissue.

      (3) Importantly, the lack of V1br from beta cells does not invalidate observations that AVP affects calcium in beta cells, but it does indicate that these effects are mediated a) indirectly, downstream of alpha cell V1br or b) via an unknown off-target mechanism (less likely). The different peak efficacies in Figure 4G would also suggest that they are not mediated by the same receptor.

      We agree with the reviewer that the absence or very low abundance of V1bR transcripts in beta cells in published transcriptomic datasets would not invalidate the observation that AVP modulates beta-cell Ca<sup>2+</sup> activity. It does, however, raise the important question of whether this modulation is mediated indirectly through alpha-cell V1bR activation, through V1bR expression in beta cells that is difficult to resolve transcriptomically, or through another mechanism. To address this more directly, we have now added RNAscope data, which confirm relatively stronger V1bR transcript enrichment in glucagon-positive alpha cells, but also show a broader V1bR transcript signal within the islet and pancreas. Thus, while our data support alpha cells as the dominant V1bR-positive endocrine population, they do not support a strict absence of V1bR-associated signaling capacity in the beta cell compartment.

      We also agree that different peak efficacies in alpha and beta cells could be interpreted as evidence for distinct receptors or indirect mechanisms. However, we favor a different interpretation: the apparent efficacy of AVP depends strongly on the physiological state in which the cells are tested. This is particularly evident in beta cells, where the AVP efficacy peak shifts with glucose concentration, suggesting that the beta-cell response is shaped by the metabolic and Ca<sup>2+</sup>-handling context rather than by receptor occupancy alone. In this framework, the same V1bR/Gq-dependent input can generate different downstream Ca<sup>2+</sup> outcomes in alpha and beta cells because the two cell types operate in different dynamic regimes.

      We have therefore revised the manuscript to acknowledge this dilemma more explicitly. We now state that beta cell effects of AVP could include indirect alpha cell-dependent components, but given the magnitude and statedependence of the beta cell Ca<sup>2+</sup> response it is unlikely to be driven by alpha cell activation. Instead, our preferred interpretation is that AVP/V1bR signaling acts within the intact islet as a context-dependent perturbation of the collective beta cell Ca<sup>2+</sup> system, with IP3R-dependent mechanisms being modulated by glucose-dependent changes in beta cell excitability and intracellular Ca<sup>2+</sup> handling.

      (4) The rationale for the use of forskolin across almost all traces is unclear. It is motivated by a desire to 'study the AVP dependence of both alpha and beta cells at the same time'. As best as I can determine, the design choice to conduct all studies under sustained forskolin stimulation is related to the permissive actions of AVP on hormone secretion in response to cAMPgenerating stimuli. The permissive actions by AVP that are cited are on hormone secretion, which in many cell types requires activation of both calcium and cAMP signaling. Whether the activation of V1br and subsequent calcium response is permitted by cAMP is unclear. I believe the argument the authors are making here is that the activation of beta cell calcium by AVP is permitted by forskolin. i.e., the cAMP stimulated by it in beta cells. However, the design does not account for the elevation of cAMP in alpha cells and subsequent release of glucagon, particularly upon co-stimulation with AVP, which permits glucagon release by activating a calcium response in alpha cells. This glucagon could then activate beta cells. If resolving the mechanism of action is the goal, often less is more. The activation of Gaq-mediated calcium is not cAMP dependent (although the downstream hormone secretion clearly often is). As was shown, AVP does not activate calcium in beta cells in the absence of cAMP. The experiments in Figures 1, 2, and 4 should have been completed in the absence of cAMP first.

      We agree with the reviewer that the use of forskolin needs to be explained more clearly, and we have revised the manuscript accordingly. Our rationale was based on the established permissive role of cAMP in AVP-dependent endocrine responses, but we acknowledge that this does not necessarily imply that the upstream V1bR/Gq-mediated Ca<sup>2+</sup> response itself is cAMP-dependent. The reviewer is also correct that forskolin elevates cAMP broadly and therefore may affect both alpha and beta cells, including the possibility that AVP-enhanced alpha cell activation and glucagon release secondarily influence beta cell activity.

      In fact, our initial experiments were performed without forskolin and revealed an important difficulty: stimulatory glucose alone can increase cAMP levels to a variable extent, as also supported by our previous work on epinephrine signaling, thereby shifting the apparent peak efficacy of AVP stimulation. Thus, forskolin was originally used to reduce this variability and create a more defined cAMP-permissive background in which alpha and beta cell responses could be compared in the same slice. However, we agree that this design works against isolation of beta cell-autonomous AVP effects.

      Within a scope of another study we have done an independent series of more focused experiments using GLP-1 receptor stimulation, which preferentially increases cAMP signaling in beta cells compared with the broad cAMP elevation produced by forskolin. We have clarified that the modulation of the AVP-dependent pathway by GLP-1 and related ligands at largely supports beta cell-autonomous AVP effects. It is part of ongoing work and will be reported independently, because a full mechanistic dissection of cAMP–AVP interactions goes far beyond the scope of the present study.

      (5) It is unexpected that epinephrine in Figure 2 does not activate the alpha cell calcium? A recent paper from the same group (Sluga et al) shows robust calcium activation in alpha cells in a similar prep by 1 nM epinephrine, which is similar to the dose used here.

      We thank the reviewer for pointing this out, but we would like to clarify that epinephrine did significantly activate alpha-cell Ca<sup>2+</sup> activity in our experiments, as shown in Fig. 3F. This result is consistent with our previous study by Sluga et al., where low nanomolar epinephrine robustly activated alpha cell Ca<sup>2+</sup> signals in the pancreatic slice preparation. The main point of the present comparison was therefore not that epinephrine is inactive in alpha cells, but that AVP produces a substantially stronger and reproducible alpha cell Ca<sup>2+</sup> response under comparable experimental conditions. This is also consistent with the data of van der Meulen et al., supporting the view that AVP/V1bR signaling is a particularly potent activator of alpha cell activity. We have revised the text to make this comparison clearer and to avoid the impression that epinephrine failed to activate alpha cells in our preparation.

      (6) Figure 8 suggests a pharmacological activation of beta cell V1bR in the low pM range. How do the authors reconcile this comparison with the apparent absence of an effect of AVP stimulation at low pM to low nM doses in beta cells (Figure 4A)? I note that there are changes over time with sustained beta cell stimulation with 8 mM glucose, but these changes are relatively subtle, gradual, and quite likely represent the progression of calcium behaviors that would have occurred under sustained glucose, irrespective of these very low AVP concentrations. I will note that the Kd of the V1bR for AVP is around 1 nM, with tracer displacement starting around 100 pM according to the data in figure 5B, which is hard to reconcile with changes in beta cell calcium by AVP doses that start 10-100-fold lower than this dose at 1 and 10 pM (Figure 8).

      We agree that the interpretation of low-pM AVP effects requires caution, particularly when compared with reported V1bR binding affinities. The apparent discrepancy between Fig. 4A and Fig. 8 most likely reflects differences in experimental design, stimulation context, and readout sensitivity. In Fig. 4A, we assessed acute AVP effects under conditions in which beta cell Ca<sup>2+</sup> responses are relatively threshold-dependent and where low AVP concentrations produced little or no activation. In contrast, Fig. 8 analyzes prolonged beta cell population dynamics during sustained stimulation with 8 mM glucose, a physiological stimulatory context in which even weak modulatory inputs may become detectable at the level of collective Ca<sup>2+</sup> activity.

      Importantly, the strongest and statistically significant effect was observed at 100 pM AVP, while lower pM concentrations showed only a trend. We therefore do not interpret the low-pM range as evidence for robust direct pharmacological activation of beta cell V1bR. Rather, these data suggest that AVP may exert permissive or modulatory effects within an already active beta cell network, where glucose-dependent excitability, receptor-effector coupling, and Ca<sup>2+</sup> amplification mechanisms can enhance the apparent efficacy of weak inputs. This interpretation is consistent with the known permissive role of AVP in endocrine responses, where AVP may not act as a primary activator alone but can increase the efficacy of other physiological stimuli.

      We have also clarified that sustained 8 mM glucose alone does not account for these effects, since Suppl. Fig. 1 shows no comparable time-dependent progression of Ca<sup>2+</sup> behavior under sustained glucose stimulation alone. Thus, we now present the low-concentration AVP effects as subtle, contextdependent modulation within the physiological stimulatory range, rather than as evidence for direct beta cell activation at concentrations below the expected receptor affinity range.

      Reviewer #3 (Public review):

      Summary:

      This work aims to better understand the role of arginine vasopressin (AVP) in the control of islet hormone secretion. This builds on previous literature in this area reporting on the actions of AVP to stimulate islet hormones. The gap in literature being addressed by these studies is primarily focused on the glucose-dependency of AVP on both insulin and glucagon secretion. A secondary objective is to explore the role of individual receptors with the use of newly generated peptides and existing tools. The methods include the use of Ca2+ imaging in pancreas slices from mice, with additional outcomes including insulin secretion in some areas. The conclusions presented are that AVP acts through V1b receptors in both alpha- and beta-cells, that this activity occurs in the high cAMP environment, and is glucose dependent.

      Strengths:

      The area of research is emerging with plenty of room for new contributions. The concept of AVP stimulating islet hormone secretion is important and deserving of further insight. The use of pancreas tissue to image primary cells makes the experiments physiologically relevant. The advancement of novel tools in this area should be helpful to other groups investigating the actions of AVP.

      We would like to thank the reviewer for recognizing the potential of our emerging area of research.

      Weaknesses:

      The conclusions are only modestly supported by the data and lack experimental depth and rigor. The rationale for only conducting studies at high cAMP conditions is not entirely clear and limits the conclusions that can be made. The use of Ca2+ is helpful, but it is a surrogate for hormone secretion. Additional measurements of hormone secretion are needed to enhance the robustness of these conclusions. Consideration of paracrine effects between alpha- and beta-cells is only superficially made and is likely essential in the context of the experimental design. For instance, there is clear literature that alpha-cells secrete several factors that work in paracrine interactions on beta-cells and autocrine actions back on alpha-cells. Conducting these studies in a high cAMP context only completely overlooks these interactions, skewing the interpretations made by the investigators. Finally, the clarity of the experiments and results could be significantly enhanced.

      We thank the reviewer for this balanced assessment and for emphasizing several issues that are central to the interpretation of our study. We agree and now explicitly state in the Limitations, that Ca<sup>2+</sup> oscillations are a surrogate readout for hormone secretion and that currently used stimulation protocols are not optimized to directly quantify the relationship between Ca<sup>2+</sup> dynamics and secretory output. To address this limitation, we have now expanded the functional part of the study by adding complete insulin and glucagon secretion measurements during AVP concentration ramps. These new data provide a stronger functional framework for interpreting the Ca<sup>2+</sup> imaging results, while also clarifying that Ca<sup>2+</sup> activity and secretion cannot be assumed to correlate linearly under all stimulation protocols.

      We have also revised the rationale for the high-cAMP experimental condition. The original aim was to reduce variability arising from glucose-dependent endogenous cAMP signaling and to study alpha and beta cell responses in a common permissive background. However, we agree that broad forskolin stimulation complicates the interpretation of cell-autonomous versus paracrine mechanisms. Independent experiments within a scope of another study demonstrate that using GLP-1 co-stimulation, which provides a more beta-cell-oriented cAMP-permissive condition and supports the interpretation that AVP can modulate beta-cell collective activity in a manner that is not solely secondary to alpha-cell activation.

      We fully agree that paracrine interactions within the islet are physiologically important and must be considered, particularly in intact pancreatic slices. Nevertheless, the rapid onset of the AVP effects observed in beta cell Ca<sup>2+</sup> activity argues against a mechanism mediated predominantly by slower indirect paracrine loops. High AVP concentrations, as shown in Fig. 5, significantly shorten the intervals between Ca<sup>2+</sup> events in both alpha and beta cells, but that the activity of the two cell populations remains largely noncoordinated. This temporal dissociation does not exclude paracrine modulation altogether, but it argues against a simple alpha-cell-driven explanation for the beta-cell response.

      We have revised the manuscript to state these points more clearly and to moderate conclusions where the data support modulation rather than definitive cell-autonomous receptor action. We believe that the added secretion experiments and alpha/beta event-timing analysis substantially strengthen the physiological interpretation of the study, and we thank the reviewer for raising these issues.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) The paragraph discussing the benefits of slice physiology over islets is not reflective of how most - if not all of your colleagues who do islet experiments conduct these. Many labs have reported for years high-quality GSIS experiments, synchronous calcium responses, and a plethora of studies detailing the mechanism of hormone and neurotransmitter actions using islet models, and have done so well. Slice physiology is a unique and helpful model that can have advantages over other models. This particular reviewer uses both models in their lab and each has benefits and - inevitably - drawbacks. Many of the possible drawbacks cited for islet studies apply equally to slices, including the possibility of altered gene expression, lack of innervation, and circulation. Added drawbacks are the exposure to higher levels of pancreatic enzymes from the slice, which require co-culture with enzyme inhibitors.

      We agree with the reviewer and have revised these limitations accordingly. Our intention was not to imply that isolated islet preparations are generally inferior, since they have provided a highly productive and rigorous experimental platform for GSIS, synchronized Ca<sup>2+</sup> dynamics, and mechanistic studies of hormonal and neurotransmitter regulation with standardized protocols with all their positive and negative sides. We modified the presentation of pancreatic slices as a complementary model with specific advantages, particularly preservation of local tissue architecture, while also acknowledging their limitations. The revised text therefore avoids a comparative hierarchy between slices and isolated islets and instead emphasizes that both models have distinct strengths and drawbacks depending on the experimental question.

      (2) If you want to demonstrate direct actions on beta cells, deconstructing the islet would be a better way to go. Less complicated, not more. Dissociated beta cells, instead of slices, were used just to prove or disprove the hypothesis of direct beta cell effects of AVP.

      We agree that dissociated beta cells can be a useful reductionist model to test whether AVP is capable of acting directly on individual beta cells. However, this approach would also remove the collective beta-cell activity that is central to the physiological question addressed in the present study. Since our data indicate that AVP effects emerge within the intact islet as rapid and extensive changes in coordinated Ca<sup>2+</sup> dynamics, dissociation would not necessarily provide a more informative model for understanding these responses. The fast onset and magnitude of the beta cell response argue against a predominantly indirect non-autonomous mechanism, and this interpretation is further supported by the GLP-1 co-stimulation experiments, which are more consistent with beta cell-autonomous modulation. We have therefore clarified in the revised manuscript that dissociated-cell experiments would be valuable for a narrowly defined receptor-cell autonomy question, but would not resolve the collective islet dynamics that are the focus of this work.

      (3) If you want to sustain the claim of beta cell expression of V1br, you would have to demonstrate this far more directly by staining (if appropriate antibodies exist), by beta cell-specific deletion of V1br, or by highly selective, well-validated pharmacology. This should include a demonstration of Gaqdependence in isolated beta cells.

      We have added RNAscope in situ hybridization data to the revised manuscript to assess V1b receptor mRNA expression within the pancreas and islet. These new data show a broader expression pattern of V1b receptor transcripts within the islet than originally assumed, suggesting that AVP signaling may not be restricted to a single endocrine cell population. At the same time, the RNAscope analysis confirms previous reports of higher AVP receptor expression in glucagon-positive alpha cells. We have added these results to Figure 1 and revised the corresponding Results and Discussion sections to clarify that the observed functional responses may reflect both direct effects on beta cells and indirect intra-islet effects mediated through alpha-cell signaling.

      Minor

      (1) O'Carroll et al. should be cited in the context of islet permissive actions of AVP/cAMP. PMID: 18434353, although that paper offers no evidence that the AVP-dependent potentiation of insulin release is mediated directly by beta cells. It does confirm dependence on PKC.

      We agree and have added O’Carroll et al. in the revised manuscript in the context of AVP/cAMP-dependent permissive actions on islet hormone secretion. We therefore use it as support for the broader concept of AVP-dependent amplification of secretion in a permissive signaling context, rather than as direct evidence for beta-cell-autonomous V1bR signaling.

      (2) Figure 2 E-H, glucose concentration mislabeled.

      We thank the reviewer for pointing this out. The glucose concentration label has been clarified: panels E–H show pooled data from separate experiments performed at 8 mM glucose, whereas panel D shows a representative experiment performed at 9 mM glucose.

      (3) The insulin secretion in 4E is difficult to interpret without a low-glucose control. If this is hard to do in a slice preparation, a separate static islet secretion experiment would help here. The possibility that the inhibition of insulin secretion traces back to the activation of delta cells by AVP could be considered - I struggle to come up with a plausible mechanistic explanation why AVP (which activates calcium in alpha and in beta cells in the presence of 8 mM G plus forskolin according to your data) would inhibit insulin secretion.

      We agree that the original insulin secretion experiment was difficult to interpret without a clearer low-glucose reference condition. To address this, we have added new insulin release experiments in which glucose was lowered to a non-stimulatory range between AVP concentrations, followed by sequential stimulation with 8 mM glucose and 500 nM forskolin in the same slices. These new data provide a broader dynamic range for assessing insulin secretion and allow a more direct comparison between AVP-dependent Ca<sup>2+</sup> modulation and secretory output.

      We also agree that AVP-dependent inhibition of insulin secretion requires careful interpretation. One possible explanation is not simply activation of delta cells, but a failure of the beta cell collective to maintain coordinated activity at very high AVP concentrations. In the Ca<sup>2+</sup> imaging data, high AVP concentrations increase activity in many beta cells, but numerous cells within the islet fail to keep pace with the collective oscillatory rhythm, leading to fragmented and less synchronized population activity. Thus, despite increased frequency of Ca<sup>2+</sup> oscillations in the islet, the integrated beta cell output may become less efficient for insulin secretion. We have added this interpretation to the revised manuscript and now discuss delta cell activation as a possible contributing mechanism, but not as the primary explanation supported by our current data.

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes a chemical screen for activators of the eIF2 kinase GCN2 (EIF2AK4) in the integrated stress response (ISR). Recently, reported inhibitors of GCN2 and other protein kinases have been shown at certain concentrations to paradoxically activate GCN2. The study uses CHO cells and ISR reporter screens to identify a number of GCN2 activator compounds, including a potent "compound 20." These activators have implications for the development of new therapies for ISR-related diseases. For example, although not directly pursued in this study, these GCN2 activators could be helpful for the treatment of PVOD, which is reported for patients with certain GCN2 loss-of-function mutations. The identified activators are also suggested to engage with the GCN2 directly and can function while devoid of GCN1, a co-activator of GCN2.

      Strengths:

      The manuscript appears to be a largely rigorous study that flows in a logical manner. The topic is interesting and significant.

      Weaknesses:

      Portions of the manuscript are not fully clear. Some experimental presentation and design concerns should be addressed to support the stated conclusions.

      We thank the reviewer for their supportive comments. We agree that portions of the manuscript were not fully clear and that some aspects of the experimental presentation and design required clarification.

      To address this, we have revised the manuscript to make the experimental logic more transparent. In particular, we now explain more clearly the rationale for the screening strategy, including the use of histidinol as a canonical GCN2 activator, latrunculin A as a modulator of PPP1R15A-mediated eIF2α dephosphorylation, and tunicamycin as a PERK-dependent ER-stress control. We also clarify why a submaximal concentration of histidinol was used: this was intended to reveal compounds that enhance ISR signalling when GCN2 is partially activated.

      We have clarified the use of the two ATF4 reporter systems. The ATF4–NanoLuc reporter was used for sensitive primary screening, whereas the ATF4–luc2 reporter was used as a more stringent orthogonal assay to prioritise robust ISR activators. We now state explicitly why some initial hits were not retained after testing in the second reporter line, and why the NanoLuc system was subsequently used again for mechanistic experiments.

      Finally, we have revised the presentation of the orthogonal validation steps to make clearer how they support the stated conclusions. These include assays designed to distinguish GCN2-dependent ISR activation from indirect activation through ER stress, additional analysis of GCN2 dependence, and clearer interpretation of biochemical and docking data.

      We hope that these revisions address the reviewer’s concern that the experimental design and data presentation needed to be made clearer in order to support the conclusions.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Zhu, Emanuelli, and colleagues describe a novel pharmacological activator of the Integrated Stress Response kinase GCN2. The work is conclusive and biochemically solid. This work significantly adds to the pharmacological arsenal targeting the ISR and, in particular, GCN2.

      Strengths:

      Strong biochemistry, novel molecular activator of GCN2 (GCN1 independent).

      Weaknesses:

      The rationale for the screen is not exploited in the results (e.g., pathogenic GCN2 mutants), and lots of cell-based read-outs are not endogenous.

      We thank this reviewer for their positive assessment of the work. We address the three major concerns in turn below.

      Major points

      (1) Regarding the justification of the work. Since the authors justify the screen for GCN2 activators with loss-of-function mutants associated with diseases, it would be of interest to evaluate whether the best compounds identified in the study are indeed able to prompt activation of those mutants (or at least of the most prevalent). This approach could actually go in parallel with the docking experiments carried out in the last figure of the manuscript, where mutants could be modelized as well.

      To address this point, we tested whether the lead compounds could activate disease-associated GCN2 variants linked to pulmonary veno-occlusive disease. In contrast to GCN2iB, the new compounds did not activate these variants. We now state this explicitly in the manuscript, thereby clarifying that although the compounds identify a new mode of GCN2 activation, they do not rescue the pathogenic GCN2 variants tested here.

      Results

      “Moreover, in contrast to GCN2iB (17), the current compounds did not activate disease-associated GCN2 variants linked to PVOD [data not shown].”

      (2) The compounds are only tested using « artificial » proximal signaling outputs. It would be interesting to evaluate whether the best identified compounds are capable of prompting endogenous eIF2alpha phosphorylation in cellular models.

      We thank the reviewer for this suggestion. Detecting eIF2α phosphorylation following activation of GCN2 is technically challenging and typically produces weaker signals compared to activation of other ISR kinases, such as PERK (e.g. by thapsigargin). For this reason, many studies rely on downstream reporter assays to monitor GCN2 activity. To address the reviewer’s concern, we have now included an orthogonal readout of ISR activation by assessing global translation using a puromycin incorporation assay. Using this approach, we show that compound 20 significantly reduces translation, and importantly, this effect is attenuated in GCN2-deficient cells, supporting a GCN2-dependent mechanism.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      (3) Other GCN2 activators (other than GCN2iB, e.g., HC-7366) were recently identified. In this context, it would be of interest to carry out a small benchmarking study to evaluate how the compounds identified in the current study perform against the previously identified molecules.

      We thank the reviewer for this suggestion. In response, we obtained HC-7366 and assessed its activity alongside our compounds in the CHO ATF4::NanoLuc reporter assay. In this system, compound 20 demonstrated greater potency than HC-7366 (see reviewer figure below). However, we note that HC-7366 showed relatively limited activity in CHO cells in our hands, despite previously reported strong effects in other cellular systems and in vivo models. This context-dependent activity makes direct benchmarking difficult. Accordingly, we have included this comparison in the revised manuscript and discuss this limitation in the Discussion.

      Discussion

      “We also evaluated the reported GCN2 activator HC-7366 in our CHO ATF4::NanoLuc reporter system. In this context, HC-7366 showed limited activity relative to compound 20, despite its reported efficacy in other cellular systems and in vivo (data not shown). This highlights potential context dependence in small‑molecule activation of GCN2 and limits direct cross-study comparison.”

      Author response image 1.

      ISR activation by compound 20 and GC-7366 in CHO cells

      Normalised fold-change in ATF4 signal in CHO ATF4::NanoLuc reporter cells treated for 19 hours with Compound 20 or HC-7366. DMSO was used as vehicle control. (representative experiment, mean ± SEM, n=3 technical replicates).

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors describe the results of a high-throughput screen for small-molecule activators of GCN2. Ultimately, they find 3 promising compounds. One of these three, compound 20 (C20), is of the most interest both for its potency and specificity. The major new finding is that this molecule appears to activate GCN2 independent of GCN1, which suggests that it works by a potentially novel mechanism. Biochemical analysis suggests that each binds in the ATP-binding pocket of GCN2, and that at least in vitro, C20 is a potent agonist. Structural modeling provides insight into how the three compounds might dock in the pocket and generates testable hypotheses as to why C20 perhaps acts through a different mechanism than other molecules.

      We agree that GCN1-independent activation suggests a potentially distinct mechanism of action. While we are currently unable to define the mechanistic basis underlying the GCN1-independence of compound 20, prior work provides some relevant context. Recent studies have shown that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions, notably in the context of the GCN2 E26A mutant [Carlson, 2023]. This observation raises the possibility that, under certain conditions, engagement of the kinase domain, potentially via the ATP-binding pocket, may bypass the requirement for GCN1. However, in our system, we did not observe GCN1-independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow or context-dependent window for such activity, or differences between wild-type and mutant GCN2. These findings suggest that GCN1-independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds, although further work will be required to define the underlying mechanism for compound 20. We have added the following to the main text:

      Discussion

      “Recent work suggests that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions using a GCN2 E26A mutant (9). In our hands, we did not observe GCN1‑independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow concentration window for GCN1‑independent activation or context‑dependent effects of the E26A mutation. These findings raise the possibility that GCN1‑independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds.”

      Strengths:

      Of the 3 compounds identified by the authors, C20 is the most interesting, not just for its intriguing mechanistic distinction as being GCN1-independent (shown genetically in two distinct cell lines, CHO and 293T in Figure 4, and in contrast to other GCN2 activators) but also for its potency. In in-cellulo assays, compound 21 appears as more of an ISR enhancer than an activator per se, and although compound 18 and compound 21 lead to upregulation of the ISR targets (Figure 2), that degree of upregulation is probably not significantly different from that induced by those compounds in Gcn2-/- cells. For C20, the effect appears stronger (although it is unclear whether the authors performed statistical analysis comparing the two genotypes in Figure 2D). In Figure 3, only C20 activates the ISR robustly in both CHO and 293T. Ultimately, C20 might be a tool for providing mechanistic insight into the details of GCN2 activation and regulation, and could be exploited therapeutically.

      Prompted by this suggestion, we assessed C20 in two additional commonly used cell lines: human colon carcinoma HCT116 cells and African green monkey COS7 cells. C20 showed no activity in these models. In contrast, primary mesothelioma cells (Mesobank T12) exhibited robust PPP1R15A induction in response to the compound. The following text has been added to the manuscript.

      Results

      “We went on to examine downstream cellular consequences of GCN2 activation in multiple models. While compounds did not induce detectable ISR signalling in HCT116 or COS‑7 cells under the conditions tested, induction of PPP1R15A was observed in Mesobank T12 primary mesothelioma cells, indicating context-dependent biological responses… [data not shown].”

      Weaknesses:

      There are some limitations to the existing work. As the authors acknowledge, they do not use any of the compounds in animals; their in vivo efficacy, toxicity, and pharmacokinetics are unknown. But even in the context of the in cellulo experiments, it is puzzling that none of the three compounds, including C20, has any effects in HeLa cells when Neratinib does. It's beyond the scope of this paper to address definitively why that is, but it would at least be reassuring to know that C20 activates the ISR in a wider range of cells, including ideally some primary, non-immortalized cells. In addition, the ISR is a complex, feedback-regulated response whose output varies depending on the time point examined. The in cellulo analysis in this paper is limited to reporter assays at 18 hours and qRT-PCR assays at 4 and 8 hours. A more extensive examination of the behaviour of the relevant ISR mRNAs and proteins (eIF2, ATF4, CHOP, cell viability, etc.) for C20 across a more extensive time course would give the reader a clearer sense of how this molecule affects ISR output.

      We thank the reviewer for this insightful suggestion. To address the need for a more comprehensive assessment of ISR signalling, we have extended our analysis across a broader time course and incorporated additional functional readouts. Using the ATF4-NanoLuc reporter, compounds 20 and 21 exhibit peak ISR activation at approximately 6-8 h in wild-type cells. In parallel, we assessed global mRNA translation using puromycin incorporation and found that compound 20, but not compound 21, induces a progressive reduction in translation over this period, which is dependent on GCN2. While we agree that direct measurement of upstream ISR markers such as eIF2α phosphorylation can be informative, detection of GCN2-mediated eIF2α phosphorylation is technically challenging and often less robust than activation of other ISR kinases (e.g. PERK). For this reason, we have prioritised orthogonal downstream functional readouts, including reporter activity and translational output, to capture ISR pathway engagement. These additional data provide a clearer picture of the kinetics and functional consequences of compound-induced ISR activation and have been incorporated into the revised manuscript.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      I also find it a bit strange that the authors describe C20 as "demonstrat(ing) weak inhibition of ... PKR" - the measured IC50 is ~4 μM, which is right around its EC50 for GCN2 activation. This raises the confounding possibility that C20 would simultaneously activate GCN2 while inhibiting PKR. While perhaps inhibition of PKR is not relevant under the conditions when GCN2 would be activated either experimentally or therapeutically, examining in cells the effects of C20 on GCN2 and PKR across a dose range would shed light on whether this cross-reactivity is likely to be of concern.

      We thank the reviewer for highlighting compound 20 as the most interesting lead compound and for recognising its apparent ability to activate GCN2 independently of GCN1. The reviewer identified several limitations relating to cell-type specificity, the temporal behaviour of ISR activation, and possible PKR cross-reactivity.

      In response to the concern about cell-type specificity, we tested compound 20 in additional cellular models. Compound 20 did not induce detectable ISR signalling in HCT116 or COS-7 cells under the conditions tested, consistent with the reviewer’s observation that activity is not universal across cell types. However, we observed induction of PPP1R15A in primary Mesobank T12 mesothelioma cells. We have therefore revised the manuscript to present compound 20 activity as cell-context dependent rather than broadly generalisable across all cell types.

      To address the reviewer’s concern that the ISR output was examined only at limited time points, we extended the time-course analysis for compounds 20 and 21. Using live-cell ATF4::NanoLuc reporter measurements, both compounds showed maximal reporter activation at approximately 6-8 hours. We also measured translational output by puromycin incorporation and found that compound 20, but not compound 21, caused a progressive GCN2-dependent reduction in translation. These data provide a clearer view of the kinetics and functional consequences of compound 20-mediated ISR activation.

      The reviewer also noted that compound 20 inhibits PKR in vitro at concentrations close to those required for GCN2 activation in cells. We agree that this is an important potential liability. Because PKR signalling was not robustly or reproducibly inducible in our CHO-based reporter system, we were unable to perform a reliable cellular dose-response analysis of PKR engagement in the present study. We have therefore revised the Discussion to acknowledge kinase cross-reactivity, including possible PKR inhibition, as an important limitation and an issue for future development of this chemical series.

      Finally, we have moderated our mechanistic interpretation of compound 20. Although the data support direct engagement of GCN2 and suggest a mechanism distinct from canonical GCN1-dependent activation, we now discuss GCN1-independent activation more cautiously and in the context of prior reports that GCN2iB can display GCN1-independent activity under specific experimental conditions.

      Discussion

      “While the functional relevance of PKR inhibition in our cellular systems is uncertain, these observations highlight the potential for kinase cross-reactivity, which will be important to address in future studies.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):<br /> (1) The description of the chemical screen for Gcn2 activators is not sufficiently clear and detailed. a) Briefly and early on provide the rationales (modes of action) for using histindinol and latruculin A. Explain further the rationale in Figure 2, outlining the purpose for the combined compound + submaximal dose of histindinol.

      The text has been amended.

      Results

      “Histidinol activates GCN2 by inhibiting histidyl‑tRNA synthetase, leading to the accumulation of uncharged tRNAHis. This uncharged tRNA binds to GCN2, relieving its autoinhibition and activating the kinase (35). Latrunculin A sequesters G‑actin, thereby inhibiting PPP1R15A activity (36, 37). Tunicamycin inhibits protein glycosylation in the endoplasmic reticulum (ER), resulting in activation of PERK (38).”

      “This submaximal concentration was used to allow detection of compounds that enhance ISR signalling when GCN2 is partially activated.”

      b) What is unique about the second CHO: ATF4-luc2 reporter line? Why do only 89 out of the original 130 compounds induce the ISR in this line versus the original CHO: ATF4-Nanoluc cell line? This is confusing for the reader about how compounds were triaged for characterization.

      The ATF4‑Nanoluc and ATF4‑luc2 reporter lines differ only in the luciferase used, but this has important practical consequences. The Nanoluc reporter is substantially more sensitive, so it was used for the primary screen to detect even weak ISR activation. The luc2 reporter has lower sensitivity and a narrower dynamic range, making it a more stringent orthogonal assay. As a result, not all hits from the Nanoluc screen (130 compounds) reproduced in the luc2 line; the 89 compounds retained are those that robustly activate the ISR under these more stringent conditions. This step was therefore used to prioritise stronger, more reproducible activators for downstream characterisation.

      Results

      “While primary screening was performed in ATF4‑Nanoluc lines for maximal sensitivity, hits were subsequently re-tested in a second CHO ATF4::luc2 reporter line as a more stringent orthogonal assay to prioritise robust ISR activators. Of the 130 hits identified in the sensitive Nanoluc screen and passing early toxicity assessment, 89 were confirmed in the luc2 assay, consistent with enrichment for higher-amplitude ISR activators under more stringent detection conditions.”

      c) The study uses a second CHO reporter line in the flow scheme (CHO:ATF4-luc2) and then switches back to an ATF4-Nanoluc line to establish GCN2 dependence. What is the rationale for switching back to the original reporter line?

      The luc2 reporter line was used as a more stringent, orthogonal validation step to prioritise robust ISR activators. For subsequent mechanistic studies, including assessment of GCN2 dependence, we returned to the ATF4‑Nanoluc line because its higher sensitivity and simpler single‑reagent assay format are better suited to multi‑point measurements and comparative analyses. In effect, the luc2 reporter was used for triage, whereas the Nanoluc system was retained for mechanistic characterisation and downstream screening.

      Results

      “In subsequent mechanistic studies, the ATF4‑Nanoluc reporter was again used to take advantage of its higher sensitivity and simpler assay format for multi‑condition comparisons.”

      d) The rationale for the first orthogonal screen described in the results section to identify inducers of ER stress is not clearly explained. The compounds were already determined to be dependent on GCN2 prior to this test, and one would have thought that this criterion would have covered ER stress and alternative eIF2 kinase activators.

      We agree with the reviewer that, in principle, establishing GCN2 dependence should reduce the likelihood of capturing compounds acting through alternative eIF2α kinases. However, we performed this orthogonal ER stress screen to address two practical considerations. First, high‑throughput screening is inherently prone to false positives, as it is typically conducted at a single concentration and time point, and compound libraries may contain degraded or chemically inconsistent material. We therefore used a lower‑throughput, more controlled ER stress assay with freshly sourced compounds and additional readouts (e.g. CHOP and XBP1) to improve confidence in the hits. Second, despite prior evidence of GCN2 dependence, ER stress signalling via PERK converges on the same downstream endpoints: eIF2α phosphorylation and ATF4 induction. We therefore wished to explicitly exclude compounds that activate the ISR indirectly via ER stress. In practice, this proved important, as the orthogonal assay did identify compounds that induced ER stress, which we subsequently excluded from the lead set.

      Results

      “Although hits were prioritised for GCN2 dependence, we performed an additional orthogonal screen to exclude compounds that activate the ISR indirectly via ER stress, which converges on the same downstream outputs.”

      (2) A major point of the manuscript is that there is GCN1 independence for the small molecule activation of GCN2, and this has not yet been reported. One report for this GCN1 independence is reference 9 [Carlson … Wek 2023] (Figure 5). In this report, low doses of GCN2iB that can activate GCN2 (although by the present manuscript at much lower levels than the identified new compounds) induce ATF4 expression in cells expressing an E26A mutant of GCN2 that is suggested to negate GCN1 binding and enhancement of GCN2 activity. Halofuginone induction of ATF4 expression was thwarted by the GCN2 E26A mutant.

      We thank the reviewer for highlighting this important point. We agree that Carlson et al. (2023) suggest that, under certain conditions, GCN2iB can activate GCN2 independently of GCN1 using the E26A mutant. In our experiments, however, we did not observe GCN1‑independent activation with GCN2iB under the conditions tested, i.e. similar low doses. One possible explanation is that the GCN1‑independent activity reported by Carlson et al. occurs only within a narrow concentration range; their observations were made at very low compound concentrations, whereas higher concentrations may engage additional regulatory mechanisms. In our study, we used concentrations optimised for robust ISR activation, which may mask such effects. We also note that the E26A mutation (E18A in yeast), originally identified by two‑hybrid analysis, disrupts the GCN2-GCN1 interaction but may not completely eliminate all modes of functional coupling under all conditions. Taken together, these observations raise the possibility that GCN1‑independent activation represents a context‑dependent mechanism that may be unmasked only under specific experimental conditions or by particular classes of compounds.

      We have revised the Discussion to acknowledge this prior report explicitly and to clarify how our findings relate to it.

      Discussion

      “Recent work suggests that the ATP-competitive modulator GCN2iB can activate GCN2 independently of GCN1 under specific conditions using a GCN2 E26A mutant (9). In our hands, we did not observe GCN1‑independent activation with GCN2iB at the concentrations tested. This discrepancy may reflect a narrow concentration window for GCN1‑independent activation or context‑dependent effects of the E26A mutation. These findings raise the possibility that GCN1‑independent activation of GCN2 may occur under specific conditions or with distinct classes of compounds.”

      (3) The authors state that the ISR was exaggerated in Ppp1r15a KO cells. It would be helpful to include statistical analyses to support this statement.

      Thank you for enabling us to be more precise. New text added:

      Results

      “Activation of the ISR by tunicamycin was exaggerated in the Ppp1r15a<sup>-/-</sup> cells owing to their defective dephosphorylation of eIF2a (wild type vs Ppp1r15a<sup>-/-</sup>, p<0.05).”

      (4) The results state that Chop and Ppp1r15a mRNAs were measured following 4 hours of treatment with compound 18, 20, or 21, but Figure 2D shows treatment from 0 to 8 hours? It appears that the compound still induces these mRNAs in GCN2 KO cells, possibly with delayed kinetics. A lengthened time course study would help determine if this is indeed the case.

      We thank the reviewer for this careful observation. To address the reviewer’s point regarding delayed or GCN2‑independent signalling, we have extended our analysis using compounds 20 and 21, which are more tractable experimentally (solubility). Using the ATF4‑Nanoluc reporter, both compounds show peak ISR activation at ~6–8 h in wild-type cells over an extended time course. In parallel, functional readouts of mRNA translation (puromycin incorporation) demonstrate that compound 20, but not 21, induces a progressive, GCN2‑dependent reduction in translation over this period. These clarify the temporal aspects of signalling by these two compounds.

      Results

      “Studies with compound 18 were limited by poor aqueous solubility; therefore, time‑course analyses focused on compounds 20 and 21. To assess ISR activation over an extended period, live‑cell luciferase measurements were performed using CHO cells stably expressing an ATF4::Nanoluc-PEST reporter. Both compounds elicited maximal reporter activation between 6 and 8 h (Figure S1A&B). Compound 20, but not 21, induced a significant GCN2‑dependent reduction in mRNA translation, as measured by puromycin incorporation, with a progressive effect observed up to 7 h (Figure S1C-F).”

      Legend

      “Supplementary Figure S1. Kinetics of responses to compounds 20 and 21

      (A-B) Wild-type CHO cells stably expressing the ATF4::nanoLuc-PEST reporter were treated with Nano-Glo and either (A) 13mM compound 20 or (B) 13mM compound 21. Median bioluminescence (fold change normalised to DMSO control) ± 95% confidence. Representative experiment (n=4 technical repeats). (C-F) Representative immunoblot of lysates from wild-type or Eif2ak4<sup>-/-</sup> CHO cells treated with 10μM compound 20 or 7.5μM 21 for the indicated times. Immediately before harvesting, cells were treated with 10μg/mL puromycin to label newly synthesised polypeptides. “-“ indicates cells not incubated with puromycin. “U” cells were treated with puromycin but without test compound. “CHX” represents the cycloheximide control (100μg/mL). Molecular size in kDa. (E-F) Quantification of puromycinylated proteins normalised to GAPDH. Mean ± SEM. CHO WT (black) and Eif2ak4<sup>-/-</sup> cells (turquoise. N = 4 independent experiments. Two-way ANOVA with Šídák's multiple comparisons test; ***: p ≤ 0.001.”

      (5) In the section describing the differences between cell lines in the ability of compounds to induce the ISR, this is difficult for the reader to interpret, as no controls are included. How does histidinol (or other canonical inducers of the ISR) behave in the three reporter assays (CHO, 293T, and HeLa)?

      As requested, we now provide ATF4::Nanoluc reporter activation (3mM, 20 hours because of this drug’s slow kinetics)

      Results

      “To benchmark ISR activation in these models, each cell type was treated with 3mM histidinol (Figure S2). Reporter activation was most robust in CHO cells, followed by 293T cells, then HeLa cells.”

      Discussion

      “Moreover, histidinol-induced ISR activation showed a clear hierarchy across cell lines, with CHO cells being the most responsive and HeLa cells the least.”

      Legend

      “Supplementary Figure S2. Cell-type differences in response to histidinol

      Fold-change of ATF4::NanoLuc reporter signal in HEK293T, HeLa and CHO cells transiently transfected with reporter and treated for 20 hours with 3mM histidinol. Fold-change calculated relative to vehicle control. Mean ± SEM).”

      We thank the reviewer for this important question. However, we respectfully disagree that a direct correspondence between the concentrations required for target engagement in the BRET assay and for ISR activation in functional assays should necessarily be expected. BRET (including NanoBRET) is a target engagement assay that measures compound binding to the protein in intact cells, typically by competition with a labelled tracer, and thus reports on apparent intracellular affinity and occupancy rather than downstream biological effect (Robers 2019, PMID 30519940). By contrast, ISR activation is a functional readout that reflects amplification through signalling networks, and can be influenced by multiple additional variables including pathway non-linearity, feedback, and kinase regulation. Consequently, it is well established that potencies derived from target engagement assays do not always align with those measured in functional assays. For example, intracellular kinase profiling studies using NanoBRET have demonstrated systematic potency offsets between binding/engagement measurements and downstream cellular activity, arising from factors such as intracellular ATP competition and pathway context (PMID Capener 2026, PMID 41495225). More generally, target engagement assays provide a quantitative measure of binding, whereas functional assays measure biological outcome, and these readouts need not coincide because they capture distinct aspects of a compound’s mechanism of action. Accordingly, we interpret our BRET data as evidence of direct interaction with GCN2 in cells, rather than as a predictor of the concentration required to activate the ISR. The observation that higher concentrations are required in the BRET assay is therefore not unexpected and does not argue against a requirement for kinase-domain engagement in ISR activation. Instead, it reflects the different mechanistic endpoints captured by the two assay formats.

      We will clarify this point explicitly in the revised manuscript.

      Results

      “The concentrations required to detect target engagement in NanoBRET assays did not directly mirror those required for ISR activation, reflecting the distinction between ligand binding and downstream pathway output.”

      (7) In Figure 5E, the authors suggest that compounds 18 and 20 are non-competitive inhibitors of GCN2 since the Vmax increases with increasing ATP concentration. What is the Km for ATP in the absence or presence of compound 18 or 20? It would be helpful to include progress curves as supplementary data to support the Vmax plots in Fig. 5E. Consider providing more specific units (currently arbitrary units) for the y-axis.

      We thank the reviewer for this insightful comment and agree that our original wording overstated the mechanistic interpretation of these data. In particular, the use of the term “non‑competitive” is not well supported by the current analysis and may be misleading, especially given that our data are consistent with binding within or proximal to the ATP-binding pocket. We have therefore revised the text to remove this designation and instead describe the data more conservatively in terms of changes in apparent Vmax, without assigning a specific inhibition mechanism. With respect to kinetic analysis, we agree that full determination of K<sup>m</sub> values and inclusion of progress curves would provide a more rigorous mechanistic interpretation. However, given the primary focus of this manuscript on identifying and characterising small‑molecule activators of GCN2 in cells, we believe that a detailed steady‑state kinetic analysis would be beyond the scope of the current study. We have therefore moderated our conclusions accordingly and now present these data as preliminary kinetic observations rather than definitive evidence of inhibition modality.

      Results

      “Compounds 18 and 20 altered the apparent kinetic parameters of GCN2, including an increase in the observed V<sub>max</sub>; however, these data do not allow assignment of a specific inhibition modality.”

      (8) The full-length GCN2 assay presented in Figure 5G appears to be unresponsive to uncharged tRNA, a known regulator of GCN2. The statement that compound 20 induces eIF2 phosphorylation to a greater extent than tRNAs is true for the in vitro assay, but arguably is because the in vitro assays do not recapitulate the in vivo arrangement.

      We thank the reviewer for this comment. We respectfully disagree that the assay is unresponsive to uncharged tRNA. In our hands, GCN2 does exhibit activation in response to tRNA; however, the magnitude of this effect is modest (~4‑fold) compared to the substantially stronger activation observed with compound 20 (~40‑fold). As a result, the tRNA response can appear compressed when both are plotted on the same scale. We agree with the reviewer that the in vitro assay does not fully recapitulate the in vivo regulatory environment, where factors such as GCN1 and ribosome association are known to potentiate GCN2 activation. This limitation likely explains the relatively weaker response to tRNA under our assay conditions and was a key motivation for incorporating cellular assays in our study. Interestingly, the marked difference in activation magnitude between tRNA and compound 20 in vitro raises the possibility that compound-mediated activation may, at least in part, bypass regulatory features that normally constrain GCN2 activity in a GCN1‑dependent manner. While we have not directly tested this hypothesis, we will temper the wording and include this as a speculative point in the Discussion.

      Results

      “Of note, uncharged tRNA produced a modest (~4‑fold) activation of GCN2 under these conditions, whereas compound 20 induced substantially greater (~40‑fold) activation.”

      Discussion

      “The markedly greater activation observed with compound 20 compared with uncharged tRNA in vitro raises the possibility that such compounds may partially bypass regulatory constraints on GCN2 activation, including those normally mediated by GCN1.”

      (9) Using purified GCN2 kinase domain at low ATP concentrations (10 μM), compounds 18 and 20 were shown not to inhibit GCN2 up to concentrations of 3 μM. In previous assays, much higher concentrations of compounds 18 and 20 were used to inhibit GCN2. Why were different concentrations used? This makes this interpretation of this data difficult for the reader to draw conclusions.

      We thank the reviewer for this comment and agree that the use of different concentration ranges across assays may not have been sufficiently clear. The kinase‑domain assay performed at low ATP (10 μM) was specifically designed to assess whether compounds 18 and 20 have a propensity to inhibit GCN2 under conditions that sensitise detection of ATP‑competitive effects and facilitate comparison with related eIF2α kinases. This assay was therefore optimised for detecting inhibition, rather than activation. In contrast, the higher concentrations used in other experiments were selected to robustly measure ISR activation in cellular or full‑length protein contexts, where higher compound exposure is required to observe downstream signalling outputs. These two assay systems therefore address distinct mechanistic questions—targeting inhibition under controlled biochemical conditions versus activation in more complex functional settings—and are not directly comparable in terms of concentration–response relationships. We will revise the manuscript to clarify this distinction and to emphasise that the kinase‑domain assay was not intended to define the activation potency of the compounds.

      Results

      “This assay was performed at low ATP concentrations to sensitise detection of ATP-competitive inhibition and was not optimised to detect compound-mediated activation of GCN2.”

      (10) In silico docking studies support the binding of compound 20 in the ATP-binding pocket of the GCN2 kinase domain. How is this compatible with the stated ATP non-competitive mechanism?

      We agree with the reviewer that our previous description was misleading. The designation of compounds 18 and 20 as “ATP non‑competitive” is not supported by the available data and is inconsistent with the docking results suggesting binding within the ATP‑binding pocket. We have therefore revised the manuscript to remove this terminology and to describe the kinetic behaviour more cautiously, without assigning a specific mode of inhibition.

      (11) The study does not appear to feature biological assays demonstrating the effects of GCN2 activation. For example, does compound 20 reduce translation or growth of cells in a GCN1/GCN2-dependent manner, or do the compounds overcome PVOD mutations akin to the authors' Hum Mol Genet 2024 Aug 18;33(17):1495-1505 article?

      We thank the reviewer for this suggestion. We agree that defining downstream biological consequences of GCN2 activation is an important goal. We did explore this using several cellular systems; however, these effects were context-dependent and not consistently observed across models. Specifically, compounds did not induce a detectable ISR in HCT116 or COS‑7 cells under the conditions tested, despite responsiveness of these systems to canonical activators such as histidinol. In contrast, we did observe induction of PPP1R15A in primary mesothelioma cells, indicating that biological responses can be elicited in certain cellular contexts. We also tested whether these compounds could rescue disease-associated GCN2 variants linked to PVOD, as previously reported for GCN2iB, but did not observe activation of these mutants. These findings suggest that while the compounds robustly activate GCN2 signalling in reporter assays, downstream biological outputs are context-dependent and may require specific cellular conditions or co-factors. Given the variability across systems, we have limited our conclusions to ISR activation and have not generalised broader biological effects. We will clarify this point in the revised manuscript.

      Results

      “We went on to examine downstream cellular consequences of GCN2 activation in multiple models. While compounds did not induce detectable ISR signalling in HCT116 or COS‑7 cells under the conditions tested, induction of PPP1R15A was observed in Mesobank T12 primary mesothelioma cells, indicating context-dependent biological responses. Moreover, in contrast to GCN2iB (17), the current compounds did not activate disease-associated GCN2 variants linked to PVOD [data not shown].”

      (12) For the control of neratinib and other activators linked with ATP binding that are suggested to be dependent on GCN1 in Fig. 4, include reporter induction by drug treatment in Gcn2-/- cells. It would be helpful to be clear in the list of compounds between those suggested to be direct activators versus those that may create stress that leads to GCN2 activation.

      We thank the reviewer for this suggestion. In response, we have performed additional reporter assays in both WT and GCN2<sup>-/-</sup> cells to assess the dependence of drug-induced ISR activation on GCN2. These experiments reveal clear GCN2 dependence for reporter induction upon treatment with sunitinib, NXP800, WEE1-in-4, Debio0123, gefitinib, and erlotinib, as signal is markedly reduced in GCN2⁻/⁻ cells compared to WT.

      In contrast, dovitinib and AZD1775 show less clear dependence, with relatively low reporter signal even in WT cells (notably lower than observed in Fig. 4C), limiting interpretation. Interestingly, dabrafenib induces stronger reporter activity in GCN2<sup>-/-</sup> cells than in WT, indicating that its effects are independent of GCN2 and may reflect activation of alternative stress or signalling pathways.

      Results

      “To further assess the mechanism of compound-induced ISR activation, we evaluated reporter responses in GCN2-deleted cells (Supplementary Figure S3). Several compounds, including sunitinib, NXP800, WEE1-in-4, Debio0123, gefitinib, and erlotinib, showed reduced reporter activity in GCN2-deficient cells, consistent with GCN2-dependent activation. In contrast, dovitinib, AZD1775 and dabrafenib produced weaker or inconclusive responses even in the paired wild-type lines, limiting analysis. These data allow us to distinguish compounds consistent with direct or GCN2-dependent activation from those more likely to induce ISR indirectly through cellular stress upstream of GCN2. These findings support a distinction between compounds that activate the ISR through GCN2-dependent mechanisms and those that likely act indirectly via alternative stress pathways.”

      Legend

      “Supplementary Figure S3. GCN2-dependence of ISR activation by putative GCN2 agonists

      Normalised fold-change in ATF4 signal in CHO WT (purple) and Gcn2-/- (blue) ATF4::NanoLuc reporter cells treated for 19 hours with a panel of ATP-competitive kinase inhibitors reported to activate GCN2 (at 1 and 3µM, sunitinib used at 3 and 10µM, AZD1175 used at 0.3 and 1µM, gefitinib and erlotinib used at 3 and 10µM). DMSO was used as vehicle control. (n=3; mean ± SEM).”

      Minor Comments:

      (1) In the introduction, the authors state that "Several type 1 and 1.5 kinase inhibitors can activate GCN2 at low concentrations while inhibiting at higher concentrations". What is the evidence that type I inhibitors can activate GCN2?

      Thank you for this query. We believe our initial phrasing was open to misinterpretation. The new wording is

      Introduction

      “Several kinase inhibitors classified as type 1 or type 1.5 with respect to their canonical targets can activate GCN2 at low concentrations while inhibiting it at higher concentrations; however, their binding mode to GCN2 remains undefined.”

      (2) A CHO:ATF4-Nanoluc translation reporter screen was used to screen 123K compounds and divide them into a pool of inhibitors and a pool of activators. The criteria used to make this distinction are not sufficiently described in the manuscript, and screen results are not supplied as supplementary data.

      We thank the reviewer for highlighting the need for greater clarity regarding the screening criteria and reproducibility. Compounds from the primary screen (BioAscent library of 123,222 drug-like compounds performed by the ALBORADA Drug Discovery Institute) were classified based on Z-score thresholds, with activators defined as those with Z-score > 3. To assess robustness, the primary screen data were re-analysed independently. In an initial analysis of 121,000 compounds (excluding plates failing quality control), 6,521 compounds met the activator threshold. Of these, 6,461 overlapped with the original hit list. The small number of discrepancies included: (i) compounds absent from the analysed dataset (e.g. originating from failed plates), and (ii) compounds with Z-scores close to the threshold (typically between −3 and −3.02), consistent with minor analytical variation (e.g. rounding). Overall, these analyses show a high degree of concordance in hit identification, with differences restricted to borderline cases near the selection threshold. We have clarified these criteria below.

      Results

      “Compounds were classified based on Z-score thresholds derived from the primary screen, with activators defined as those with Z-score > 3. An initial analysis of 121,000 compounds, excluding plates failing quality control, identified 6,521 activators, of which 6,461 overlapped with the subset selected for follow-up screening. Minor discrepancies were restricted to compounds absent from the analysed dataset (e.g. originating from failed plates) or those with Z-scores close to the threshold (3 to 3.02), consistent with limited analytical variation. Using this approach, we assembled a subset of 6,461 compounds enriched for potential ISR activators and screened these in 384-well format at 10 µM for 16 hours.”

      Reviewer #3 (Recommendations for the authors):

      (1) The legend to Figure 4 should read "Compound 20 displays GCN1 independence", not "dependence".

      Thank you for spotting this error. We have made the correction.

      (2) The text describes Figure 2D as examining 4h, but it also examines 8h.

      We have amended the text to:

      “After validating their effects using the assays described above, we next confirmed activation of the ISR by these compounds at the transcriptional level at 4 and 8 hours.”

      (3) Unless I'm misreading Table S1, compound Z134826202 is listed as activating the ISR, but the authors describe it in the text as inactive.

      Thank you. We have corrected this error.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the referees for noting the substantive revisions and for the praise of the work. While we each have somewhat different weightings of likelihood, we feel the appraisals are fair and reasonable.


      The following is the authors’ response to the original reviews.

      eLife Assessment

      The authors addressed an important biological question, namely the role of glutamine metabolism in humoral responses, and they obtained solid conclusions. The strength of this study is that the authors used state-of-the-art transgenic mouse models together with in vitro analysis, thereby providing significant insights into the question posed. The following would strengthen the manuscript: i) adding more in-depth functionality/physiological relevance in the discussion part, and ii) regarding the experiments, the inclusion of more appropriate controls and a clearer and more accurate description of the methods.

      We are grateful for the decision of the Editors to select this submission for in-depth peer review and to the Reviewing Editor and referees for the thoughtful and constructive comments.

      We mostly agree with the specific comments and evaluation of strengths of what the work adds as well as with indications of limitations and caveats that apply to the breadth of conclusions. We have edited the text to be more clear and provide more details about certain aspects of the Methods and Legends. In addition, although we try to avoid Discussion sections that are unduly long or have flights of fancy, we will add to the Discussion as well as edit it for directness about potential relevance, basic explorations of mechanisms, and functionality.

      The revised manuscript also contains new data, some of it dealing with comments of the referees, other additions representing work done while the manuscript was under review. While we would be inclined to do more, the sad practical problem is one of limits placed by both the absence of any grant funds and the institution's terminations (RIFs) of the two experimenters in the lab.

      While we believe the original data interpretable as presented originally, up to a point it nonetheless is good to enhance scope or have even better data and add refinements about some of the technical issues. Ultimately, the question becomes "when is enough enough?"

      In the detailed point-by-point response below, we outline changes prompted by the reviewers. We also comment on a few points more expansively that would be suitable for the paper itself, and offer some skepticism or disagreement, (longer and more detailed explanations.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Cho et al. present a comprehensive and multidimensional analysis of glutamine metabolism in the regulation of B cell differentiation and function during immune responses. They further demonstrate how glutamine metabolism interacts with glucose uptake and utilization to modulate key intracellular processes. The manuscript is clearly written, and the experimental approaches are informative and well-executed. The authors provide a detailed mechanistic understanding through the use of both in vivo and in vitro models. The conclusions are well supported by the data, and the findings are novel and impactful. I have only a few, mostly minor, concerns related to data presentation and the rationale for certain experimental choices.

      Detailed Comments:

      (1) In Figure 1b, it is unclear whether total B cells or follicular B cells were used in the assay. Additionally, the in vitro class-switch recombination and plasma cell differentiation experiments were conducted without BCR stimulation, which makes the system appear overly artificial and limits physiological relevance. Although the effects of glutamine concentration on the measured parameters are evident, the results cannot be confidently interpreted as true plasma cell generation or IgG1 class switching under these conditions. The authors should moderate these claims or provide stronger justification for the chosen differentiation strategy. Incorporating a parallel assay with anti-BCR stimulation would improve the rigor and interpretability of these findings.

      We edited the manuscript to be clear that total splenic B cells were used in this set-up figure and the rest of the paper. In addition, we performed new experiments to improve this "set-up figure (Fig. 1)" and moved the older data using alternative experimental conditions to a supplemental figure, Figure 1 - supplement 1. We also used new conditions that included styles of stimulating proliferation and differentiation - to foster an increased sense of generality. The findings in no way change the supported conclusions of the work. Specifically, we used mitogenic stimulation with anti-IgM <sup>+</sup> anti-CD40, all with BAFF, IL-4, and IL-5 in addition to the anti-CD40 stimulation of the original manuscript, bearing in mind excellent work from Aiba et al, Immunity 2006; 24: 259-268, and similar papers. In addition, we added a panel with representative flow cytometric profiles. These new data are presented in Figure 4 - supplement 1 (panels ae).

      To be transparent and add to a more open public discussion (using the virtues of this forum), the senior author and colleagues would caution about whether any in vitro conditions exist that warrant complete confidence. That is the reason for proceeding to immunization experiments in vivo. That is not said to cast doubt on our own in vitro data - there are some experiments (such as those of Fig. 1a-c and associated Fig 1 - supplement 1) that only can be done in vitro or are better done that way (e.g., because of rapid uptake of early apoptotic B cells in vivo).

      For instance: Well-respected papers use the CD40LB and NB21.2D9 systems to activate B cells and generate plasma cells. Those appear to be BCR-independent and yet continue in common use. [We found that these cellular systems (CD40LB; NB21.2D9) cannot be used in experiments with a.a. deprivation or the inhibitors due to effects on the engineered stroma-like cells.] In considering BCR engagement, Reth has published salient points about signaling and concentrations of the Ab, the upshot being that this means of activating mitogenesis and plasma cell differentiation (when the B cells are costimulated via CD40 or TLR (4 or 7/8) is also artificial. Moreover, although Aiba et al, Immunity 2006; 24: 259-268 is a laudable exception, one rarely finds papers using BAFF despite the strong evidence it is an essential part of the equation of B cell regulation in vivo and a cytokine that modulates BCR signaling - in the cultures.

      (2) In Figure 1c, the DMK alone condition is not presented. This hinders readers' ability to properly asses the glutaminolysis dependency of the cells for the measured readouts. Also, CD138<sup>+</sup> in developing PCs goes hand in hand with decreased B220 expression. A representative FACS plot showing the gating strategy for the in vitro PCs should be added as a supplementary figure. Similarly, division number (going all the way to #7) may be tricky to gate and interpret. A representative FACS plot showing the separation of B cells according to their division numbers and a subsequent gating of CD138 or IgG1 in these gates would be ideal for demonstrating the authors' ability to distinguish these populations effectively.

      In the revised manuscript, we have added new experimental data (Figure 1).

      We agree that exact placement of divisions and deconvolution by FlowJow is more fraught than might be thought from presentations in many or most papers. We include the data shown to the right as representative FACS plot(s) with old and new data that illustrate the gating on CTV fluorescence. With the representative examples pasted in here and presented in Fig 1 - supplement 1f, g of the revised manuscript, we will aver that using divisions 0-6, and ≥7 was and is entirely reasonable.

      Ditto for DMK with normal glutamine. However, in the spirit of eLife transparency lacking in many other journals, this comparison is more fraught than the referee comment would make things seem. The concentration tolerated by cells is highly dependent on the medium and glutamine concentration, and perhaps on rates of glutaminolysis (due to its generation of ammonia). In practice, DMK becomes more toxic to B cells unless glutamine is low or glutaminolysis is restricted. Thus, the concentration of DMK that is tolerated and used in Fig. 1b, c can become toxic to the B cells when using the higher levels of glutamine in typical culture media (2 mM or more) - at which point the "normal conditions <sup>+</sup> DMK" "control" involves the surviving cells in conditions with far greater cell death and less population expansion than the "low glutamine <sup>+</sup> DMK". condition.

      (3) A brief explanation should be provided for the exclusive use of IgG1 as the readout in classswitching assays, given that naïve B cells are capable of switching to multiple isotypes. Clarifying why IgG1 was preferentially selected would aid in the interpretation of the results.

      On lines ~112-3 and ~182-5, we edited the text in light of the referee's suggestion that we focus the presentation of serologic data on IgG1 in the immunization experiments. We also rearranged figures and panels to be more explicit and harmonize. That said, and [Brief explanation - IgG1 provides the strongest signal and hence better signal/noise both in vitro and with the alum-based immunizations that are avatars for the adjuvant used in the majority of protein-based vaccines for humans. Perhaps for this reason, the majority of papers on molecular mechanisms seem only to analyze IgG1. Nonetheless, since molecular regulation can differ according to isotype, and the more pro-inflammatory mouse IgG2c is more pertinent to some forms of anti-pathogen immunity and some auto-immune disease models, we believe it valuable to retain these data in supplements to the related Figures.]

      (4) The immunization experiments presented in Figures 1 and 2 are well designed, and the data are comprehensively presented. However, to prevent potential misinterpretation, it should be clarified that the observed differences between NP and OVA immunizations cannot be attributed solely to the chemical nature of the antigens - hapten versus protein. A more significant distinction lies in the route of administration (intraperitoneal vs. intranasal) and the resulting anatomical compartment of the immune response (systemic vs. lung-restricted). This context should be explicitly stated to avoid overinterpretation of the comparative findings.

      We appreciate the positive assessment, and agree with the referee that it is possible the conditions of immune challenge or re-exposure may contribute to the observed differences. We edited the text of the revised manuscript accordingly [lines ~152-153; ~159-160]. Certainly, the difference in how the anti-ova response is elicited compared to the anti-NP response in the same mice or with a bit different an immunization regimen might be another factor - or the major factor - explaining why glutaminolysis was important after ovalbumin inhalations (used because emergence of anti-ova Ab / ASCs is suppressed by the NP hapten after NP-ova immunization) but not needed for the anti-NP response unless Slc2a1 or Mpc2 also was inactivated. Thank you prompting addition of this important caveat!

      Nevertheless, it seems fair to note that in Figures 1 and 2, the ASCs and Ab are being analyzed for NP and ova in the same mice, albeit with the NP-specific components not being driven by the inhalations of ovalbumin. With that in mind, when one compares the IgG1 anti-NP ASC and Ab to those for IgG1 anti-ovalbumin (ASC in bone marrow; Ab), the ovalbumin-specific response was reduced whereas the anti-NP response was not. [lines ~171-172]

      (5) NP immunization is known to be an inducer of an IgG1-dominant Th2-type immune response in mice. IgG2c is not a major player unless a nanoparticle delivery system is used. However, the authors arbitrarily included IgG2c in their assays in Figures 2 and 3. This may be confusing for the readers. The authors should either justify the IgG2c-mediated analyses or remove them from the main figures. (It can be added as supplemental information with proper justification).

      We rearranged the Figure panels to move IgM and IgG2c data to Supplemental Figures (Figure 3 - supplements 1, 2, 4, 5 in the eLife system).

      For purposes of public discourse, we note first that in contrast to the premise about weak IgG2c responses, the data [previously, Figure 3(c, g); now in the supplements] show substantial levels of NP-specific IgG2c. The referee is quite right that the class switching and in vitro ASC generation were done with IL-4 / IgG1-promoting conditions.

      To assist readers, the revised manuscript takes note of the important role of IgG2c (mouse - IgG1 in humans) in controlling or clearing various pathogens as well as in autoimmunity [lines ~182-5]. Moreover, we continue to think that these measurements add substantial value both from the standpoint of providing a better sense of generality to the loss-of-function effects, and in considering potential ways of translating the findings to B cell-dependent autoimmune conditions such as systemic lupus erythematosus.

      [As a scientific aside, we speculate that a greater or lesser IgG2c anti-NP response may arise due to different preparations of NP-carrier obtained from the vendor (Biosearch) having different amounts of TLR (e.g., TLR4) ligand. In any case, the points of presenting the IgG2c (and IgM) data were to push against the limiting boundaries of convention (which risks perpetuating a narrow view of potential outcomes) and make the breadth of results more apparent to readers.

      (6) Similarly, in affinity maturation analyses, including IgM is somewhat uncommon. I do not see any point in showing high affinity (NP2/NP20) IgMs (Figure 3d), since that data probably does not mean much.

      As noted in the reply immediately preceding this one, we appreciate this suggestion from the reviewer and moved the IgM and IgG2c to supplemental status.

      Nonetheless, in collegial discourse we disagree a bit with the referee in light of our data as well as of work that (to our minds) leads one to question why inclusion of affinity maturation of IgM is so uncommon - as the referee accurately notes. Of course a defect in the capacity to class-switch is highly deleterious in patients but that is not the same as concluding that recall IgM or its affinity is of little consequence.

      In some of the pioneering work back in the 1980's, Bothwell showed that NP- carrier immunization generated hybridomas producing IgM Ab with extensive SHM (~11% of the 18 lineages; ~ 1/3 of the IgM hybridomas) [PMID: 8487778], IgM B cells appear to move into GC, and there is at least a reasonable published basis for the view that there are GC-derived IgM (unswitched) memory B cells (MBC) that would be more likely, upon recall activation, to differentiate into ASCs. [As an example, albeit with the Jenkins lab anti-rPE response, Taylor, Pape, and Jenkins generated quantitative estimates of the numbers of Ag-specific IgM<sup>+</sup> vs switched MBC that were GC-derived (or not). [PMID: 22370719]. While they emphasized that ~90% of IgM<sup>+</sup> MBC appeared to be GC-independent, their data also indicated that ~1/2 of all GC-derived MBC were IgM<sup>+</sup> rather than switched (their Fig. 8, B vs C; also 8E, which includes alum-PE). And while we immensely respect the referee, we are perhaps less confident that IgM or high-affinity Ag-specific IgM doesn't mean that much, if only because of evidence that localized Ab compete for Ag and may thus influence selective processes [PMCID: PMC2747358; PMID: 15953185; PMID: 23420879; PMID: 27270306].

      (7) Following on my comment for the PC generation in Figure 1 (see above), in Figure 4, a strategy that relies solely on CD40L stimulation is performed. This is highly artificial for the PC generation and needs to be justified, or more physiologically relevant PC generation strategies involving anti-BCR, CD40L, and various cytokines should be shown.

      In line with our response to point (1), we tested BCR-stimulated B cells (anti-CD40 plus anti-IgM with BAFF, IL-4, and IL-5, parallel to the analyses with anti-CD40 but no BCR engagement). These results align with and reinforce the utility of the data with anti-CD40 as the sole mitogen.

      (8) The effects of CB839 and UK5099 on cell viability are not shown. Including viability data under these treatment conditions would be a valuable addition to the supplementary materials, as it would help readers more accurately interpret the functional outcomes observed in the study.

      We added presentation of data that provide cues as to relative viability / cxmsurvival under the experimental conditions used.

      [FSC X SSC as well as 7AAD or Ghost dye panels; we also generated new data that in[ further experiments scoring annexin V staining (see Fig 4 - supplement 1d, e, and Fig 5 - supplement 1e, f)].

      (9) It is not clear how the RNA seq analysis in Figure 4h was generated. The experimental strategy and the setup need to be better explained.

      Including text added at lines ~291-293 and ~582-585, the revised manuscript provides more information in the Results, Methods and Legend for Fig 4j-l. We agree entirely with the concern and apologize that in this and a few other instances we inadvertently sacrificed sufficiency of detail on the altar of attempting brevity.

      [As a synopsis: In three temporally and biologically independent experiments, cultures were harvested 3.5 days after splenic B cells were purified and cultured as in the experiments of Fig. 4a-e. Total cellular RNA was prepared from the twelve samples (three replicates for each of four conditions - DMSO vehicle control, CB839, UK5099, and CB839 <sup>+</sup> UK5099), then analyzed by RNA-seq. RNA-seq data were initially processed using the pipeline described in the Methods. For panels g & h of Fig 4, DESeq2 was used to quantify and compare read counts in the three CB839 <sup>+</sup> UK5099 samples relative to the three independent vehicle controls and identify all genes for which variances yielded P<0.05. In Fig 4g, all such genes for which the difference was 'statistically significant' (i.e., P<0.05) were entered into the indicated Immgen tool and thereby mapped to the B lineage subsets shown in the figure panels (i.e., g, h). In (g), these are displayed using one format, whereas (h) uses the 'heatmap' tool in MyGeneSet.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, the authors investigate the functional requirements for glutamine and glutaminolysis in antibody responses. The authors first demonstrate that the concentrations of glutamine in lymph nodes are substantially lower than in plasma, and that at these levels, glutamine is limiting for plasma cell differentiation in vitro. The authors go on to use genetic mouse models in which B cells are deficient in glutaminase 1 (Gls), the glucose transporter Slc2a1, and/or mitochondrial pyruvate carrier 2 (Mpc2) to test the importance of these pathways in vivo.

      Interestingly, deficiency of Gls alone showed clear antibody defects when ovalbumin was used as the immunogen, but not the hapten NP. For the latter response, defects in antibody titers and affinity were observed only when both Gls and either Mpc2 or Slc2a1 were deleted. These latter findings form the basis of the synthetic auxotrophy conclusion. The authors go on to test these conclusions further using in vitro differentiations, Seahorse assays, pharmacological inhibitors, and targeted quantification of specific metabolites and amino acids. Finally, the authors document reduced STAT3 and STAT1 phosphorylation in response to IL-21 and interferon (both type 1 and 2), respectively, when both glutaminolysis and mitochondrial pyruvate metabolism are prevented.

      Strengths:

      (1) The main strength of the manuscript is the overall breadth of experiments performed. Orthogonal experiments are performed using genetic models, pharmacological inhibitors, in vitro assays, and in vivo experiments to support the claims. Multiple antigens are used as test immunogens--this is particularly important given the differing results.

      (2) B cell metabolism is an area of interest but understudied relative to other cell types in the immune system.

      (3) The importance of metabolic flexibility and caution when interpreting negative results is made clear from this study.

      Weaknesses:

      (1) All of the in vivo studies were done in the context of boosters at 3 weeks and recall responses 1 week later. This makes specific results difficult to interpret. Primary responses, including germinal centers, are still ongoing at 3 weeks after the initial immunization. Thus, untangling what proportion of the defects are due to problems in the primary vs. memory response is difficult.

      We performed new experiments and added the data on differences prior to a boost [see below; new Fig 3d, e; etc].

      (2) Along these lines, the defects shown in Figure 3h-i may not be due to the authors' interpretation that Gls and Mpc2 are required for efficient plasma cell differentiation from memory B cells. This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. The more likely interpretation is that ongoing primary germinal centers are negatively impacted by Gls and Mpc2 deficiency, and this, in turn, leads to reduced affinities of serum antibodies.

      We have edited the wording of the conclusion to add a possibility we consider unlikely and downplay a conclusion that MBCs bearing switched BCRs are affected once reactivated. [see lines ~221-230] We also have added citations pertaining to the topic, including work from the Victora lab which seems to put the point succinctly: "Recall GCs in mice consist almost entirely of naïve B cells, whereas recall antibodies derive overwhelmingly from memory B cells." [emphasis added] [PMID: 38838672; new ref #83]. While unclear as to the reasoning - as one looks at the data - and skeptical as to the accuracy of the referee's point (2), it suggests that the matter is open to reasonable doubt. In line with the point and the edits, we also have added citation of a bioRxiv preprint from the Victora lab, which touches on the concept of what one could call boost-induced reinvigoration of a pre-existing GC [new ref #82].

      Beyond the textual changes, we performed a new series of experiments to investigate partially, and present the results in Fig 3d, e as well as Fig 3 - supplement 1d, e. Unfortunately, time before lab closure was an enemy both for the period between primary and recall immunizations in performance and multiple replication of work to extend that presented in Figure 3, panels g & h, and the related Supplemental Data (Fig 3 - supplements 4d, 5a-g). Unfortunately, it was not possible to do a longer-term memory experiment with recall immunization out at 8 weeks.

      The intriguing concerns and questions of points 1 & 2 provide a springboard for consideration of generalizations and simplifications. Germinal center durability is not at all monolithic, and instead is quite variable**. It is true that in the literature (especially with the substantially different approach of transferring BCR-transgenic / knock-in versions of an NP-biased BCR) there may be meaningful pools of IgG1 and IgG2c GC B cells. The premise (cognitive bias, perhaps?) in our interpretation is that in our previous work we measured few if any GC B cells - NP-APC-binding or otherwise - above the background (non-immunized controls) three weeks after immunization with NP-ovalbumin in alum. While recognizing that the immunogen can matter, we note for the readers and referee that Fig. 1 of the Taylor, Pape, & Jenkins paper considered above [PMID: 22370719] reported 10-fold more Ag-specific MBCs than GC B cells at day 29 post-immunization (the point at which the boost/recall challenge was performed in our Figure 3g, h. [That work did not use NP-carrier in alum to immunize, or measure the anti-NP response.]

      Viewing Fig. 3i from that perspective, the surmise of the comment is that a major contribution to the differences in both all-affinity and high-affinity anti-NP IgG1 (whose production requires differentiation into plasma cells) derived from the immunization at 4 wk stimulating persistent GC B cells as opposed to memory B cells.

      The issue and question also relate to rates of output of plasma cells or rises in the serum concentrations of class-switched Ab. To this point, our prior experiences agree with the long-published data of the Kurosaki lab in Figure 3c of the Aiba et al paper noted above (Immunity, 2006) (and other such time courses). Readers can note that the IgG1 anti-NP response (alum adjuvant, as in our work) hits its plateau at 2 wk, and did not increase further from 2 to 3 wk. The most likely interpretation is that GC are on the decline and Ab production has reached its plateau by the time of the 2nd immunization in Fig. 3h.

      Assuming we understand the comment and line of reasoning correctly, we also lean towards disagreeing with the statement " This interpretation would only be correct if the absence of Gls/Mpc2 leads to preferential recruitment of low-affinity memory B cells into secondary plasma cells. Our evidence shows that both low-affinity as well as high-affinity anti-NP Ab (IgG1) were reduced due to combined gene-inactivation after the peak primary response (Fig. 3h; also, see the new data in Fig 3 and Fig 3 - supplement 1). Recent papers show that affinity maturation is attributable to greater proliferation of plasmablasts with high-affinity BCR. Accordingly, the findings with loss of GLS and MPC function are quite consistent with the interpretation that much of the response after the second immunization draws on MBC differentiation into plasmablasts and then plasma cells, where the proliferative advantage of high-affinity cells is blunted by the impaired metabolism. Notwithstanding these issues, the revised manuscript includes the alternative, if less likely, interpretation proposed by the review [lines ~221-230].

      **In some contexts, of course, especially certain viral infections or vaccination with lipid nanoparticles carrying modified mRNA, germinal centres are far more persistent; also, in humans even the seasonal flu vaccine

      (3) The gating strategies for germinal centers and memory B cells in Supplemental Figure 2 are problematic, especially given that these data are used to claim only modest and/or statistically insignificant differences in these populations when Gls and Mpc2 are ablated. Neither strategy shows distinct flow cytometric populations, and it does not seem that the quantification focuses on antigen-specific cells.

      The revised manuscript improves these aspects of the presentation, using old and new data. See Fig 3 - supplement 3a, c; Fig 3 - supplement 4a. We note for readers that many other papers in the best journals show plots in which the separation of, say, GC-Tfh from overall Tfh is based on cut-off within what essentially is a continuous spectrum of emission as adjusted or compensated by the cytometer (spectral or conventional).

      The revised manuscript presents results from new experiments that deal with the subset of GC B cells whose BCRs bind NP-APC with enough affinity to retain a positive signal after washing. These new data are presented in Fig 3 - supplement 3c & 3e. In practice, the new findings suggest that the metabolic requirement applied more to the NP-binding B cells than the overall GC B cell population.

      (4) Along these lines, the conclusions in Figure 6a-d may need to be tempered if the analysis was done on polyclonal, rather than antigen-specific cells. Alum induces a heavily type 2-biased response and is not known to induce much of an interferon signature. The authors' observations might be explained by the inclusion of other ongoing GCs unrelated to the immunization.

      We apologize for ambiguity or insufficient clarity and, as noted above, have edited the text to be more clear that the in vitro experiments do not represent GC B cells and that the RNA-seq data were from experiments that did not involve alum and were not an Ag (SRBC)-specific subset.

      New text in the Results, an expanded Legend, and tweaking the Methods make it more readily clear that the RNA-seq data (and hence the GSEA) involved immunizations with SRBC (not the alum / NP system. That said, we note that the hapten-carrier experiments in which the immunogen was adjuvantized with alum actually generated a robust IgG2c (type 1-driven) response along with the type 2-enhanced IgG1 response, in line with what has been reported by others with alum-adjuvanted vaccination.

      Reviewer #3 (Public review):

      Summary:

      In their manuscript, the authors investigate how glutaminolysis (GLS) and mitochondrial pyruvate import (MPC2) jointly shape B cell fate and the humoral immune response. Using inducible knockout systems and metabolic inhibitors, they uncover a "synthetic auxotrophy": When GLS activity/glutaminolysis is lost together with either GLUT1-mediated glucose uptake or MPC2, B cells fail to upregulate mitochondrial respiration, IL 21/STAT3 and IFN/STAT1 signaling is impaired, and the plasma cell output and antigen-specific antibody titers drop significantly. This work thus demonstrates the promotion of plasma cell differentiation and cytokine signaling through parallel activation of two metabolic pathways. The dataset is technically comprehensive and conceptually novel, but some aspects leave the in vivo and translational significance uncertain.

      Strengths:

      (1) Conceptual novelty: the study goes beyond single-enzyme deletions to reveal conditional metabolic vulnerabilities and fate-deciding mechanisms in B cells.

      (2) Mechanistic depth: the study uncovers a novel "metabolic bottleneck" that impairs mitochondrial respiration and elevates ROS, and directly ties these changes to cytokinereceptor signaling. This is both mechanistically compelling and potentially clinically relevant.

      (3) Breadth of models and methods: inducible genetics, pharmacology, metabolomics, seahorse assay, ELISpot/ELISA, RNA-seq, two immunization models.

      (4) Potential clinical angle: the synergy of CB839 with UK5099 and/or hydroxychloroquine hints at a druggable pathway targeting autoantibody-driven diseases.

      We agree and thank the referee for the positive comments and this succinct summary of what we view as contributions of the paper.

      Weaknesses:

      (1) Physiological relevance of "synthetic auxotrophy"

      The manuscript demonstrates that GLS loss is only crippling when glucose influx or mitochondrial pyruvate import is concurrently reduced, which the authors name "synthetic auxotrophy". I think it would help readers to clarify the terminology more and add a concise definition of "synthetic auxotrophy" versus "synthetic lethality" early in the manuscript and justify its relevance for B cells.

      We edited the Abstract, Introduction, and Discussion to try to do better on this score. Conscious of how expansive the prose and data are even in the original submission, we appear to have taken some shortcuts that we will try to rectify or at least mitigate. Thank you for highlighting this need to improve on key concepts !!

      Specifically, the revised text expands a bit on the notion that synthetic auxotrophy represents effects on differentiation that go beyond additional mechanisms of reducing division efficiency and a modest impact on selective death. [see the 10th - 11th lines in Abstract and lines ~84-85, Introduction] Even though decreased population expansion is observed and new evidence supports a model in which the altered metabolism contributes to enhanced death in vivo, at equal division numbers the frequency of CD138<sup>+</sup> progeny is lower once glutaminolysis and mitochondrial pyruvate are reduced by either genetic or pharmacological means.

      This comment of the review raises interesting semantic questions about what represents "physiological relevance". The fundamental point is to explore a basic science question - what, if any, are limits to metabolic flexibility? In principle, shouldn't B cells be able to use fatty acid metabolism to generate enough ATP and provide the backbones for biosynthesis during growth? Put a different way, the point is that a basic curiosity to understand why decreasing glucose influx did not have an even more profound effect than what was observed, combined with curiosity as to why glutaminolysis was dispensable in relatively standard vaccine-like models of immunize/boost, provided a springboard to identification of new vulnerabilities. The manuscript shows one physiological limitation (and hence vulnerability). Be that as it may, the revised text of the Discussion section more clearly addresses this issue (lines ~531-549 at the end of the Discussion).

      While the overall findings, especially the subset specificity and the clinical implications, are generally interesting, the "synthetic auxotrophy" condition feels a little engineered.

      CAR-T cells are 'a little engineered' (or more than a little) and yet they do seem to have had an impact on understanding the centrality of B cells in various autoimmune conditions as well as in the direction of cancer therapy research. So it is a matter of balancing this perspective of the referee against the strengths they highlight in points 1, 2, and 4. In editing the revision, we try to expand and be more explicit about this in the Discussion of the revised manuscript.

      In brief, even were the money not all gone, we would not believe that expanding the heft of this already rather large manuscript and set of data would be appropriate. As matters stand, a basic new insight about metabolic flexibility and its limits leads to evidence of a way to reduce generation of Ab and a novel impairment of STAT transcription factor induction by several cytokine receptors. The vulnerability that could be tested in later work on B cell-dependent autoimmunity includes the capacity to test a compound that already has been to or through FDA phase II in patients together with an FDA-approved standard-of-care agent.

      Therefore, the findings strongly raise the question of the likelihood of such a "double hit" in vivo and whether there are conditions, disease states, or drug regimens that would realistically generate such a "bottleneck".

      Hence, the authors should document or at least discuss whether GC or inflamed niches naturally show simultaneous downregulation/lack of glutamine and/or pyruvate. The authors should also aim to provide evidence that infections (e.g., influenza), hypoxia, treatments (e.g., rapamycin), or inflammatory diseases like lupus co-limit these pathways.

      Again, we appreciate some 'licensing' to be more expansive and explicit, and will try to balance editing in such points against undue tedium or tendentiously speculative length in the Discussion. In particular, we will note that a clear, simple implication of the work is to highlight an imperative to test CB839 in lupus patients already on hydroxychloroquine as standard-of-care, and to suggest development of UK5099 (already tested many times in mouse models of cancer) to complement glutaminase inhibition.

      As backdrop, we note that the failure to advance imaging mass spectrometry to the capacity to quantify relative or absolute (via nano-DESI) concentrations of nutrients in localized interstitia is a critical gap in the entire field. Techniques that sample the interstitial fluid of tumour masses or in our case LN as a work-around have yielded evidence that there can be meaningful limitations of glucose and glutamine, but it needs to be acknowledged that such findings may be very model-specific and, as can be the case with cutting-edge science, are not without controversy. That said, yes, we had found that hypoxia reduced glutamine uptake but given the norms of focused, tidy packages only reported on leucine in an earlier paper [PMID27501247; PMCID5161594].

      Beyond all that, another impetus to and inspiration for these experiments stems from quite data that we generated in a model of short-term protein-restricted diet (loosely akin to kwashiorkor in humans), based on an excellent publication showing that such a regimen quickly led to lower circulating glutamine and mTORC1 activity (**). In brief, we found that a low-protein diet did, in our experiments, preferentially lower glutamine but - importantly - led to reduced Ab responses (which would match what we have modeled here). The findings were not a well-enough connected evidentiary component to include in the "story" but I'll append slides with the relevant data to this Response to Reviews for the referee's perusal (and anyone else who reads this online discourse).

      It would hence also be beneficial to test the CB839 + UK5099/HCQ combinations in a short, proof-of-concept treatment in vivo, e.g., shortly before and after the booster immunization or in an autoimmune model. Likewise, it may also be insightful to discuss potential effects of existing treatments (especially CB839, HCQ) on human memory B cell or PC pools.

      We certainly agree that the suggestions offered in this comment are important next steps and the right approach to test if the findings reported here translate toward the treatment of autoimmune diseases that involve B cells, interferons, and pathophysiology mediated by auto-Ab. As practical points, performance and replication of such studies would take more time than the year allotted for return of a revised manuscript to eLife and in any case neither funds nor a lab remain to do these important studies.

      Concrete evidence for our concurrence was embodied in a grant application to NIH that was essential for keeping a lab and doing any such studies. [We note, as a suggestion to others, that an essential component of such studies would be to test the effects of these compounds on B cells from patients and mice with autoimmunity]. Perhaps unfortunately for SLE patients, the review panelists did not agree about the importance of such studies. However, it can be hoped that the patent-holder of CB839 (and perhaps other companies developing glutaminase inhibitors) will see this peer-reviewed preprint and the public dialogue, and recognize how positive results might open a valuable contribution to mitigation of diseases such as SLE.

      (2) Cell survival versus differentiation phenotype

      Claims that the phenotypes (e.g., reduced PC numbers) are "independent of death" and are not merely the result of artificial cell stress would benefit from Annexin-V/active-caspase 3 analyses of GC B cells and plasmablasts. Please also show viability curves for inhibitor-treated cells.

      This comment leads us to see that the wording on this point may have been overly terse in the interests of brevity, and thereby open to some odd misunderstanding. The CD138<sup>+</sup> events are scored among VIABLE CELLS, so a decrease in the %CD138<sup>+</sup> at similar division number represents an effect independent from (or beyond) survival and division-counting. Accordingly, we expanded the text of the Abstract and elsewhere in the manuscript, to be more clear. In addition, we added data from new experiments addressing death in vitro and among GC-phenotype B cells in vivo. To clarify in this public context, it is not that an increase in death (along with the reported decrease in cell cycling) can be or is excluded. The point is that beyond any such increase, and taking into account division number (since there is evidence that PC differentiation and output numbers involve a 'division-counting' mechanism), the frequencies of CD138<sup>+</sup> cells and of ASCs among the viable cells are lower, as is the level of Prdm1-encoded mRNA even before the big increase in CD138<sup>+</sup> cells in the population.

      (3) Subset specificity of the metabolic phenotype

      Could the metabolic differences, mitochondrial ROS, and membrane-potential changes shown for activated pan-B cells (Figure 5) also be demonstrated ex vivo for KO mouse-derived GC B cells and plasma cells? This would also be insightful to investigate following NP-immunization (e.g., NP+ GC B cells 10 days after NP-OVA immunization).

      We performed a series of new experiments to have enough biologically independent replications for meaningful and statistical analyses. The new results, added in as Fig 5 - supplement 1, showed that the combined pathway interruption by loss-of-function increased ROS, mtROS, and death (annexin V / 7AAD) upon analyzing GCphenotype B cells immediately upon harvest. The findings align well with the data in Fig 5 (cultured B cells).

      (4) Memory B cell gating strategy

      I am not fully convinced that the memory-B-cell gate in Supplementary Figure 2d is appropriate. The legend implies the population is defined simply as CD19+GL7-CD38+ (or CD19+CD38++?), with no further restriction to NP-binding cells. Such a gate could also capture naïve or recently activated B cells. From the descriptions in the figure and the figure legend, it is hard to verify that the events plotted truly represent memory B cells. Please clarify the full gating hierarchy and, ideally, restrict the MBC gate to NP+CD19+GL7-CD38+ B cells (or add additional markers such as CD80 and CD273). Generally, the manuscript would benefit from a more transparent presentation of gating strategies.

      In considering the referee's viewpoint, we further expanded the supplemental data displays to include more of the gating and analytic schemes, which we believe should mitigate one concern noted here. In addition, we now include flow data from the non-immunized control mice that had been analyzed concurrently in the experiments.

      Third and finally, we performed new experiments and analyses in which the focus was the frequencies of memory-phenotype (IgD<sup>neg</sup> GL7<sup>neg</sup> CD38<sup>+</sup> / CD38<sup>hi</sup> aka CD38<sup>+</sup><sup>+</sup>) NPbinding B cells after immunization. While this time, as opposed to previously, the NP-APC staining met our standard for interpretability, the gist of the findings was that the two independent repeat experiments yielded a split decision and a degree of variability. With time being up due to the funds running out, we have elected to delete the issue and the data panel in question.

      That said, it bears noting that in the previous figure panel, the labeling indicated that the gating included the important criterion that cells be IgD<sup>neg</sup>, which excludes the vast majority of naive B cells but measures memory-phenotype B cells independent from consideration of whether or not they were NP-binding.

      [In principle marginal zone (MZ) B cells might fall within this gate. However, the MZ B population is unlikely to explain the differences shown.

      (5) Deletion efficiency - [The] mRNA data show residual GLS/MPC2 transcripts (Supplementary Figure 8). Please quantify deletion efficiency in GC B cells and plasmablasts.

      Even were there resources to do this, the degree of reduction in target mRNA (Gls; Mpc2) renders this question superfluous. To the best of our understanding, the proteins (for which there might be some phenotypic lag) are translated from RNA. Might there be a small subpopulation of B cells (or their PC progeny) with only one, or even neither, allele converted from fl to D? Yes, but they would be a minor subset in light of the magnitude of mRNA reduction, in contrast to our published observations with Slc2a1. As to plasmablasts and plasma cells, the pre-existing populations make such an analysis misleading, while the scarcity of such cells recoverable with antigen capture techniques is so low as to make both RNA and genomic DNA analyses questionable. We also refer readers to the supplemental figure that presents the results of experiments testing the issue one might infer from the question about extents of deletion in PC (i.e., how much counter-selection might have occurred by the PC stage).

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This interesting study adapts machine learning tools to analyze movements of a chromatin locus in living cells in response to serum starvation. The machine learning approach developed is useful, the experiments are well controlled, and the data are solid. The study would be greatly strengthened by testing key predictions made using perturbation experiments. This work will be of interest to those studying chromosome biology and gene expression patterns.

      We thank eLife for this nice assessment. We indeed believe that the presented machine learning approach will be useful for many types of research questions, and this was the main aim of this manuscript.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Redchuk et al. explore the dynamic properties of chromatin upon serum starvation using machine learning approaches. They use CRISPR-tagging to visualize a region on chromosome 1 in human cells and show that in their system, chromosome 1, but not the previously reported chromosomes 10, 13, and X, undergo a change in radial position upon serum starvation. Live cell imaging showed a position change towards the periphery after serum starvation. They then apply a machine learning algorithm for the analysis of the imaging data, which reveals changes in nuclear area during serum starvation and longer displacements of the chromosome 1 locus near the nuclear periphery. Differential behavior of homologues is also reported.

      Strengths:

      (1) The study of chromatin dynamics is an interesting and important area of research.

      (2) The use of machine learning approaches to analyze live cell imaging data is timely.

      (3) With serum starvation, the authors use a simple, well-controllable model system.

      Weaknesses:

      (1) This study only provides limited new insight into chromatin dynamics.

      We respectfully disagree with this conclusion. To the best of our knowledge, our study is the first to provide any insights into chromatin dynamics upon serum starvation. Previous studies are solely based on studies in fixed cells, and the dynamics have remained unexplored. Moreover, for example the notion that homologous chromosomes show differential dynamic behavior is novel and will likely have implications and relevance to many chromatin-based processes beyond the example studied here.

      (2) It was not immediately evident what the use of machine learning approaches added to this study. It appears that the main conclusions could have been reached by conventional analysis.

      First, we would like to point out that the other reviewer found our machine learning analysis pipeline a major strength of our manuscript. Indeed, analyzing single features and assessing their impact on the studied phenomenon could have been achieved relatively easily by conventional analysis. However, this analysis would have ignored the interactions (some of which were not intuitively obvious) between different features and thereby limited the knowledge gain from the experiment.

      Unbiased analysis of the interactions between the different features would have been already very difficult and time-consuming with conventional approaches. We believe that our analysis pipeline, especially with the Shapley values, addresses the key issue of combinatorial explosion prominent to multiparametric data, such as imaging data, and helps the researcher to navigate complex datasets.

      (3) There are several specific technical points:

      (a) It was not clear what the CRISRP-Sirius probes actually labelled. The chromosome 1 sgRNA sequence is provided, but I could not find information as to which region(s) of the chromosome are actually labelled (size, location, etc.).

      We have added a schematic as Supplementary Figure 1A to show the region of the chromosome that is labelled. In addition, the target sequence, together with the relevant references can be found in the Materials and methods (page 16). Please see also below Reviewer #1 (Recommendations for the authors) point 4a.

      (b) The authors visualize a relatively small region of chromosome 1 but make conclusions regarding the entire chromosome. Additional probes on the same chromosome should be used.

      Related to this point, the discussion of why the authors are unable to reproduce the prior findings of relocation of chromosomes 10, 13, and X is not satisfying. It would be worth comparing the FISH-based painting of entire chromosomes, which generated the results suggesting relocation of these chromosomes, with the point-labelling method used here.

      We agree that our approach to labeling chromosome 1 is very different than the FISH-based probes utilized before. However, we also feel that we discuss this aspect, and the difference between our and previous results, which may also stem from the used cell model, in quite a detail in the first paragraph of the results (page 4). Also, we are very careful throughout the manuscript to indicate that here we study the dynamics of a specific chromosome loci, not the entire chromosome, and have further amended the text to emphasize this. In the future, it would be very interesting to study the dynamics of also other loci of chromosome 1. As indicated also below in response to reviewer 2, we have failed to identify further gRNAs that would reliably and reproducibly label further chromosome 1 loci, suggesting that we would need to change the labeling system entirely. Unfortunately, this is not in the scope of this manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 1.

      (c) The study lacks controls. Since in their hands chromosomes 10, 13, and X do not change position, they should be used as a negative control in all experiments demonstrating a shift in the location of chromosome 1.

      We disagree that our study lacks controls, since we use telomeres as controls throughout the manuscript. Please see also below Reviewer #1 (Recommendations for the authors) point 2,3.

      (d) I did not find information about the spatial or temporal resolution of the imaging modality. This is important to assess whether the observed changes in position, relative to time, are meaningful.

      To estimate the spatial resolution, we have added new data using fixed cells (Supplementary figure 1E; corresponding text in results on page 5); temporal resolution is indicated in Materials and methods (page 17). Please see also below Reviewer #1 (Recommendations for the authors) point 4d.

      (e) The authors analyze surprisingly early timepoints (up to 40 minutes) of serum starvation. Would these results look different if longer serum starvation timepoints of several hours were analyzed?

      We chose to analyze early time points of serum starvation based on the previous literature reporting the chromosome relocation within the first 15 minutes of starvation. Indeed, the results might look very different later during serum starvation, since we already observe differences between 0-20 min vs 20-40 min into starvation (see for example Figure 5A-D). Analyzing further time points is not in the scope of this manuscript.

      (f) The authors can do a better job of explaining what the biological meaning of the various parameters (DistR, TDist, etc.) they measure is.

      We have amended Table 1 to describe the measured features more clearly. Please see also below Reviewer #1 (Recommendations for the authors) point 4e.

      (g) I did not understand the reasoning for the authors' conclusion of differential behavior of homologues. Please explain this better, or idealy use more direct labeling methods that identify the individual homologues.

      The differential behavior of homologues is best demonstrated in Figure 6H, which shows that in serum-containing media, the peripheral homolog has equal probability of being faster or slower compared to its homolog. However, the distribution changes upon starvation, with the peripheral loci being more frequently the faster homolog. We completely agree that further studies are needed to understand this phenomenon better, but changing the labeling method is not in the scope of this manuscript.

      (h) In many figures, statistical analysis of the data is missing, including, but not limited to, Figures 1B, C, G, Figures 4, 5, 6.

      We have added a Supplementary table to include inferential statistics. See also below Reviewer #1 (Recommendations for the authors) point 4b.

      (i) No information is provided throughout the manuscript as to how many cells were analyzed in each experiment. This should be indicated in every figure legend.

      The number of analyzed loci or nucleus is indicated in every figure. See also below Reviewer #1 (Recommendations for the authors) point 4c.

      Reviewer #2 (Public review):

      Summary:

      The study demonstrates that CRISPR-Sirius provides a powerful approach to investigating chromosome dynamics in living cells during environmental stress. By focusing on serum starvation, the authors show that this process induces global nuclear changes, including a reduction in nuclear area and increased morphological dynamism, while at the same time driving specific reorganization of chromosome 1. Chromosome 1 relocates toward the nuclear periphery and displays distinctive patterns of motion, maintaining overall motility but punctuated by occasional long-distance displacements, particularly near the nuclear envelope. Importantly, the analysis reveals that homologous copies of chromosome 1 do not behave uniformly: peripheral loci become more mobile and responsive to starvation, whereas central homologs remain comparatively stable, often associated with nucleolar subcompartments. By integrating live imaging with machine learning and explainable AI analysis, the study highlights the complexity of nuclear organization and provides valuable insights into how chromosome-specific and locus-specific responses to stress are orchestrated within the three-dimensional nuclear landscape.

      Strengths:

      The study uses live-cell imaging to investigate the dynamics of loci during starvation. Livecell tracking and data interpretation are carried out using machine learning and AI models, which is a major strength.

      Weaknesses:

      The manuscript is at times difficult to follow, partly because the methodological descriptions are highly specialized, especially for non-expert biologists. In addition, the observations are not tested for a mechanistic basis. Experiments that could provide deeper insights are missing, for example, why chromosome 1 moves, why the peripheral homologue dislocates, or why a "long jump" is observed at the periphery even though the speed of the loci does not change. It is also unclear whether a displacement of 0.5 μm is functionally meaningful.

      We appreciate the comment about the readability of our manuscript, and have seriously evaluated this point. We also completely agree that it would be interesting and important to understand the mechanistic and functional basis of the observed changes in chromatin dynamics take place upon serum starvation. However, we feel that it is not in the scope of the present manuscript. See also below Reviewer #2 (Recommendations for the authors) points 3,7-11.

      Recommendations for the authors:

      Reviewing Editor Comments:

      I would like to first offer my congratulations on a very interesting study; second, I would like to encourage you to test a few key predictions using a perturbation experiment. Two reviewers with deep expertise in this area were supportive of the work, and both noted that such an addition would greatly increase the impact and visibility of this work in the field. I welcome a revision that addresses this seminal point. Thank you for sending your work to eLife!

      We thank eLife for the positive assessment. We have aimed to address all of the reviewers comments and suggestions. However, we feel that some of the suggestions are not in the scope of this particular manuscript, since they would require setting up a different chromatin labeling system.

      Reviewer #1 (Recommendations for the authors):

      The following experiments would strengthen the study:

      (1) Please label additional regions on chromosome 1 so as not to rely on a single point to represent the behavior of the entire chromosome.

      This is an excellent suggestion, but unfortunately, despite our extensive efforts, we have failed to identify further gRNAs that would reliably label chromosome loci with the CRISPR-Sirius system. Changing the labeling system is not in the scope of the presented manuscript.

      (2) Please use chromosomes 10, 13, or X as a negative control since these chromosomes do not change position in the authors' hands.

      (3) Please compare the behavior of the homologues to that of either random loci or control loci on 10, 13, or X to assess whether the differential behavior observed for chromosome 10 is a specific effect.

      Related to points 2 and 3, we opted to use telomeres as controls in this study. Throughout the manuscript, the behavior of chromosome 1 loci is compared to telomeres, demonstrating the specific effect of serum starvation on chr 1. For example, Figure 5A and 5B show that when analyzing mean locus displacement, chr1 and telomeres show the opposite behavior.

      (4) In addition:

      (a) Please provide detailed information on the sequence and location of the probes used.

      We have added a schematic showing the location of the probes as Supplementary Figure S1A. In addition, the sequences are indicated in Materials and methods (page 16).

      (b) Please provide a statistical analysis in all graphs.

      To make statistical analysis more comprehensive, we have added supplementary table 1, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. See also below Reviewer #2 (Recommendations for the authors) point 5.

      (c) Please provide throughout the manuscript in each figure legend information as to how many cells were analyzed in each experiment.

      The number of analyzed loci (or nucleus) is indicated in each graph.

      (d) Please provide information on the spatial and temporal resolution of the imaging modality.

      The imaging settings are indicated in Materials and methods, including the temporal resolution of 0.25 frames per second (page 17). To estimate spatial resolution, and especially its relationship with the observed repositioning of the chromosome loci, we performed experiments in fixed cells, using an optically identical set-up as utilized for live imaging. Unfortunately, the microscope utilized for live imaging was taken out of use by the core facility after submission of the original draft of this manuscript, but we used a microscope with essentially a similar set-up. The data from fixed cells is now presented as Supplementary figure 1E and discussed in results on page 5. This analysis indicates that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error.

      (e) Please better explain what the various measured parameters mean in biological terms.

      We have amended Table 1 to provide better explanation of the measured parameters.

      (f) Please add a scale bar to Figure 6I.’

      Scale bar has been added to figure 6I.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) SHAP analysis identified nuclear area (MA) and its change (sA) as the most predictive features of starvation state, while motility features (MD, MaxD, TD) showed strong interactions with nuclear morphology. Discrete features, such as displacement outliers and homolog subclassification by speed/proximity, influenced classification, particularly in MLP models. Could the authors clarify why morphological and motility features act in combinatorial and context-dependent ways? A biological interpretation of this interdependence would strengthen the study.

      Unfortunately, we do not have a good biological interpretation for this. The fact that some interactions are context-dependent indicates that there could be subpopulations of cells/analyzed loci. For example, we found that the predictive value of nuclear area was high in a subgroup of low motility loci (Figure 4F and Supplementary figure 4D-F). We do not believe that adding more speculation would strengthen the study.

      (2) The manuscript shows that chromosome 1 moves toward the periphery within the first 20-40 minutes of serum withdrawal. However, it remains unclear whether the locus eventually "touches" the periphery and whether it subsequently stabilizes or retracts. It would be valuable to compute the time point of minimal nuclear distance and examine whether this is transient or sustained.

      With the experimental set-up utilized here, we imaged the loci for only two minutes at random time point within the first 40 minutes of the starvation. Hence extracting the time point of minimal nuclear distance is not meaningful from this dataset. As we discuss in the manuscript, following the dynamics of the same locus for longer periods of this would be very interesting in the future. However, this is not in the scope of the present manuscript.

      (3) The manuscript is at times difficult to follow, partly because methodological descriptions are highly detailed in the main text. Consider moving more of the methodological content into Supplementary Methods and emphasizing the main results and interpretations in the main text for clarity.

      We have carefully evaluated this point. Most methodological descriptions in the manuscript relate to the machine learning models and their explanation with SHAP. As we feel that this combination is an essential part of the manuscript, and likely the aspect that can have widest impact beyond chromatin dynamics studies, we feel that the background and our reasoning related to the chosen methods are important.

      (4) The distinction between the first 20 minutes and the latter 40-minute window is intriguing. Could these different time scales be paralleled with early versus delayed gene expression responses to serum starvation? A discussion of this temporal connection would add biological depth.

      This is an intriguing idea, and we have added a short note on this in the discussion (page 13). However, as we do not know how the U2OS cells utilized here respond transcriptionally to serum starvation, we are hesitant to speculate too much.

      (5) If the observed interpretations are robust, could this be demonstrated more explicitly through statistical principles or reproducibility tests across independent datasets?

      To provide further evidence of the robustness of our findings, we have 1) added new data to estimate the spatial resolution (Supplementary figure 1E) and 2) expand the statistical analysis as supplementary table 1. Regarding the spatial resolution (see also the response to reviewer 1), our experiments on fixed cells demonstrate that the change in minimal distance to the nuclear edge, reported in our study under serum starvation in live samples (0.32 and 0.5 micron), is more than one order of magnitude above the static error. Descriptive statistics data are shown on figures as kernel density estimation, confidence intervals and bootstrapped changes distributions. To make statistical analysis more comprehensive, we added a supplementary table, showing the results of inferential statistics, namely, two-sided Mann-Whitney (MW) U-test. MW test was used as a non-parametric statistic, with null hypothesis assuming the samples are coming from the same distribution. Null hypothesis was rejected at the p-value below 0.05. In most cases (bold font in table) MW test results were in accordance with the descriptive statistics confirming the conclusions in the study. In case of exceptions (MD, TD for telomeres and TDist), the results were reported, for example, as an “appearing trend” to reflect descriptive statistics while highlighting certainty levels.

      (6) Figure labeling is difficult to follow. Please include abbreviation explanations directly in the figure panels or legends for clarity.

      Abbreviations have been added to figure legends. Adding them to figures themselves would have made the figures too busy.

      (7) The manuscript reports higher displacement at the nuclear periphery. Can the authors explain why displacement amplitudes increase near the periphery and how this relates to nuclear architecture?

      We speculate in the manuscript (results, page 11; discussion, page 14) that actually the lower displacement observed with the central locus may, at least partially, result from anchoring this locus to the nucleolus (Fig 6I). Nevertheless, alternative explanations, such as differences in transcriptional and/or chromatin states may exist (see also the response to point 11), and this is now mentioned in the discussion (page 14).

      (8) How is the movement of chromosome 1 directed specifically toward the periphery, rather than being random fluctuations? This point requires clarification.

      This is an important question, but unfortunately our data does not provide an answer to this, and suggesting any mechanism would be pure speculation. Nevertheless, our results agree with previous studies utilizing fixed cells that also demonstrated movement of chromosome 1 towards nuclear periphery (Mehta et al., 2010), arguing against random fluctuation.

      (9) Only chromosome 1, and not the other tested chromosomes, undergoes this relocalization. Could the authors elaborate on why some chromosomes but not others display this behavior?

      Previous studies (Mehta et al., 2010) utilizing chromosome paints in fixed cells actually show the relocalization of several chromosomes upon serum starvation. The fact that we observed the relocalization of only chr1 loci is likely due to the labeling method and/or the cell model utilized in this study. This is quite explicitly discussed in the first paragraph of results (page 4).

      (10) The magnitude of these movements appears relatively small (0.5 micron). Can the authors discuss whether such small but reproducible displacements are likely to be biologically meaningful in terms of nuclear function or gene regulation?

      At the moment, our experimental set up allows us to analyze the dynamics of only a small portion of chr1, which indeed shows an average 0.5 micron displacement towards the nuclear periphery. Based on the chromosome painting data from fixed cells, the displacement at the level of whole chromosome is significantly larger. As mentioned in the discussion (page 13), the functional implications of radial repositioning of chromosomes upon serum starvation is not known. Therefore further discussion on the relevance of the magnitude reported here would be pure speculation.

      (11) Peripheral homologs of chromosome 1 became faster and more dynamic under starvation. Why might these loci be more prone to movement? Could this be linked to differences in transcriptional activity or chromatin state between central and peripheral homologs?

      At the moment we favour the idea that the central homolog is constrained by its anchorage to the nucleolus (Figure 6I). However, transcriptional activity and/or chromatin state may also play a role, and this possibility is now mentioned in the discussion on page 14.

      Minor points:

      (1) Figure legends use inconsistent capitalization and panel labels. These should be standardized across all figures for better readability.

      We apologize for these inconsistencies, and have aimed to standardize all labeling.

      References

      Mehta, I.S., Amira, M., Harvey, A.J., and Bridger, J.M. (2010). Rapid chromosome territory relocation by nuclear motor activity in response to serum removal in primary human fibroblasts. Genome Biol 11, R5.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We are grateful to all the reviewers for dedicating time to review our manuscript and for providing insightful comments and suggestions. We have revised our manuscript in line with the reviewers' feedback. The major revisions include characterization of Dcp-1 overexpression-induced cell death, demonstration of the involvement of autophagy in Dcp-1 activation, characterization of the interaction between full-length Bruce and cleaved Dcp-1. We have introduced new figures (Figure 1 – figure supplement 1, Figure 2 – figure supplement 2, Figure 3 – figure supplement 1, Figure 5 – figure supplement 1), new panels (Figures 1D, Figure 3C, Figure 4I, J) and a new table (Table S2). The previous Figure 5 – figure supplement 1 has been relocated to Figure 4 – figure supplement 2.

      With all concerns and suggestions from the reviewers addressed, our conclusion—that Bruce suppresses autophagy-regulated caspase activity and wing tissue growth in Drosophila— is now more robustly supported. We are confident that our revised manuscript makes a significant contribution to the fields of cell death, autophagy, and developmental biology, as it provides a new conceptual framework for understanding non-lethal caspase regulation. We remain hopeful that the reviewers will find it suitable for publication in eLife.

      Reviewer #1 (Public review):

      Summary:

      The authors clearly demonstrate that overexpressed Dcp-1, but not Drice, is activated without canonical apoptosome components. Using TurboID-based proximity labeling, they revealed distinct proximal proteomes, among which Sirtuin 1, an Atg8a deacetylase, which promotes autophagy, was specifically required for Dcp-1 activation. Additionally, the show that autophagy-related genes, including Bcl-2 family members Debcl and Buffy, are required for Dcp1 activation. Using structure-based prediction using AlphaFold3, they identified that Bruce, an autophagy-regulated inhibitor of apoptosis, acts as a Dcp-1-specific regulator acting outside the apoptosome-mediated pathway. Finally, they show that Bruce suppresses wing tissue growth. These findings indicate that non-lethal Dcp-1 activity is governed by the autophagy-Bruce axis, enabling distinct non-lethal functions independent of cell death.

      Strengths:

      This is an excellent paper with very good structure, excellent quality data and analysis.

      Weaknesses:

      This reviewer did not identify any weaknesses or recommendations for revision.

      We sincerely thank the reviewer for their highly positive evaluation of our work. We are pleased that the reviewer found the overall structure, data quality, and analyses to be strong, and that they clearly recognized the key findings of our study. No changes to the manuscript were required in response to this review.

      Reviewer #2 (Public review):

      Summary:

      The Drosophila executioner caspase Dcp-1 has established roles in cell death, autophagy, and imaginal disc growth. This study reports previously unrecognized factors that work together with Dcp-1. Specifically, the authors performed a turboID-based proximal ligation experiment to identify factors associated Dcp-1 and Drice. Dcp-1-specific interactors were further examined for their genetic interaction. The authors report autophagy-related genes, including Debcl and Buffy, to be required for Dcp-1 activation. In addition, the authors present evidence of an interaction between Bruce and Dcp-1. Bruce-expression blocks the Dcp-1 overexpression phenotype. Inhibition of effector caspases or overexpression of Bruce commonly reduced wing growth, suggesting a relationship between the two proteins.

      Strengths:

      On the positive side, the study identifies new Dcp-1-interacting proteins and provides a functional link between Dcp-1 and Sirt1, Fkbp59, Debcl, Buffy, Atg2, and Atg8a.

      Weaknesses:

      The data supporting the Dcp-1/Bruce interaction are not strong, even though the title of this manuscript highlights Bruce. For example, the authors' turboID data does not support Dcp1/Bruce interaction. The case for the interaction is based on a single experiment that overexpresses a truncated Bruce transgene in S2 cells.

      We sincerely thank the reviewer for their constructive and detailed evaluation of our manuscript. We appreciate the positive assessment that our study identifies new Dcp-1-associated factors and provides functional links between Dcp-1 and Sirt1, Fkbp59, and multiple autophagy-related genes, including Debcl, Buffy, Atg2, and Atg8a. We also thank the reviewer for clearly pointing out concerns regarding the strength and interpretation of the evidence connecting Bruce and Dcp-1. In the revised manuscript, we have addressed these concerns in two major ways. First, we provided additional experimental evidence explaining why TurboID-mediated labeling did not identify Bruce. Specifically, we showed that the majority of TurboID-tagged Dcp-1 expressed in wing imaginal discs remains in its full-length form, which is unlikely to engage Bruce. Second, and more importantly, we now demonstrated that endogenously expressed full-length Bruce interacts with cleaved Dcp-1 in wing imaginal discs. These new data provide strong support for a physiologically relevant interaction between Bruce and cleaved Dcp-1. Detailed descriptions of these experiments and results are provided in the point-by-point responses in the “recommendations for the authors” section. Together, these newly added data substantially strengthen the evidence for the Dcp-1/Bruce interaction and support the focus of the original manuscript title.

      Reviewer #2 (Recommendations for the authors):

      (1) The title of the manuscript highlights Dcp-1/Bruce interaction, even though the evidence there is not strong. The evidence for Dcp-1/Sirt1 and Dcp-1/Fkbp59 is stronger. How about changing the title to highlight these other Dcp-1 interactions?

      We thank the reviewer for the thoughtful suggestion. We agree that several Dcp-1-associated factors identified in our study, particularly Sirt1 and Fkbp59, are supported by functional evidence. Specifically, our data show that Sirt1 and Fkbp59 are required for Dcp-1 overexpression-mediated activation. However, Bruce differs from these factors in both the scope and the nature of its effects on Dcp-1. Bruce is not only shown to specifically suppress Dcp-1 activity, but also to suppress wing tissue growth, indicating a broader physiological role in modulating non-lethal Dcp-1 function. Importantly, we further demonstrate that Bruce can specifically physically interact with cleaved Dcp-1. In addition, in this revised manuscript, we show that using the endogenously mStayGold::V5-tag knock-in-tagged Bruce allele, cleaved Dcp-1, induced by overexpression of Dcp-1::VENUS in wing imaginal discs, can be co-immunoprecipitated with full-length Bruce (new Figure 4I, J). These results support a physical interaction between full-length Bruce and activated Dcp-1 in vivo, consistent with a direct inhibitory role. Based on these findings, we decided to retain Bruce in the manuscript title, as it is the only factor for which both physiological and functional interactions with Dcp-1 are supported by multiple independent lines of evidence.

      (2) The case for Dcp-1/Bruce interaction is not strong because the Dcp-1 turboID fails to identify Bruce. In fact, the Dcp-1 turboID approach may not have been effective, as it failed to detect many established interactions, including Diap1 (Wang et al. 1999 PMID 10481910; Tenev et al., 2006 PMID 15580265). The authors may want to comment on this.

      We thank the reviewer for raising this important point. We agree that Bruce, as well as DIAP-1, was not identified in our TurboID-MS labeling dataset (Figure 2C, Table S1). Previous studies have shown that DIAP1 interacts with Dcp-1 and Drice only after exposure of the IAP-binding motif (IBM) at the neo-N-terminus of the large executioner caspase subunit following cleavage (Tenev et al., 2005). Similarly, our co-immunoprecipitation analyses show that Bruce interacts specifically with cleaved Dcp-1, but not with full-length Dcp-1. In the revised manuscript, we confirmed by western blot that the majority of endogenously expressed Dcp-1 in wing imaginal discs is present in the full-length pro-form (new Figure 2 – figure supplement 2A). Thus, the failure to identify Bruce and DIAP1 by TurboID-MS using full-length Dcp-1 as bait is expected, as this approach primarily labels interactors of the inactive, full-length form of Dcp-1. To evaluate whether our proximity labeling approach was nevertheless effective, we compared our TurboIDMS dataset with a previously published immune-affinity purification (IAP)-MS dataset generated using catalytically inactive, C-terminally V5-tagged Dcp-1 overexpressed in Drosophila 1(2)mbn cells (Choutka et al., 2017). Although the experimental conditions differ in several respects, we observed a substantial overlap between the TurboID-MS-mediated and IAP-MS-mediated interaction lists (new Figure 2 – figure supplement 2B, new Table S2). Importantly, SesB, one of the best-characterized Dcp-1 interactors located in mitochondria (DeVorkin et al., 2014), was also identified in our mass spectrometry dataset (new Figure 2 – figure supplement 2B, new Table S2). Based on these analyses, we now more explicitly describe the experimental context and limitations of the TurboID approach, clarifying that it preferentially labels interactors of full-length Dcp-1 in the revised manuscript. We also incorporate comparisons with prior studies to further support the validity of our mass spectrometry experiments in the revised manuscript

      (3) The best experimental evidence for Bruce/Dcp-1 interaction can be found in Figure 4H. But here, they see a weak interaction only when a truncated Bruce construct is overexpressed in S2 cells. Whether Dcp-1 interacts with Bruce in a physiological setting remains unsupported.

      We thank the reviewer for the important comment. We agree that, in the original manuscript, the biochemical evidence for the Bruce/Dcp-1 interaction relied primarily on experiments using an overexpressed truncated Bruce construct in S2 cells and therefore did not sufficiently establish whether this interaction occurs in vivo, especially in wing imaginal discs. To address this concern, we performed additional experiments to examine the Bruce/Dcp-1 interaction. In the background of the mStayGold::V5-tag knocked-in Bruce allele, we overexpressed Dcp-1::VENUS using WPGal4 driver to induce Dcp-1 activation and tested whether full-length Bruce under endogenous expression interacts with cleaved Dcp-1 in wing imaginal discs. Following immunoprecipitation with anti-V5 antibody-conjugated magnetic agarose, we found that cleaved Dcp-1 signal was enriched by co-immunoprecipitation (new Figure 4I, J). These new data demonstrate that Bruce associates with cleaved Dcp-1 in vivo and thus support the physiological relevance of the Bruce/Dcp-1 interaction. We have clarified this point in the revised manuscript and included the corresponding data.

      (4) The genetic interaction between Bruce and Dcp-1 is interesting, but the interpretation becomes complicated because Bruce inhibits Reaper, and at the same time, Dcp-1 genetically interacts with Reaper, Hid, and Grim (Figures 1E, F, G). Thus, it remains unclear if the genetic interaction between Bruce/Dcp-1 is due to a direct interaction between Bruce/Dcp-1 or alternatively, because Bruce inhibits Reaper and Grim.

      We thank the reviewer for the comment. The primary function of Reaper, Hid, and Grim (RHG proteins), collectively referred to as IAP antagonists, is to directly interact with inhibitor of apoptosis proteins (IAPs), most notably DIAP-1 (Kornbluth and White, 2005; Ryoo and Baehrecke, 2010), leading to the inhibition of DIAP-1 function. RHG proteins have not been shown to directly inhibit caspases. Because inhibition of RHG proteins results in the stabilization of DIAP-1, it is likely that the effects observed upon RHG gene knockdown are mediated through DIAP-1. Consistent with this idea, overexpression of DIAP-1, while less potent than Bruce, can also suppress Dcp-1 activation (Figure 5B, C). However, we also acknowledge that Bruce suppresses Reaper- and Grim-dependent, but not Hid-dependent, cell death (Vernooy et al., 2002). In addition, Bruce directly targets Reaper through non-lysine ubiquitination, promoting its degradation (Domingues and Ryoo, 2012). Thus, it is possible that Bruce overexpression suppresses Reaper and thereby strengthens DIAP-1 function, which could indirectly contribute to the inhibition of Dcp-1 activation. Nevertheless, because the effect of Bruce overexpression is stronger than that of DIAP-1 overexpression (Figure 5B, C), and together with our physical interaction data of Bruce with cleaved Dcp-1, we propose that Bruce most likely inhibits Dcp-1 directly to attenuate its activation.

      (5) In general, the manuscript could benefit from highlighting the strong data on Sirt1 and Fkbp59, while clearly acknowledging the limitations of the Bruce/Dcp-1 interaction.

      We thank the reviewer for the comment. As described above, in the revised manuscript we now demonstrate that endogenously expressed full-length Bruce physically interacts with cleaved Dcp-1 in wing imaginal discs (Figure 4I, J). These new data provide strong support for a physiologically relevant interaction between Bruce and cleaved Dcp-1. Based on this evidence, we decided to highlight Bruce in the manuscript, as it is the only factor for which both physiological and functional interactions with Dcp-1 are supported by multiple independent lines of evidence.

      Reviewer #3 (Public review):

      Summary:

      The present paper by Shinoda et al. from the Miura group builds upon findings reported in an earlier study by the same team (Shinoda et al., PNAS, 2019), which identified a nonapoptotic role for the Drosophila executioner caspase Dcp-1 in promoting wing tissue growth. That earlier work attributed this function primarily to Dcp-1 and to Decay, a caspase structurally related to executioner caspases, but not to DrICE, the principal apoptotic executioner caspase. The authors further proposed that this non-apoptotic caspase activity operates independently of the initiator caspase Dronc.

      In the current study, the authors both corroborate aspects of their previous findings and extend the investigation to mechanisms regulating Dcp-1 in this context. They identify roles for the giant IAP Bruce, two BCL-2 family members, and autophagy-related components in modulating nonapoptotic Dcp-1 activity. Moreover, they show that Bruce binds to a BIR-like peptide exposed upon Dcp-1 cleavage, but not to DrICE. The study further suggests that low levels of Dcp-1 activity promote wing tissue growth, whereas excessive activity induces cell death, as evidenced by impaired wing development following Dcp-1 overexpression. Overall, the manuscript provides several intriguing insights into the non-apoptotic regulation of the comparatively weak apoptotic executioner caspase Dcp-1 and complements the group's earlier work. However, several concerns remain regarding certain interpretations of the data and the experimental rigour of some of the results.

      Strengths:

      A major strength of the work is its systematic genetic and biochemical approaches, which combine tissue-specific manipulation with protein interaction mapping to explore how Dcp-1 is regulated. The identification of several regulatory factors, including an inhibitor of cell death protein and components linked to autophagy, provides a coherent framework for understanding how Dcp-1 activity might be tuned.

      Weaknesses:

      The evidence supporting some key claims remains incomplete. In particular, the type of cell death form induced when Dcp-1 is overexpressed is not clearly established, and additional tests would be needed to distinguish between the different cell death types.

      Likely impact:

      The study contributes to a growing body of work showing that proteins traditionally associated with cell death can have broader roles in tissue development. This conceptual advance is likely to be of interest to researchers studying growth control and tissue maintenance.

      We sincerely thank the reviewer for their thoughtful and constructive evaluation of our study. In response to these concerns, we have performed additional experiments to clarify the nature of the cell death induced by Dcp-1 overexpression. Based on the detection of cleaved Dcp-1, the detection of executioner caspase activity, and TUNEL assay, we now conclude that excessive Dcp-1 expression induces typical executioner caspase activity-dependent apoptotic cell death. Detailed explanations and experimental results are provided in the point-by-point responses below. Overall, we believe that these additions strengthen the manuscript by clarifying the dual roles of Dcp-1 in promoting tissue growth at low activity levels while triggering apoptosis when excessively activated.

      Specific points:

      (1) Nature of the wing ablation phenotype

      A central concern is whether the wing ablation phenotype observed upon Dcp-1 overexpression truly reflects apoptotic cell death. The authors show in Figure 1c that nuclei in cells overexpressing Dcp-1, but not DrICE, zymogens are highly condensed, which is suggestive of apoptosis. However, it is equally plausible that this phenotype reflects a form of non-apoptotic, Dcp-1-dependent cell death (e.g. autophagy-dependent cell death). This distinction could be readily addressed using TUNEL labelling and direct caspase activity assays. The latter would be particularly informative, as it remains unclear whether zymogen Dcp-1 is capable of cleaving standard effector caspase reporters in vivo. Does the anti-cleaved Dcp-1 antibody detect Dcp-1 activation following overexpression of the Dcp-1 zymogen?

      We thank the reviewer for this important point regarding the nature of cell death. We agree that nuclear condensation alone is not sufficient to conclude apoptotic cell death, and we therefore performed additional experiments. First, we performed TUNEL staining and detected robust TUNEL-positive signals in wing imaginal discs upon Dcp-1 overexpression (new Figure 1D), supporting apoptotic DNA fragmentation. Second, to directly test whether Dcp-1 overexpression leads to executioner caspase activity in vivo, we used two independent executioner caspase activity probes, GC3Ai (Schott et al., 2017; Zhang et al., 2013) and CD8::PARP::VENUS (Williams et al., 2006). Both probes showed clear executioner caspase activity-positive signals in wing imaginal discs upon Dcp-1 overexpression (new Figure 1 – figure supplement 1C–F), demonstrating that Dcp-1 overexpression leads to executioner caspase activity capable of cleaving standard substrates in vivo. In addition, staining with an anti-cleaved Dcp-1 antibody was positive upon Dcp-1 zymogen overexpression (new Figure 1 – figure supplement 1B), indicating that the overexpressed Dcp-1 zymogen is converted into its active form. Consistent with this result, western blot analysis revealed that Dcp-1 zymogen overexpression results in the appearance of a cleaved Dcp-1 (new Figure 1 – figure supplement 1A). Importantly, consistent with our original observation that the wing ablation phenotype is suppressed by expression of the caspase inhibitor p35, we further showed that p35 overexpression completely abolished the appearance of cleaved Dcp-1 in western blot (new Figure 1 – figure supplement 1A), suggesting that Dcp-1 activation is mediated by self-cleavage. Taken together, these new results demonstrate that Dcp-1 zymogen overexpression induces typical executioner caspase activity-dependent apoptotic cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (2) Role of Decay

      In their earlier study, the authors identified Decay as another caspase influencing wing growth, albeit more modestly than Dcp-1. It is therefore unclear why this line of investigation was not pursued further in the current work. This omission is notable, as Decay is not implicated in apoptosis and, to date, no substantial physiological function has been assigned to this caspase in any system. At a minimum, this point should be discussed explicitly.

      We thank the reviewer for the comment regarding the role of Decay. In our previous study (Shinoda et al., 2019), we demonstrated that both Dcp-1 and Decay promote wing tissue growth in a non-lethal manner. In the present study, however, we focused our analysis on Dcp-1. This decision was based on both technical and biological considerations. From a technical perspective, we had established TurboID knock-in lines and UAS overexpression lines for Dcp-1, Drice, and Dronc, whereas corresponding genetic tools are not available for Decay. From a biological standpoint, Dcp-1 exerts a stronger effect on wing growth than Decay, as shown in our previous work, and exhibits a dual functional spectrum: Dcp-1 promotes tissue growth at low activity levels, whereas excessive activation induces overt cell death. By contrast, Decay has not been implicated in cell death in wing imaginal discs (Kondo et al., 2006). Given that a central aim of the present study was to dissect how executioner caspase activity is differentially regulated to support nonlethal functions versus apoptotic cell death, we therefore focused on the two executioner caspases that are known to participate in apoptosis, Dcp-1 and Drice. We agree with the reviewer that Decay remains an intriguing caspase with largely unexplored physiological roles, and further investigation into its regulation and function will be an important direction for future studies. Importantly, Decay has been shown to mediate Hid-induced cell death in the DIAP1- and apoptosome-independent manner in differentiating photoreceptors and accessory cells of the eye (Leulier et al., 2006). In addition, although not required for cell death, Decay accounts for most of the caspase activity during metamorphic midgut programmed cell death, which is executed by autophagy (Denton et al., 2009). Thus, similar to Dcp-1, Decay might be an executioner caspase that can be regulated independently of the canonical apoptosome-mediated pathway, potentially involving autophagy-Bruce axis, and thereby contributing to the regulation of tissue growth. We have now discussed this point in the revised manuscript.

      (3) Figure 2: Proximity labelling analysis

      The authors use TurboID-mediated proximity labelling to reveal distinct Dcp-1- and DrICEassociated proteomes across tissues, with a particular focus on the wing disc. They further demonstrate that RNAi-mediated knockdown of the Dcp-1-associated proteins Sirt1 and Fkbp59 suppresses the wing ablation phenotype induced by Dcp-1 overexpression, suggesting that these factors are required for Dcp-1 activity. However, it should be clarified whether Bruce was identified as a Dcp-1 interactor in the proximity labelling dataset, given its proposed central regulatory role. In addition, further discussion of Fkbp59, its known functions and how it might mechanistically influence Dcp-1 activity would be valuable.

      We thank the reviewer for the comment regarding the TurboID-based proximity labeling analysis and the interpretation of the identified Dcp-1-associated factors. With respect to Bruce, we clarify that Bruce was not identified as a Dcp-1 interactor in the TurboID proximity labeling dataset. Our co-immunoprecipitation analyses in S2 cells indicate that Bruce interacts specifically with cleaved Dcp-1, but not with the full-length, inactive form. In the revised manuscript, we confirmed by western blot that the majority of endogenously expressed Dcp-1 in wing imaginal discs exists in the full-length pro-form (new Figure 2 – figure supplement 2A). Therefore, the failure to detect Bruce in the TurboID experiment using full-length Dcp-1 as bait is expected, as this approach primarily labels proteins proximal to the inactive form of Dcp-1. To examine the Bruce/Dcp-1 interaction under more physiological conditions, we performed additional in vivo experiments. Using the mStayGold::V5-tag knock-in allele of Bruce, we overexpressed Dcp1::VENUS using WP-Gal4 driver to induce Dcp-1 activation and assessed whether endogenously expressed full-length Bruce associates with Dcp-1 in wing imaginal discs. Following immunoprecipitation with anti-V5 antibody-conjugated magnetic agarose, we found that cleaved Dcp-1 signal was enriched by co-immunoprecipitation (new Figure 4I, J). These new data demonstrate that Bruce associates selectively with the cleaved, active form of Dcp-1 in vivo, thereby supporting the physiological relevance of the Bruce/Dcp-1 interaction. We have clarified this point in the revised manuscript and included the corresponding data.

      FK506-binding proteins (FKBPs) are a conserved group of proteins known to bind FK506, an immunosuppressive drug. FKBPs contain FK domains, which correspond to peptidyl cis-trans isomerase (PPIase) domains. Drosophila Fkbp59 is an orthologue of the mammalian FKBP4 and FKBP5, both of which possess a C-terminal tetratricopeptide repeat (TPR) domain that functions independently of the PPIase domain by mediating protein-protein interactions. The mammalian orthologues of Drosophila Fkbp59 function as Hsp90 co-chaperones (GharteyKwansah et al., 2018). Importantly, loss of Fkbp59 results in pupal lethality (Iki et al., 2020), which precludes further mechanistic analysis on Dcp-1 activation using adult wing phenotypes. To date, the involvement of Fkbp59 in caspase regulation has not been reported. Given that Fkbp59 functions as a co-chaperone, it may facilitate Dcp-1 activation by promoting proper folding, stability, or subcellular positioning of Dcp-1 or its regulatory factors. Importantly, Dcp1 proximal proteins are enriched in chaperone-related factors, including CCT2, CCT8, Droj2, CG16817, Fkbp59, Sgt1, and nudC; seven out of sixteen identified proximal proteins are chaperone-related. These observations suggest that Dcp-1 activity may be regulated by chaperone proteins or that Dcp-1 activity may be spatially restricted to regions enriched in chaperone machinery. Further analysis of the relationship between Dcp-1 activity and chaperone-related proteins will be important to elucidate the mechanisms and functions underlying non-lethal Dcp1 activation.

      (4) Figure 3: Autophagy-related factors

      Given that Sirt1 is known to promote autophagy, the authors next examine autophagy-related proteins and identify roles for Atg2, Atg8a, Debcl, and Buffy in Dcp-1 activation. Notably, these proteins do not promote cell death in the Hid-induced canonical apoptotic pathway. However, it is important to determine whether knockdown of Debcl, Buffy, Atg2, or Atg8a alone affects wing development in the absence of Dcp-1 overexpression, to exclude the possibility that these perturbations independently impair wing formation.

      We thank the reviewer for the comment. To address whether knockdown of Debcl, Buffy, Atg2, or Atg8a independently affects wing development, we performed RNAi-mediated knockdown of each gene using the WP-Gal4 driver in the absence of Dcp-1 overexpression. Under these conditions, knockdown of Debcl, Buffy, Atg2, or Atg8a did not cause any detectable defects in wing morphology (new Figure 3 – figure supplement 1A), indicating that these autophagy-related factors specifically function to suppress Dcp-1-mediated cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (5) Evidence for canonical autophagy

      The involvement of autophagy would be more convincingly demonstrated by testing additional core autophagy genes, such as Atg7, Atg5, and Atg12, as well as performing a combined knockdown of Atg8a and Atg8b. Moreover, direct assessment of autophagy at the cellular level using established genetic reporters would substantially strengthen the conclusions.

      We thank the reviewer for the constructive comment regarding the involvement of canonical autophagy. To further strengthen the evidence that autophagy is required for Dcp-1 activation, we examined additional core autophagy-related genes that function at distinct steps of the autophagy process, in addition to the previously tested Atg2, which mediates autophagosomal membrane expansion, and Atg8a, a core component directly associated with autophagosomal membranes. Specifically, we performed knockdown of genes including FIP200/Atg17, which is required for the initiation of autophagosome formation; Atg9, which is required for autophagosomal membrane nucleation; Atg5, which is required for autophagosomal membrane expansion through Atg12-Atg5-Atg16 ubiquitin-like conjugation system; and Stx17, which is required for autophagosome-lysosome fusion (Umargamwala et al., 2024). Because Atg8b is known to be specifically expressed in the male germline and is dispensable for autophagy, at least in fat body cells (Jipa et al., 2021), we did not further examine Atg8b in wing imaginal discs. Using WPGal4 driver, knockdown of each of these genes significantly suppressed Dcp-1-induced wing ablation phenotype (new Figure 3 – figure supplement 1C), supporting a requirement for canonical autophagy components across multiple stages of autophagosome biogenesis in Dcp-1 activation. Importantly, knockdown of these autophagy-related genes alone did not affect wing morphology in the absence of Dcp-1 overexpression (new Figure 3 – figure supplement 1B), as observed previously for Atg2 and Atg8a, suggesting the suppressive effects are specific to Dcp-1 overexpression-dependent cell death. Together, these results indicate that inhibition of autophagy at any of several key steps can suppress Dcp-1-dependent cell death, demonstrating that intact canonical autophagy is required for Dcp-1 activation. In addition, to directly assess autophagy at the cellular level, we monitored autophagosome formation using mCherry::Atg8a reporter. Upon overexpression of Dcp-1::VENUS in the wing pouch region, we observed a clear accumulation of Atg8a-positive puncta in wing imaginal discs (new Figure 3C), demonstrating that Dcp-1 overexpression induces autophagy in vivo. Together, these results provide both genetic and cellular evidence that canonical autophagy is activated upon Dcp-1 overexpression and is required for Dcp-1-dependent cell death. We have clarified this point in the revised manuscript and included the corresponding data.

      (6) Figures 4-5: Functional consequences

      It would be informative to determine whether Synr, Debcl, or Buffy influence wing size on their own and whether their overexpression enhances wing growth.

      We thank the reviewer for the suggestion regarding the functional consequences of Synr, Debcl, and Buffy on wing size. As requested, we knocked down Debcl or Buffy using WP-Gal4 driver and found that this led to reduced wing size (new Figure 5 – figure supplement 1A), indicating that endogenous Debcl and Buffy promote wing growth potentially through regulating endogenous Dcp-1 activity. We have included the corresponding data in the revised manuscript. Because Synr RNAi did not show any detectable effect on the Dcp-1 overexpression-induced phenotype (Figure 3A, B), we did not further examine the effect of Synr knockdown on wing development alone. Overexpression of Synr was not examined in this study. However, Synr overexpression has previously been reported to induce cell death in wing imaginal discs, resulting in malformed adult wings (Ikegawa et al., 2023), suggesting that increased Synr expression is likely to have deleterious rather than growth-promoting effects. Because Debcl and Buffy are both required for Synr-induced cell death, overexpression of Debcl or Buffy may lead to similar phenotypes. Therefore, we did not test Debcl or Buffy overexpression in the wing imaginal discs.

      (7) Terminology and interpretation of cell death

      Taken together, the results suggest that Dcp-1 zymogen overexpression induces a form of nonapoptotic cell death, potentially autophagy-dependent or related. The reviewer does not understand the authors' insistence on referring to this process as apoptosis. The authors should be more cautious in their terminology: there is no canonical versus non-canonical apoptosis; there is simply apoptosis. Without stronger evidence, these effects should not be described as apoptotic cell death.

      We thank the reviewer for the important comment on terminology and interpretation of the cell death phenotype. As explained in our response to comment #1, we have performed additional experiments to clarify the nature of the cell death induced by Dcp-1 overexpression. Based on the detection of cleaved Dcp-1, the detection of executioner caspase activity, and TUNEL assay, we now conclude that excessive Dcp-1 expression induces typical executioner caspase activity-dependent apoptotic cell death. At the same time, as explained in our response to comment #5, we provide both genetic and cellular evidence that canonical autophagy is activated upon Dcp-1 overexpression and promotes Dcp-1 activation. We recognized that the phrase “Dcp1 activity-regulating alternative apoptosis signaling pathway” used in Figure 5L could be misleading, as it may imply the existence of an “alternative apoptosis”. To avoid this confusion, we have revised the figure legend to read “autophagy-facilitated alternative caspase activation pathway.”

      Reviewer #3 (Recommendations for the authors):

      Figure 1c should be annotated more clearly so that it is evident that the images shown are grouped by genotype.

      We thank the reviewer for the helpful suggestion. We have added lines to Figure 1C to improve clarity by indicating that the images are grouped by genotype.

      References

      Choutka C, DeVorkin L, Go NE, Hou Y-CC, Moradian A, Morin GB, Gorski SM. 2017. Hsp83 loss suppresses proteasomal activity resulting in an upregulation of caspase-dependent compensatory autophagy. Autophagy 13:1573–1589.

      Denton D, Shravage B, Simin R, Mills K, Berry DL, Baehrecke EH, Kumar S. 2009. Autophagy, not apoptosis, is essential for midgut cell death in Drosophila. Curr Biol 19:1741–1746.

      DeVorkin L, Go NE, Hou Y-CC, Moradian A, Morin GB, Gorski SM. 2014. The Drosophila effector caspase Dcp-1 regulates mitochondrial dynamics and autophagic flux via SesB. J Cell Biol 205:477–492.

      Domingues C, Ryoo HD. 2012. Drosophila BRUCE inhibits apoptosis through non-lysine ubiquitination of the IAP-antagonist REAPER. Cell Death Differ 19:470–477.

      Ghartey-Kwansah G, Li Z, Feng R, Wang L, Zhou X, Chen FZ, Xu MM, Jones O, Mu Y, Chen S, Bryant J, Isaacs WB, Ma J, Xu X. 2018. Comparative analysis of FKBP family protein: evaluation, structure, and function in mammals and Drosophila melanogaster. BMC Dev Biol 18:7.

      Ikegawa Y, Combet C, Groussin M, Navratil V, Safar-Remali S, Shiota T, Aouacheria A, Yoo SK. 2023. Evidence for existence of an apoptosis-inducing BH3-only protein, sayonara, in Drosophila. EMBO J 42:e110454.

      Iki T, Takami M, Kai T. 2020. Modulation of Ago2 loading by Cyclophilin 40 endows a unique repertoire of functional miRNAs during sperm maturation in Drosophila. Cell Rep 33:108380.

      Jipa A, Vedelek V, Merényi Z, Ürmösi A, Takáts S, Kovács AL, Horváth GV, Sinka R, Juhász G. 2021. Analysis of Drosophila Atg8 proteins reveals multiple lipidation-independent roles. Autophagy 17:2565–2575.

      Kondo S, Senoo-Matsuda N, Hiromi Y, Miura M. 2006. DRONC coordinates cell death and compensatory proliferation. Mol Cell Biol 26:7258–7268.

      Kornbluth S, White K. 2005. Apoptosis in Drosophila: neither fish nor fowl (nor man, nor worm). J Cell Sci 118:1779–1787.

      Leulier F, Ribeiro PS, Palmer E, Tenev T, Takahashi K, Robertson D, Zachariou A, Pichaud F, Ueda R, Meier P. 2006. Systematic in vivo RNAi analysis of putative components of the Drosophila cell death machinery. Cell Death Differ 13:1663–1674.

      Ryoo HD, Baehrecke EH. 2010. Distinct death mechanisms in Drosophila development. Curr Opin Cell Biol 22:889–895.

      Schott S, Ambrosini A, Barbaste A, Benassayag C, Gracia M, Proag A, Rayer M, Monier B, Suzanne M. 2017. A fluorescent toolkit for spatiotemporal tracking of apoptotic cells in living Drosophila tissues. Development 144:3840–3846.

      Shinoda N, Hanawa N, Chihara T, Koto A, Miura M. 2019. Dronc-independent basal executioner caspase activity sustains Drosophila imaginal tissue growth. Proc Natl Acad Sci U S A 116:20539–20544.

      Tenev T, Zachariou A, Wilson R, Ditzel M, Meier P. 2005. IAPs are functionally non-equivalent and regulate effector caspases through distinct mechanisms. Nat Cell Biol 7:70–77.

      Umargamwala R, Manning J, Dorstyn L, Denton D, Kumar S. 2024. Understanding developmental cell death using Drosophila as a model system. Cells 13:347.

      Vernooy SY, Chow V, Su J, Verbrugghe K, Yang J, Cole S, Olson MR, Hay BA. 2002. Drosophila Bruce can potently suppress Rpr- and Grim-dependent but not Hid-dependent cell death. Curr Biol 12:1164–1168.

      Williams DW, Kondo S, Krzyzanowska A, Hiromi Y, Truman JW. 2006. Local caspase activity directs engulfment of dendrites during pruning. Nat Neurosci 9:1234–1236.

      Zhang J, Wang X, Cui W, Wang W, Zhang H, Liu L, Zhang Z, Li Z, Ying G, Zhang N, Li B. 2013. Visualization of caspase-3-like activity in cells using a genetically encoded fluorescent biosensor activated by protein cleavage. Nat Commun 4:2157.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors describe an early diverging vertebrate KCNE gene present in jawless lampreys that they denote KCNE0.

      Three forms of the protein are isolated from different lampreys, which have 95% homology to each other, but only moderate homology to KCNE1-6.

      Co-expression with lamprey KCNQ1 produced a non-inactivating current, whereas co-expression with mammalian KCNQ1 resulted in less modulation. Introduction of a tetra-leucine motif from KCNE4 into KCNE0 reduced current on co-expression with KCNQ1, conferring an inhibitory effect.

      Strengths:

      This is an interesting and uncontroversial report of a new KCNE isoform from lower vertebrates that gives insight into the evolutionary progression of the sequence and functional properties of the accessory protein.

      Thank you for reviewing our manuscript and for your constructive comments. Our point-to-point responses are shown below.

      Weaknesses:

      (1) No error bars visible for lamprey Q1 isoforms (open symbols) in Figure 2G. No statistical comparison was provided to indicate whether lamprey Q1 isoform V1/2s are significantly different (nor in Supplementary Table 1).

      (2) There is the same issue in Figures 3 and 4. No appropriate statistical comparison is made between V1/2s for different truncations of PmKCNE0 (Figure 3), or between KCNQ1 species isoforms with and without PmE0.

      We thank you for these helpful comments. Based on your suggestions, we revised the presentation of error bars in Fig. 2G and in other panels showing G–V or F–V relationships (Figs. 2J, 3F, 3N, 4C, 4F, 4I, 4L, and 5E; Supplementary Fig. 5D) to make the SEM bars clearer. We also added statistical comparisons of V<sub>1/2</sub> values among the three lamprey KCNQ1 orthologs in Fig. 2G and among truncation-series constructs in Figs. 3F and 3N using one-way ANOVA followed by Tukey–Kramer multiple-comparison tests. For Fig. 4, we added statistical comparisons between KCNQ1 species isoforms expressed with or without PmKCNE0 (Figs. 4C, 4F, and 4I), and between PmKCNQ1 expressed alone or with human KCNE1 or KCNE3 (Fig. 4L), using unpaired two-tailed Welch’s t-tests. These statistical comparisons are included in Supplementary Table 1.

      Reviewer #2 (Public review):

      Summary:

      This study functionally characterizes a single KCNE-like gene, kcne0, from a jawless vertebrate. The authors conducted multiple experiments, including TEVC, VCF, RT-PCR, and RNA-seq to show that KCNQ1 and kcne0 exhibited a broadly overlapping organ distribution in lamprey species, and KCNE0 produced a constitutively active current when co-expressed with lamprey KCNQ1, similar to the effects of human KCNE3 on KCNQ1. This modulation was species-specific, as co-expression of KCNE0 with other species' KCNQ1 was less effective. Moreover, the authors found that truncating the N-terminal had a more significant reduction of the modulatory effects than truncating the C-terminal of KCNE0. Interestingly, the introduction of the tetra-leucine motif from human KCNE4 into KCNE0 conferred KCNE0 with comparable effects of human KCNE4 on KCNQ1.

      Strengths:

      The authors clearly introduced an early-diverging member of the KCNE family, and convincingly demonstrated the function of this gene, KCNE0. The results are supported by experiments of multiple approaches and are clearly written. The work is significant and will interest readers from the extended research area.

      Weaknesses:

      No major concerns were identified with the manuscript in general.

      We thank you for the positive assessment of our work and for the constructive suggestions. Our point-by-point responses are provided below.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What is the physiological role of this KCNE0 and lamprey KCNQ1 in the lamprey species? While the authors mention that the physiological roles of KCNE0 are the next focus, it is preferable to discuss some of the potential functional significance of this newly characterised KCNE.

      We thank you for this helpful suggestion. We agree that discussing the potential physiological roles of KCNQ1–KCNE0 complexes in lamprey strengthens the manuscript. We have therefore expanded the Discussion to raise the possibility that, given the broad tissue distribution of kcne0 transcripts and the ability of KCNE0 to render lamprey KCNQ1 constitutively active, KCNQ1–KCNE0 complexes may contribute to general ion homeostasis, potentially analogous to the epithelial K<sup>+</sup> recycling function of mammalian KCNQ1–KCNE3, rather than to the highly specialized KCNQ1–KCNE1 function in the mammalian heart and inner ear (page 14, lines 263–267).

      (2) Human KCNQ1 has 676 amino acids, but LcKCNQ1 contains just 507 amino acids. The species-specific regulatory effects of KCNE0 may not only be attributed to KCNE0 itself but might also be influenced by the species of KCNQ1. Some discussion on this possibility will be helpful.

      We thank you for raising this important point. We agree that the species-specific regulatory effects observed in our cross-species pairing experiments are unlikely to be determined by KCNE0 alone and may also be influenced by species-specific features of the KCNQ1 α-subunit. To address this point, we expanded the Discussion to note that the KCNQ1 proteins used in this study vary in amino-acid length, largely reflecting differences in the cytoplasmic C-terminal region, which may affect KCNQ1–KCNE compatibility and thereby influence channel gating and coupling to KCNE subunits (pages 12–13, lines 227–238). We also updated Supplementary Fig. 4 to include LrKCNQ1 and LcKCNQ1, and clarified that the LcKCNQ1 construct used in this study encodes 644 amino acids.

      (3) In Supplementary Figure 3, bands corresponding to LcKCNQ1 (507 amino acids, Supplementary Figure 1) were not seen.

      We thank you for pointing out this potentially confusing point. Supplementary Fig. 3 shows RT-PCR products amplified from tissue cDNA, not full-length amplification of the LcKCNQ1 ORF. As stated in the Methods section (pages 19–20, lines 374–399), the primers used for RT-PCR in Fig. 1F and Supplementary Fig. 3 were different from those used for cloning the full-length LcKCNQ1 cDNA shown in Supplementary Fig. 1. Therefore, a band corresponding to the full-length LcKCNQ1 coding sequence was not expected in Supplementary Fig. 3.

      To clarify this point, we revised the figure legends and indicated the RT-PCR primer-binding sites with orange arrows in Supplementary Figs. 1 and 2.

    1. Author response:

      We thank the reviewers for their constructive and careful assessment of our manuscript. We are encouraged that both reviewers recognised the value of the empirical contribution: the semantic advantage is robust across experiments, the two new experiments are pre-registered, and the drift-diffusion modelling provides an informative decomposition of behavioural performance. At the same time, both reviewers raise an important and convergent point: the manuscript currently places too much interpretive weight on non-decision time and sometimes moves too quickly from decision-model parameters to claims about the representational format of working memory.

      We agree that this aspect of the manuscript should be revised. In the next version, we will substantially soften claims about adaptive reformatting and long-term-memory-like formats. We will instead frame the central contribution more precisely: semantic-category judgements show a reliable advantage at stages preceding evidence accumulation and this advantage is modulated by attentional prioritisation and interference during the maintenance interval. Our data constrain the dynamics with which different kinds of information become available for WM-guided decisions, but they do not, on their own, provide a direct measure of representational format. This hypothesis should be tested in future experiments.

      At the same time, we think the data provide stronger constraints on alternative explanations than the current manuscript makes clear. The reviewers correctly note that non-decision time is not a pure retrieval parameter, as we also note in the discussion. It can include probe encoding, response preparation, motor execution, and other processes. We will therefore avoid more explicitly equating NDT directly with retrieval latency. However, many of the alternatives raised by the reviewers, such as easier question reading or simpler response mapping for semantic probes, predict a relatively fixed semantic–perceptual offset. In our experiments, the probes and response mappings are held constant across attentional conditions, while the semantic NDT advantage changes as a function of whether the relevant item can be prioritised in advance or must be selected/reactivated at test. We will restructure the Results and Discussion to make these condition × feature interactions central to the argument.

      We will also clarify the logic of Experiment 1. We agree with Reviewer 1 that a valid retro-cue likely triggers retrieval or reactivation of the cued item. Our original phrasing, which described the valid-cue condition as reducing retrieval demands, was imprecise. The critical manipulation is better described as shifting item prioritisation/retrieval earlier in the trial. Under valid cueing, the relevant item can be prioritised before the probe appears, whereas under neutral cueing, item selection and access must occur after probe onset. We will rewrite this section accordingly.

      We will also clarify our operationalisation of semantic and perceptual categories. The present contrast is specifically between semantic category information (animate versus inanimate) and perceptual-format information (photograph versus drawing). We agree that the perceptual judgement is still categorical and does not measure fine-grained perceptual fidelity. We will therefore avoid broad claims about semantic structure or perceptual detail in general. However, as pointed in the manuscript, we believe the contrast remains meaningful: the two dimensions are orthogonal within the same stimuli, and previous work using the same feature space showed the opposite ordering during perception (Linde-Domingo et al., 2019), where perceptual-format information was available before semantic-category information. We will move this argument earlier in the manuscript and present it as converging evidence for dissociable access dynamics, while acknowledging that it does not by itself prove representational format.

      In response to the modelling concerns, we will expand the model-validation section. Specifically, we plan to add posterior predictive checks for the reported models, report model comparisons more transparently, clarify when more complex models do or do not provide practically meaningful improvements, and include sensitivity analyses using alternative parameterisations where identifiable.

      We will also make several methodological clarifications. First, because the reanalysis of Kerrén et al. (2022) forms a substantial part of the manuscript, we will add a fuller description of the original task in the main text, including how the probed item was indicated at test. Second, we will rewrite the unclear sentence describing pseudo-random stimulus selection in Experiment 1 and add a control analysis testing whether performance differs when the probed item belongs to the majority versus minority category within the trial. Third, we will clarify the stimulus repetition scheme and discuss possible long-term-memory contributions. Importantly, because semantic and perceptual probes are applied to the same items from the same trials, any repetition history or proactive-interference contribution is shared across the two probe types, although we agree that this should be discussed explicitly.

      Finally, we will revise the broader theoretical framing. We will remove or substantially qualify claims linking the present data directly to episodic memory and imagery. We will also integrate the recent literature suggested by Reviewer 2 on semantic structure, associative relations, long-term-memory contributions to working memory, and boundary conditions for semantic labelling effects. This will allow us to position the study as one piece of a broader literature on how semantic information influences WM performance, rather than as direct evidence for a general representational reformatting mechanism.

      In summary, the revised manuscript will make a narrower but stronger claim: semantic-category information shows a robust pre-accumulation advantage during WM-guided decisions, and this advantage is shaped by attentional prioritisation and interference during maintenance. We will present this as evidence about WM access dynamics and decision components, not as direct evidence that WM representations are transformed into long-term-memory-like formats.

    1. Author response:

      eLife Assessment

      The manuscript by von Velsen et al. offers valuable structural insights into the mitogen-activated protein kinase (MAPK) pathway by providing cryo-EM structures of stabilized MEK1-ERK2 kinase-substrate complexes in inactive, active, and nucleotide-free states, complemented by HDX-MS, SAXS, ITC, crystallography, and molecular dynamics. The work provides solid evidence for the overall architecture of the complex and identifies interaction sites that help explain MAPK pathway specificity. However, some mechanistic conclusions are not yet fully supported, particularly the designation of one state as an active phosphoryl-transfer configuration, the claim that substrate binding releases the MEK1 catalytic machinery, the proposed link to processive phosphorylation, and the extrapolation to disease-associated mutations.

      We would like to counter the final statement. We were very careful in our description of state 2, while we describe it as ‘active’ we clearly explain that the resolution of the reconstruction is not sufficient to define all the classical indicators of a kinase active state; however, the map is consistent with the active conformation, the complex is active in vitro, and the MEK1 variant used is the well known DD mutant that is constitutively active. While the A-loop is not observed, this is in agreement with many crystal structures of other DD mutants. We therefore decided to define this state as ‘active’ as the A-loop of ERK2 approaches the active site, the alpha-C helix has moved in and the A-loop of MEK1 no longer occludes the active site - to clarify the state we refer to the classically active confirmation as ‘fully active’. Our supporting data also show that the complex is highly dynamic during turnover, meaning we have captured MEK1 in a number of conformations on the landscape of an active state – we feel that rather than a limitation, this is an important observation in MAP2K studies. Finally, the determination of an 80 kDa complex by cryoEM to resolutions well below 4 Å is a huge technical achievement allowing the first snapshots of the MEK1-ERK2 complex to be visualised.

      Regarding the A-helix release – our observation is that the helix becomes less folded on binding of substrate. There are many studies, which we cite, that show that destabilising this helix leads to release of the catalytic machinery, see Mansour et al, 1996, Biochemistry, 35, 15529-15536 and Jindal et al 2017 J. Biol. Chem. 292, 18814-18820 for initial studies. Our observation shows that this is linked to substrate binding – a very relevant new insight that demonstrates the importance of this helix, in addition to many previous studies, but links unfolding to substrate recognition for the first time.

      For the mechanism of processive phosphorylation – it has been well established that both processive and distributive mechanisms exist. While the way that a distributive mechanism could work is obvious (complete dissociation of the two proteins), it has not been clear how a MAP2K can remain bound to its substrate and exchange nucleotides. While caution should be employed in interpreting our state 3 structure, it clearly shows what nucleotide exchange when bound to substrate can look like and that this low nucleotide affinity state is linked to disorder in the P-loop, the A-helix and substrate binding via the KIM. We would love to perform experiments that could demonstrate this but cannot at present think of an appropriate method – the reviewers did not suggest a route either.

      Finally, for the cancer-causing mutations – there are many studies demonstrating that the mutations lead to a destabilisation of the A-helix. Our study links this to substrate recognition. While this is inference, it seems justified to describe a link between substrate recognition, A-helix unfolding and disease mutations given the large body of literature describing these events.

      We are currently performing a series of in-cell activity assays that should strengthen our claims regarding the A-helix and other observations in the structure - the histidine interactions in particular.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes three conformers derived from a complex between ERK2-T185V, a variant of MEK1-DD with the KIM sequence replaced by the KIM from the p38 activator, GRA24, ADP, and AlF4-. The goal was to try to capture the complex in its active state. The results show contacts between the kinases between their N-lobe and their C-lobes that resemble MKK6-p38 complexes previously reported by the authors. Two MEK1-ERK2 conformers (States 1,3) are deemed inactive based on the lack of access of ERK-Y187 to the MEK1 active site, and the absence of ADP bound to MEK1 in State 3, while one conformer (State 2) is deemed active, but not fully active due to disorder in MEK1 activation loop (A-loop) and an essential salt bridge between strand beta3 and helix aC. HDX-MS and SAXS solution measurements and all-atom MD simulations are used to model the mutant complex and variants with WT ERK2. The study concludes that substrate recognition involves low-energy contacts with MEK, allowing substantial protein flexibility within the complex in a manner that may accommodate processive phosphorylation of ERK2.

      Strengths:

      The strengths of the work are that the findings provide important structural insights for MEK-ERK signaling and protein phosphorylation in general. These are valuable given that atomic resolution structures of kinase-substrate complexes are still limited in number. The authors succeeded in showing key contacts between subunits and conformational variations within the complex.

      Weaknesses:

      Weaknesses were that some of the conclusions about activity state, dynamics, and effects of ligand binding were less convincing. For example, that State 2 truly represents an active configuration seemed ambiguous, given the absence of Mg2+ and AlF4- in the cryoEM structure and disorder in the activation loop and the K97-E114 salt bridge. Conclusions by SAXS that ADP-AlF4 binding increases active site compaction while increasing local flexibility were not rigorously supported by HDX data, given that the latter were performed without ligand. Sections of the narrative and figures throughout were often confusing, and many assertions were made without clear explanation. Data shown in the supplementary materials were not always described in the Results, even those important for the conclusions. Figure legends and text lacked clear descriptions of specific complexes analyzed. Substantial changes are recommended to improve the readability and clarity of the work.

      We thank reviewer #1 for in-depth comments and analysis of our manuscript. However, there is a misunderstanding regarding HDX-MS and SAXS data. First, the HDX-MS data were performed on the ADP.AlF<sub>4</sub><sup>-</sup> inhibited complex - this was not made sufficiently clear in the text, and we will amend this. Secondly, we are not trying to support local flexibility observed in the SAXS data with the HDX data. The HDX data support the interactions observed in the cryoEM structure and demonstrate flexibility in the proline-rich loop, the ERK A-loop and unfolding of the MEK1 A-helix. The SAXS data demonstrate that when the transition state complex is formed, the complex is more compact but flexibility within the complex increases - as observed in the dimensionless Kratky plot, supporting our observations in the cryoEM maps. Therefore, the HDX data and SAXS data are separate observations. We thank reviewer #1 for all the comments and will rewrite the manuscript in order to increase clarity as suggested.

      Reviewer #2 (Public review):

      Summary:

      The authors used Cryo-EM to obtain a complex between MEK1 and ERK2. They used the same method as previously used by the same authors to form a stable complex between MKK6 and p38, an extra-strong KIM replacing the wild-type KIM in MEK1. Three conformers were resolved, with the highest resolution of 3.0 Å. The multiple conformers indicate more flexibility in MEK2 than in ERK2. These data suggest that nucleotide exchange is possible while maintaining MEK1-ERK2 interactions. SAXS and HDX data reinforce the idea of flexibility. They point to interactions between the two N-terminal domains between histidines at the N-terminus of helix C and between the G helices that are maintained in each of the 3 conformers, and sequence and structure suggest these histidines may be a source of specificity in MEK1-ERK2 versus MKK6-p38 interactions. A 2.2 Å structure of a complex between ERK1 (88% identical to ERK2) and the docking peptide used was also presented. Molecular dynamics simulations suggest that the MEK1-ERK2 complex can assume a fully active configuration of MEK1.

      Strengths:

      This is the first structure of a MEK1-ERK2 complex. The structural data are valuable additions to our understanding of MAP2K-MAPK interactions. The discussion points offered in the results section are palatable. These include the origins of specificity and the idea of flexibility in the MAP2K in support of a processive mechanism for the dual phosphorylation activity of MAP2Ks.

      Weaknesses:

      (1) This reviewer considers that the abstract is overstated. Specifically, this paper does not reveal the molecular details of phosphoryl transfer, nor does it demonstrate that substrate binding releases the catalytic machinery.

      (2) The discussion is in some places not supported by evidence and in others has superfluous text. Examples follow:

      - "Once the αG-helix is docked, and the C-lobe histidine triad is in place, the N-lobe interactions must then be fulfilled." The data in this paper does not suggest an order of events.

      - "If the substrate MAPK is incorrect, the N-lobe interaction will not be stabilised, preventing alignment of the MAPK A-loop with the MAP2K active site." This statement could be described as obvious.

      (3) Much of the discussion is embedded in the results, such that it is difficult to separate new facts offered by the paper from speculation.

      We thank reviewer #2 for comments and thorough analysis of our manuscript. We agree that perhaps the abstract should be toned down in terms of claims of an active conformation even though we feel that the combination of the first structure of the MEK1-ERK2 complex combined with MD simulation studies clearly demonstrate how phosphoryl transfer will occur. We are now also performing in-cell assays to support our theory of A-helix regulation. For point 2 we based the order of events on data from Juyoux et al 2023 Science, 381, 1217-1225, where in long-timescale MD simulations and experimentally validated adaptive Markov state model simulations the KIM interaction was the last to dissociate after the alpha-G helix interaction. In our MD simulations of the MEK1-ERK2 complex, the interactions formed by the N-lobe were weaker than those formed by the alpha-G. Indeed, dissociation of the N-lobe was observed in various independent simulations, whereas dissociation of the alpha-G was observed only once.

      Assuming that, as in the MKK6-p38a complex, the association proceeds along the reverse of the dominant dissociation pathway, the simulations suggest that the KIM interaction forms first, followed by the alpha-G and finally the N-lobes. While alternative association pathways are possible, this interpretation is consistent with the MD and in line with the main association and dissociation pathway observed for the MKK6-p38a complex. This is additionally supported by the observation that there is no catalytic activity if the KIM is removed, demonstrating this as the first essential recruitment event. We will expand this section to include our arguments.

      For the second example, we feel this is rather unfair. The statement that if the His-His interaction is absent, catalysis will be prevented is only obvious if one knows about the His-His interaction - this is the first structure showing pathway-specific interactions in the variable loop regions of a MAPK. If it is obvious, why has no one described these residues as important before?

      We have taken on board the comments on the style of the manuscript and will make significant changes as suggested by both reviewers.

    1. Author response:

      Reviewer #1

      (1) Causality and the role of DMS; a specific, falsifiable prediction. We agree that the paper should not leave the causal question implicit, and we will expand the Discussion to state a concrete prediction rather than a general claim of involvement. Briefly, if the accumulation signal we describe carries the animal's intended departure time, then suppressing DMS during patch occupancy should not simply shift exit times in one direction but should degrade their structure: exit-time variability should increase, and exit timing should lose its systematic dependence on reward-rate context and on the time of the most recent reward. A plausible alternative outcome is disengagement from the task altogether, which would be uninformative and would need to be controlled for. The result that would falsify our hypothesis is the one we will state explicitly: exit timing that remains as predictable, and as sensitive to context and reward history, under DMS suppression, as without it. We will also discuss the alternative the reviewer raises, that these signals are inherited from cortical inputs such as OFC or mPFC with DMS acting as a relay, and note that our data cannot presently distinguish this from a locally generated signal.

      (2) Behavior under a normative strategy. This is an interesting question and we will address it in the Discussion. Our expectation, which we will frame as a prediction rather than a result, is that an animal timing from patch entry rather than resetting at each reward would show accumulation that begins at entry and proceeds to a context-dependent threshold at the reward-rate-optimal time, rather than the reward-triggered resets we observe. In this view, the reset structure of the neural signal is a reflection of the behavioral policy rather than a property of the region. We will make clear that this is a testable prediction that our current dataset does not address.

      (3) Do these neurons also track time at the context port? We intend to answer with new analysis, and it converges with Reviewer #2's suggested analysis (below), so we treat the two together there.

      Reviewer #2

      (a) Timing versus movement initiation: activity at the context port. We take this to be a central interpretational concern. We will examine whether the step-like DMS activity extends to the context port, testing the interval between the final context-port reward and departure for the same step-like transitions and accumulation we observe in the time-investment port. We will apply the same comparison to the context-port inter-reward intervals, which addresses Reviewer #1's third point about whether these neurons also track time at the context port.

      We want to flag one feature of the task that bears on how the outcome should be read. The context port is not a timing-free epoch. Its four rewards are delivered at predictable, regularly spaced intervals, and the interval between the final reward and the animal's departure is self-timed. Departure from the context port is therefore also a self-timed action, and observing accumulation there would not by itself indicate that the signal reflects movement initiation rather than timing. What the comparison can inform is whether the accumulation is specific to a decision about when to disengage from a depleting resource, or is a more general feature of self-timed departures. This is a meaningful distinction either way, and one we will report and interpret whichever direction the result falls.

      (b) Operational definition of leaving; the exit-to-entry latency. We agree these details are necessary for interpretation and their absence is our omission. Exit is the final withdrawal from the time-investment port preceding the next context-port visit, and we will make that clear in the revised methods. Mice can and occasionally do re-enter the time-investment port without an intervening context-port visit (especially early in training). Such re-entries are not counted as exits. We will also examine whether the latency between time-investment-port exit and context-port entry relates to the accumulation slope on the corresponding interval, to test whether the neural signal relates to the decision or to the execution of the transition.

      (c) Independence of interval-level fits from the session-level model. This is a fair characterization of the procedure, and we accept the distinction the reviewer draws between a detector of consistency with a session-defined step model and an unbiased test of discrete versus continuous dynamics. We will address it with a held-out validation: estimating each unit's state parameters and transition time from one half of its intervals and testing whether the transition times recovered from the withheld half agree. The discrete-versus-continuous comparison is addressed directly by the ramp simulations under Reviewer #3's point (5) below.

      (d) The inclusion threshold for the accumulation analysis. We will add a supplementary figure reporting the total number of recorded units per session and the fraction classified as step-like, so that the seven-unit criterion can be evaluated against the recorded population rather than in the abstract. Yield varied substantially across sessions, from a handful of units to roughly one hundred, and we will show this distribution directly. We will also report the accumulation results across a range of inclusion thresholds spanning approximately five to eight simultaneously recorded step-like units, so that readers can assess sensitivity to the choice.

      Reviewer #3

      (1) Specification of the task. We agree the task description was insufficient, and we will correct this at the front of the Results and in the Methods. The time-investment port operates on a poke-and-hold basis: the mouse maintains its head in the port and rewards are delivered stochastically over time for as long as it remains, with no requirement to withdraw and re-poke. Reward delivery follows an exponentially decaying rate in time since port entry, with a time constant of eight seconds, integrating to an expected eight rewards of one microliter each (8 µL total) for indefinite occupancy; we will give the explicit function. The reward probability function in the time-investment port is identical across blocks. The high- and low-reward-rate contexts are properties of the context port alone (four rewards over five seconds versus four rewards over ten seconds), and we will make this contrast unambiguous, since it is the manipulation on which the design rests.

      (2) Derivation of the optimum, Figure 1h, and a formal test of overstaying. We appreciate this comment. The optimal residence time in our task is not obtained by the classical Charnov tangent construction. It is computed by explicit maximization of the overall reward rate over the full cycle, following the framework in Sutlief et al. (2025) and shown graphically in Figure 1h. We will expand the legend of Figure 1h so its conventions are stated explicitly, give the reward-rate-maximizing derivation as an explicit equation in the Methods, and reframe the surrounding text around reward-rate maximization as the normative principle, with MVT identified as the special case it is. We will also add the formal statistical comparison of observed residence times against the computed optimum, which the reviewer correctly notes was asserted rather than tested.

      (3) Interpretation of SVM coefficients with correlated predictors. We accept this criticism. We will not rest the ordering of predictors on coefficient magnitudes alone. We will add a variable inclusion-and-ablation analysis, reporting cross-validated classification performance as each predictor is added to and removed from the model, so that the contribution of time since last reward is assessed by its effect on performance rather than by its normalized weight. We will additionally add a complementary analysis of the leave hazard that estimates the contribution of each variable without requiring the classification framework.

      (4) The five-second post-exit window. The reviewer is right that we did not explain this choice, and right that it rests on an assumption. Our reasoning was that a single exit moment per trial leaves the decision boundary badly under-constrained, and that treating the moments immediately following an exit as conditions under which the animal would also have left is licensed by the fact that within-patch reward rate declines monotonically with time, so conditions in the counterfactual continued visit would have been strictly less favorable than those already rejected. We accept that this is an assumption rather than an observation and will state it as such. We will also report the analysis across a range of window durations so that the independence of the result on this choice is visible. To the reviewer's specific questions: both time since entry and time since last reward are computed with respect to the time-investment port throughout, and we will state this explicitly.

      (5) False positive rate against ramps and random walks. We agree this is the more informative validation, and that continuous ramping is the most relevant alternative to ours. We will generate simulated units with continuous ramp-like rate changes, matched to the firing rates and interval structure of our recorded units, and pass them through the identical classification pipeline to obtain false positive rates comparable to the Poisson analysis already reported. We will retain the flat-rate Poisson simulation and present the ramp results as additional panels of the same supplement. We will also explore random-walk dynamics. Together with the held-out validation of transition times described under Reviewer #2(c), this converts the step characterization from a single-null validation into a comparison against the relevant continuous alternatives.

      (6) Specificity of the accumulation signal versus a generic increasing function. This is a valuable challenge and we will address it directly. We will implement the trial-mismatch control the reviewer proposes, randomly reassigning accumulation trajectories to reward-to-exit intervals within session and showing the extent to which predictive performance degrades relative to the correctly matched case.

      We would also note two features of the existing results that speak to this concern, and which we will bring forward in the revision because we did not make them salient enough. First, a signal that merely tracked elapsed time would be expected to shift its starting level as well as its rate across trials with different exit times; instead, the accumulation slope is strongly related to exit time (mean r = -0.551) while the intercept is not (mean r = 0.013), indicating a variable rate from a stable origin. Second, and more to the point, the accumulation arrives at a common level at the moment of exit whether the animal leaves early or late. The rate of accumulation shifts with the animal’s policy on that trial such that the threshold is met at the intended time. 

      What makes this predictor non-trivial is its trial-by-trial correspondence to behavior, not simply that it increases over time. We will make this argument explicitly alongside the shuffle control.

      Summary

      To summarize the planned additions: (i) analysis of step-like activity and its accumulation at the context port, with the interpretive caveat noted above; (ii) validation of step detection against ramping alternatives, together with held-out estimation of transition times; (iii) a trial-mismatch control for the specificity of the accumulation predictor; (iv) an inclusion-and-ablation analysis of the behavioral predictors and a complementary hazard model of the leave decision; (v) reporting of unit yield, step-like fraction, and sensitivity of the accumulation results to the inclusion threshold; (vi) analysis of the exit-to-context-entry latency in relation to the neural signal; (vii) a formal statistical test of overstaying relative to the computed optimum; and (viii) substantial clarification of the task specification, the operational definition of exit, and the derivation of the optimal residence time, including an expanded Figure 1h legend.

      Our aim in the revision is to meet the specificity concern raised in the assessment as directly as the existing data allow, and we hope the revised manuscript will warrant reconsideration of the strength-of-evidence characterization.

      We are grateful to the reviewers for the care evident in their reports, and to you both for handling the manuscript.

      References

      Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. Theoretical Population Biology, 9(2), 129–136.

      Sutlief, E., Walters, C., Marton, T., & Hussain Shuler, M. G. (2025). The value of initiating a pursuit in temporal decision-making. eLife. https://doi.org/10.7554/eLife.99957.2.

    1. Author response:

      The following is the authors’ response to the current reviews.

      Reviewer #1 (Recommendations for the authors):

      (1) Interpretation of Syllable-Tracking in the RND Condition:

      The finding of greater syllable-tracking in the LL group compared to the HL group in the RND condition warrants cautious interpretation. Currently, there is no direct statistical evidence demonstrating greater PLV at 4 Hz in the Structured versus Random conditions for either group; readers must infer this solely from numeric differences in Figure S5 B and D. Therefore, while the interpretation on Page 14 (Lines 443-446) "successful segmentation may enhance syllable tracking via top-down predictions of the next syllable" is an interesting speculation, it feels somewhat far-reaching. Additionally, the authors should discuss whether this upregulated syllable tracking in the structured condition (which is specific to the HL group) represents an adaptive or maladaptive response.

      The reviewer correctly highlights the lack of direct  comparison between conditions (RND versus STR). We tempered our claims in the cited paragraph and insisted on the speculative nature of this part of the discussion. We also clarified that, to us, it may represent an adaptive compensatory strategy:

      Page 14, line 441: “Interestingly, our supplementary analyses (Supplementary Material Figure S4-5) suggest that syllable entrainment may be differentially affected in HL versus LL infants, depending on the statistical structure of the input stream (RND versus STR). However, as our experiment was not explicitly designed to test stream effects, these results should be interpreted with caution. Future studies could explore how successful segmentation may enhance syllable tracking via top-down predictions of the next syllable in both LL and HL infants. If confirmed, such a mechanism may improve alignment to syllable onsets, potentially constituting a compensatory process allowed by preserved segmentation abilities.”

      (2) Preservation of Statistical Learning in HL Infants:

      The text added on Pages 17-18 (Lines 562-566) regarding a "heightened dependence on bottom-up mechanisms (in autism)" does not appear to be supported by the data or by theories of implicit statistical learning. Because greater syllable-level entrainment was observed in the LL group than the HL group across both the random and structured conditions, the data actually point toward impaired bottom-up processes. Furthermore, implicit statistical learning typically involves an interplay of both bottom-up and top-down mechanisms; the implicit nature of a task does not guarantee a strictly bottom-up process. Consequently, this interpretation is not entirely convincing.

      We agree with the reviewer that the concepts of “top-down” and “bottom-up” were not fully appropriate to support our point in the cited paragraph. We should have used the concepts of implicit versus explicit learning instead, in line with previous literature suggesting increased reliance on preserved implicit learning in autism to compensate for altered explicit processes. The paragraph was slightly modified.

      Page 18, line 564: “According to these studies, autistic impairments in explicit attentional processes, such as social orienting - which are critical for bootstrapping language acquisition (70) - may result in a heightened dependence on implicit mechanisms, including statistical learning. As previously discussed, preserved word segmentation abilities may further compensate for alterations in lower-level implicit processes, such as syllable tracking.”

      Reviewer #2 (Recommendations for the authors):

      Potential typo on line 199 - I think an apostrophe is needed here.<br /> Potential typo on line 255 - do you mean Central electrodes?

      We addressed the typos spotted by reviewer.

      Line 199: variables’

      Line 255: Centro-frontal electrodes


      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports a prospective longitudinal study examining whether infants with high likelihood (HL) for autism differ from low-likelihood (LL) infants in two levels of word learning: brain-to-speech cortical entrainment and implicit word segmentation. The authors report reduced syllable tracking and post-learning word recognition in the HL group relative to the LL group. Importantly, both the syllable-tracking entrainment measure and the word recognition ERP measure are positively associated with verbal outcomes at 18-20 months, as indexed by the Mullen Verbal Developmental Quotient. Overall, I found this to be a thoughtfully designed and carefully executed study that tackles a difficult and important set of questions. With some clarifications and modest additional analyses or discussion on the points below, the manuscript has strong potential to make a substantial contribution to the literature on early language development and autism.

      Strengths:

      This is an important study that addresses a central question in developmental cognitive neuroscience: what mechanisms underlie variability in language learning, and what are the early neural correlates of these individual differences? While language development has a relatively well-defined sensitive period in typical development, the mechanisms of variability - particularly in the context of neurodevelopmental conditions - remain poorly understood, in part because longitudinal work in very young infants and toddlers is rare. The present study makes a valuable contribution by directly targeting this gap and by grounding the work in a strong theoretical tradition on statistical learning as a foundational mechanism for early language acquisition.

      I especially appreciate the authors' meticulous approach to data quality and their clear, transparent description of the methods. The choice of partial least squares correlation (PLS-c) is well motivated, given the multidimensional nature of the data and collinearity among variables, and the manuscript does a commendable job explaining this technique to readers who may be less familiar with it.

      The results reveal interesting developmental changes in syllable tracking and word segmentation from birth to 2 years in both HL and LL infants. Simply mapping these trajectories in both groups is highly valuable. Moreover, the associations between neural indices of brain-to-speech entrainment and word segmentation with later verbal outcomes in the LL group support a critical role for speech perception and statistical learning in early language development, with clear implications for understanding autism. Overall, this is a rich dataset with substantial potential to inform theory.

      Weaknesses:

      (1) Clarifying longitudinal vs. concurrent associations

      Because the current analytical approach incorporates all time points, including the final visit, it is challenging to determine to what extent the brain-language associations are driven by longitudinal relationships vs. concurrent correlations at the last time point. This does not undermine the main findings, but clarifying this issue could significantly enhance the impact of the individual-differences results. If feasible, the authors might consider (a) showing that a model excluding the final visit still predicts verbal outcomes at the last visit in a similar way, or (b) more explicitly acknowledging in the discussion that the observed associations may be partly or largely driven by concurrent correlations. Either approach would help readers interpret the strength and nature of the longitudinal claims.

      We thank the reviewer for this insightful comment. We agree that distinguishing between longitudinal predictive power and concurrent correlations at the final visit is crucial for clarifying the nature of these brain-language associations. Following the reviewer’s suggestion (a), we re-ran the two critical Partial Least Squares Correlation (PLS-c) analyses by excluding all EEG and behavioral data from the final 18–21 month visit (n = 54 recordings kept) to test whether earlier trajectories still predict the final verbal outcome.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2C–D): The PLS-c restricted to the 3- to 15-month visits still identified a single significant component (p=.001, r=.56, 64.0% explained covariance, Figure 2 -figure supplement 3). Bootstrap ratios (BSR) were: contrast (low vs. high autism likelihood) 4.1; mean age −1.3; contrast*mean-age −2.1; delta-age 10.5; contrast*delta-age 2.0; age<sup>2</sup> −6.0; contrast*age<sup>2</sup> −5.8; and notably verbal outcome 7.1; contrast*verbal-outcome −6.3.

      The latent component and its spatial electrode configuration remain highly consistent with the original analysis (Figure 2C–D). This confirms that excluding the final visit preserves the model’s predictive validity: lower syllable entrainment correlates with poorer verbal outcomes at 18–21 months, particularly in the high-likelihood group.

      (2) Late evoked response to novel words (original analysis on Figure 6): The PLS-c analysis on the ERP late time window (1500–3000 ms), excluding the final visit, also revealed one significant component (p=.002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (part-word versus word) 14.5; mean age -5.2; contrast*mean-age 6.2; delta-age -0.8; contrast*delta-age 3.5; age<sup>2</sup> -0.8; contrast*age<sup>2</sup> -12.1; verbal-outcome 20.1; contrast*verbal-outcome -7.2. The latent component closely mirrored the original analysis (Figure 6), with frontal electrodes contributing negatively and posterior electrodes positively. Minor divergences in age-related parameter contributions were observed, likely due to the absence of 18-21 month timepoints, which previously contributed to the convex/concave shapes of the group age trajectories in figure 6B (left panel).

      Crucially, both models (with and without the final visit) positively predicted verbal outcomes (Figure 6B, right panel, and Figure 6 -figure supplement 1B). However, excluding the final visit reversed the direction of the group*verbal-outcome interaction (from 3 to -7.2): This indicates that after ruling out cross-sectional correlations at 18–21 months, the early predictive value of the late ERP to word novelty is more prominently observed in high-likelihood infants, suggesting that the original result was influenced by concurrent cross-sectional correlations at the final visit. This aligns with the syllable entrainment findings (Figure 2 -figure supplement 3), as both 4 Hz neural tracking and late ERP responses to novelty predominantly predict verbal outcomes in infants at high likelihood for autism.

      We reported these supplementary analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 3 and Figure 6 -figure supplement 1. In general, most of figures that were present in Supplementary materials were moved as figure supplements to enhance readability.

      Page 8 lines 234-240 (pages and lines refer to the reviewed uploaded manuscript): “To rule out the possibility that the association between syllable entrainment and verbal outcome was driven by concurrent measures taken at 18–21 months, we re-ran the PLS-c analysis excluding EEG data from the final visit (n = 54 recordings kept). The resulting latent component remain significant (p = .001) and showed contributions from behavioral and EEG variables that were highly similar to those observed in the previous analysis, with a verbal outcome BSR of 7.1 and a group’verbal-outcome interaction BSR of −6.3 (Figure 2 -figure supplement 3).”

      Page 12 lines 387-394: “As we did for neural entrainment to syllables, we conducted a new analysis on late ERP to word novelty, excluding EEG data from the final visit. This PLS-c yielded one significant latent component (p = .002, r=.74, 33.5% explained covariance, Figure 6 -figure supplement 1) with globally similar EEG parameter contributions and age trajectory modelling. Verbal outcome still significantly contributed to the latent component (BSR=20.1), with a negative verbal outcome*group interaction (BSR=-7.2). These results suggest that, after ruling out cross-sectional correlations at 18–21 months, the late ERP to word novelty predominantly predicts verbal outcomes in high-likelihood infants for autism.”

      Page 17 lines 547-548: “As with syllable entrainment, the late ERP to novel words primarily predicted verbal outcomes in high-likelihood (HL) infants.”

      Page 18 lines 588-590: “Likewise, the absence of a late ERP orientation response in HL participants may represent an early neural signature of altered attention to novelty that can be used both as a non-invasive predictor of language development and as a potential target for early intervention.’

      (2) Incorporating sleep status into longitudinal models

      Sleep status changes systematically across developmental stages in this cohort. Given that some of the papers cited to justify the paradigm also note limitations in speech entrainment and word segmentation during sleep or in patients with impaired consciousness, it would be helpful to account for sleep more directly. Including sleep status as a factor or covariate in the longitudinal models, or at least elaborating more fully on its potential role and limitations, would further strengthen the conclusions and reassure readers that these effects are not primarily driven by differences in sleep-wake state.

      The reviewer is highlighting here a limitation of our study design that comprised sleeping status that varied from one timepoint to another among participants. To rule out any confounding effect of wake status (coded as a binary variable: sleeping or awake during recording) on analyses comparing groups, a linear mixed-effect model with repeated measures was fitted finding no significant difference between high- and low-likelihood participants (p=.769, reported at page 20, lines 646-647). However, as rightly suggested by the reviewer, this doesn’t prevent from a sleep bias on age trajectories, especially given that sleeping status significantly decreases with age in our sample.

      Including sleep status as a covariate in our analyses, as suggested by the reviewer, would be difficult to implement in our PLS-c methods, since a categorical behavioral parameter that varies within participants is not possible in the models provided by myPLS toolbox.

      As an alternative option, we re-ran all analyses that explored the condition effect on the whole sample within the sleeping participants only (n=25 recordings) to confirm that the same age-trajectories of EEG parameters were highlighted. However, negative results should be interpreted with caution since the sample is small for such a multivariate approach, resulting in modest statistical power.

      (1) Syllable entrainment (4 Hz) (original analysis on Figure 2A–B): The PLS-c identified one significant component (p <.001, r = .78, 85.1% explained covariance, Figure 2 -figure supplement 2 and Figure 3 -figure supplement 1). Bootstrap ratios (BSR) were: contrast (4hz vs. adjacent frequencies) 30.3; mean age -2.9; contrast*mean-age -2.4; delta-age 3.8; contrast*delta-age 3.4; age<sup>2</sup> -1.1; contrast* −2.5. The spatial distribution of contributing electrodes globally matched that shown in Figure 2A. The high contrast BSR (30.3) confirms robust syllable entrainment in sleeping infants. Critically, the contrast*age<sup>2</sup> parameter contributed negatively to the latent component (BSR = −2.5), confirming that the convex age trajectory of syllabic entrainment (Figure 2B) is also present in the sleeping subsample.

      (2) Word entrainment (1.3 Hz) (original analysis on Figure 3A–B): The PLS-c identified one significant component (p <.001, r = .63, 37.3% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (1.3hz vs. adjacent frequencies) 24.9; mean age -7.5; contrast*mean-age -3.2; delta-age 2.9; contrast*delta-age 0.0; age<sup>2</sup> 1.5; contrast* 1.1. The spatial distribution of significant electrodes partially overlaps with the ones in the original analysis, primarily showing fronto-central positive contribution to the latent component. The high contrast BSR confirms a robust word entrainment in sleeping participants, in line with previous studies (e.g., Flò et al, Sci Rep, 2022). However, the lack of a significant contrast* age<sup>2</sup> suggests that the U-shape age trajectory illustrated on Figure 3 might be modulated by wakefulness or due to a lack of power in the present analysis. A non-significant trend towards a U-shape pattern with a 12-month nadir is visible in sleeping participants, but additional data from sleeping 18-21 months sleeping infants would be required to confirm or refute this trend.

      (3) Early evoked response to novel words (original analysis on Figure 4): The PLS-c analysis on the ERP early time window (0–1000 ms) in sleeping participants revealed no significant component. The absence of early response to word novelty in sleeping participant might account for the lack of response observed in the whole sample, illustrated on Figure 4. To test this hypothesis, we conducted the same PLS-c in awake participants (n=58 recordings), which also yielded no significant latent component. This suggests that the lack of a measurable early response to word novelty observed in the whole sample is consistent across both sleeping and awake infants, and not driven by any of the two subsamples.

      (4) Late evoked response to novel words (original analysis on Figure 5): The PLS-c analysis on the ERP late time window in sleeping participants revealed no significant component. This suggests that sleeping participants might present a reduced or even absent late response to novel words. Given this identified effect of sleep on late ERP response, we reran the PLS-c on the late ERP window using group as contrast (original analysis on figure 6), excluding the sleeping participants to avoid any confounds. This PLS-c revealed one significant component (p = .006, r = .68, 30.5% explained covariance, Author response image 1). Bootstrap ratios (BSR) were: contrast (low versus high likelihood) 7.6; mean age -5.9; contrast*mean-age -3.8; delta-age 2.5; contrast*delta-age -3.5; age<sup>2</sup> -2.5; contrast*age<sup>2</sup> -4.9; verbal-outcome 8.5; contrast*verbal-outcome -1.0. Behavioral parameters contribute to this latent component with similar magnitude and polarity as in the original analysis. Electrode contributions are also highly consistent, with frontal negative and posterior positive contributions. This confirms that sleeping participants, despite their potentially reduced late response, did not significantly bias the results presented in Figure 6.

      Author response image 1.

      Late evoked response potential (ERP) to word novelty in awake participants. A. Design and brain saliences derived from the significant latent component. Brain topographies of bootstrap ratios (BSR) are displayed at 250ms intervals. Black dots indicate BSR > 2.3. B. Participants’ brain scores for part-word and word conditions, as a function of age (left panel) and verbal DQ (right panel). For details on brain scores, see Figure 6 -figure supplement 1. Linear fitting is used for illustrative purposes only. HL: high likelihood for autism; LL: low likelihood for autism.

      We reported these analyses in the revised manuscript as follows:

      We added Figure 2 -figure supplement 2A and Figure 3 -figure supplement 1.

      Page 7, lines 218-224: “Because some infants were asleep during the recording session, particularly at younger ages, we performed a supplementary control analysis restricted to this sleeping subsample (n = 25 recordings, Figure 2 -figure supplement 2). This PLS-c also identified a significant latent component (p < .001, r = .78, 85.1% explained covariance), with a significant contrast effect (BSR = 30.3) and a significant negative contrast*age<sup>2</sup> interaction (BSR = −2.5). These findings confirm that the convex age trajectory observed in the main analysis remains present and observable even in sleeping infants.”

      Page 8, lines 252-259: “We further investigated word entrainment in sleeping participants (n=25), which yielded one significant latent component (p<.001, r=.63, 37.3% explained covariance, Figure 3 -figure supplement 1). Centro-frontal electrode contributed to this component, with a high contrast BSR (24.9), confirming a similar word entrainment pattern in the sleeping subsample. The contrast*age<sup>2</sup> was also positive but not significant (1.1), suggesting a trend toward a U-shape age trajectory with a 12-month nadir in sleeping infants. Additional 18-21 month recording would be required to confirm this trend.”

      Page 11, lines 347-349: “The same PLS-c, conducted separately in sleeping (n=25) and awake subsamples (n = 58), yielded no significant latent component, indicating a consistent absence of early response to word novelty in both sleeping and awake infants.”

      Page 11-12 lines 368-370: “The same PLS-c in the sleeping subsample yielded no significant latent component, suggesting that sleep may reduce or even abolish the late response to word novelty.”

      Page 12 lines 382-385: “Given that no late response was detected in sleeping participants, we re-ran the PLS-c analysis using group as a contrast in the awake subsample (n=58). This yielded one significant latent component (p=.006, r=.68, 30.5% explained covariance), with behavioral and electrode contributions highly overlapping with those in Figure 6.”

      Page 18 lines 590-592: “This potential biomarker might nevertheless be modulated by participants’ sleep status, warranting careful consideration of vigilance state in future studies.”

      (3) Use of PLS-c and potential group × condition interactions

      I am relatively new to PLS-c. One question that arose is whether PLS-c could be extended to handle a two-way interaction between group and condition contrasts (STR vs. RND). If so, some of the more complex supplementary models testing developmental trajectories within each group (Page 8, Lines 258-265) might be more directly captured within a single, unified framework. Even a brief comment in the methods or discussion about the feasibility (or limitations) of modeling such interactions within PLS-c would be informative for readers and could streamline the analytic narrative.

      The reviewer raises a valid concern regarding the capacity of PLS-c to accommodate multi-way interactions among categorical and continuous variables. While PLS-c has no inherent theoretical constraints on the number of predictor terms (they can even exceed the sample size in number), practical limitations arise from model stability and interpretability when the ratio of predictors to sample size becomes excessive. As noted by Geladi and Kowalski (1986), exceeding ~10% of the sample size with predictors increases noise sensitivity and overfitting.

      In our study, the PLS-c analyses already reach this ~10% limit, with a maximum of nine predictors for a sample size of n=83. Attempting to integrate both group and condition as contrasts — along with necessary age parameters to account for developmental trajectories — would result in 12 predictors (or 15 if verbal outcome is included). Specifically, the model would require behavioral terms for Group, Condition, Group*Condition, Mean-age, Group*Mean-age, Condition*Mean-age, Delta-age, Group*Delta-age, Condition*Delta-age, Age<sup>2</sup>, Group*Age<sup>2</sup>, Condition*Age<sup>2</sup>, Verbal-outcome, Group*Verbal-outcome, and Condition*Verbal-outcome.

      Although a unified multivariate model capturing the complex dynamics at play in our sample is theoretically appealing, the substantial risk of overfitting precludes its feasibility. Therefore, we opted to use only one categorical predictor per PLS-c analysis to maintain model parsimony and reliability. However, a larger sample could overcome this limitation, allowing a stable and unified model of longitudinal EEG data that simultaneously captures age trajectories, group, clinical outcome, and condition.

      Reference:

      Geladi, P., & Kowalski, B. (1986). Partial least-squares regression: A tutorial. Analytica Chimica Acta, 185, 1–17. https://doi.org/10.1016/S0003-2670(00)82582-3

      We added the following comment in the method section:

      Page 24, lines 773-777: “We limited the number of behavioral variables to nine to mitigate noise sensitivity and overfitting risks associated with exceeding the 10% sample size threshold (Geladi & Kowalski, 1986). This limitation precluded the implementation of a single PLS-c model incorporating group, condition (STR vs. RND), age, and their interactions.”

      (4) STR-only analyses and the role of RND

      Page 8, Lines 241-245: This analysis is conducted only within the STR condition. The lack of group difference observed here appears consistent with the lack of group difference in word-level entrainment (Page 9, Lines 292-294), suggesting that HL and LL groups may not differ in statistical learning per se, but rather in syllabic-level entrainment. As a useful sanity check and potential extension, it might be informative to explore whether syllable-level entrainment in the RND condition differs between groups to a similar extent as in Figure 2C-D. In other work (e.g., adults vs. children; Moreau et al., 2022), group differences can be more pronounced for syllable-level than for word-level entrainment. Figure S6 seems to hint that a similar pattern may exist here. If feasible, including or briefly reporting such an analysis could help clarify the asymmetry between the two learning measures and further support the interpretation of syllabic-level differences.

      The reviewer points to the interesting pattern highlighted in supplementary figure S6, suggesting that group differences in syllabic entrainment might be modulated by the structure of the stream (STR versus RND). Such modulatory effect of stream structure on entrainment to syllables has been suggested by many studies, like Moreau et al (2022), as pointed by the reviewer, and seems at play in our sample, as illustrated on supplementary figure S5 (decline in the 4hz PLV that exceeds the size of confidence intervals, ~90 s after STR onset).

      Following the reviewer’s suggestion, we ran a PLS-c testing group effect on 4hz PLVs in each stream:

      (1) in the RND stream: the analysis yields one significant component (p<.001, r=.49, 52.7% explained covariance, Author response image 2A-B). Bootstrap ratios (BSR) are: contrast (low versus high likelihood) 6.9; mean age -1.0; contrast*mean-age 0.1; delta-age 11.8; contrast*delta-age 0.1; age<sup>2</sup> -4.9; contrast*age<sup>2</sup> 4.6; verbal-outcome 9.7; contrast*verbal-outcome -0.6. Interestingly, the model still highlights a strong link between syllable tracking and group, suggesting that RND also discriminate between HL and LL. However, RND syllable tracking doesn’t appear to be linked to group x verbal-outcome as we observed in Figure 2C-D.

      (2) In the STR stream, we obtained one significant latent component (p=.002, r=.51, 57.8% explained covariance, Author response image 2C-D). Bootstrap ratios (BSR) are: contrast 2.2; mean age -1.6; contrast*mean-age -1.1; delta-age 6.1; contrast*delta-age -0.7; age<sup>2</sup> -4.4; contrast*age<sup>2</sup> 0.5; verbal-outcome 8.2; contrast*verbal-outcome -7.3. Here, the strong association between syllable tracking and group x verbal-outcome is similar to the model presented in Figure 2C-D.

      Taken together, these results suggest that the apparent STR/RND dissociation illustrated in Figure S6 might primarily reflect a Group*Verbal-outcome divergence, with syllable tracking in the STR stream being related to verbal outcome mainly in high likelihood for autism.

      Author response image 2.

      Syllable entrainment within RND (A-B) and STR (C-D).

      These results were reported in the revised manuscript in the Result section (Time course of the entrainment along experiment subheader), implying a slight reframing of the result presentation of supplementary analysis S6. Author response image 2 was added in supplementary material as Figure S5.

      Page 10, lines 307-319: “The group, age and verbal outcome parameters were mainly correlated (BSR>2.3) with the neural entrainment occurring~90 seconds after the onset of the STR stream, coinciding with the time participants began tracking word boundaries (Supplementary material, S3). This result suggests that the group differences in syllable entrainment, as shown in Figure 2C-D, as their associations with verbal outcome, are modulated by the structure of the stream (STR versus RND). We ran one additional PLS-c for each stream separately, using group as contrast. In both streams, the PLS-c yielded a significant LC (p<.001 for RND and p=.002 for STR), with a positive group effect (BSR>2.3) in both LC (Supplementary material, S5). Most strikingly, the group*verbal outcome parameter reached significance exclusively within the STR latent component (BSR:-7.3). These results suggest that while syllable tracking is generally decreased in HL infants across both streams, its association with verbal outcome is prominently driven by the stream containing words (STR).”

      Page 14, lines 443-446: “This temporal overlap suggests that successful segmentation may enhance syllable tracking via top-down predictions of the next syllable, improving alignment to syllable onsets in LL infants as well as in HL with better verbal outcome.’

      (5) Multi-speaker input and voice perception (Page 15, Lines 475-483)

      The multi-speaker nature of the speech input is an interesting and ecologically relevant feature of the design, but it does add interpretive complexity. The literature on voice perception in autism is still mixed: for example, Boucher et al. (2000) reported no differences in voice recognition and discrimination between children with autism and language-matched non-autistic peers, whereas behavioral work in autistic adults suggests atypical voice perception (e.g., Schelinski et al., 2016; Lin et al., 2015). I found the current interpretation in this paragraph somewhat difficult to follow, partly because the data do not directly test how HL and LL infants integrate or suppress voice information. I think the authors could strengthen this section by slightly softening and clarifying the claims.

      We acknowledge the reviewer’s concern regarding the potential ambiguity in the cited paragraph. To address this, we have revised the text to explicitly clarify the aims of our study and its design. Furthermore, we now emphasize the speculative and post-hoc nature of the hypotheses and interpretations presented, thereby ensuring transparency regarding the limitations of our findings.

      Page 16 lines 520-530), as follows: “HL infants, on the other hand, did not show this transient disruption. In this group, word entrainment remained stable over time. To account for this unexpected finding, we followed up on the post-hoc hypothesis proposed above: a reduced sensitivity to social and vocal cues observed in HL infants may have spared segmentation abilities by limiting the interference introduced by speaker variability. If this post-hoc hypothesis holds true, LL and HL infants would differ not in their intrinsic ability to learn statistical regularities per se, but rather in how they integrate or suppress competing cues (such as speaker changes) during the segmentation process. It is important to note, however, that the present study was not designed to isolate and evaluate the specific impact of speaker changes on word segmentation. Consequently, this interpretation remains speculative, and additional research is required to further address this question.”

      (6) Asymmetry between EEG learning measures

      Page 16, Lines 502-507 touches on the asymmetry between the two EEG learning measures but leaves some questions for the reader. The presence of word recognition ERPs in the LL group suggests that a failure to suppress voice information during learning did not prevent successful word learning. At the same time, there is an interesting complementary pattern in the HL group, who show LL-like word-level entrainment but does not exhibit robust word recognition. Explicitly discussing this asymmetry - why HL infants might show relatively preserved word-level entrainment yet reduced word recognition ERPs, whereas LL infants show both - would enrich the theoretical contribution of the manuscript.

      We concur with the reviewer’s observation that our findings imply a theoretically significant double dissociation between HL and LL groups, specifically concerning the asymmetries between word-level neural entrainment and word recognition mechanisms. We believe this point was partly addressed in the subsequent paragraph, where we stated that “in contrast” to LL, HL infants “showed no clear ERP difference between novel and familiar triplets”, while “both groups showed similar word neural entrainment during learning”. We further explored potential explanations for this apparent dissociation, such as a possible deficit in novelty orientation that may be specific to HL infants and unrelated to statistical learning itself. We cited Liu et al (2023) as a reference showing the dissociation between mechanisms underlying implicit versus explicit traces of statistical learning. We acknowledge that we can discuss more in depth the potential preservation of statistical learning in HL infants. We have incorporated the following discussion in the reviewed manuscript, supported by relevant references:

      Pages 17-18, lines 562-566: “Interestingly, this dissociation between spared implicit versus impaired explicit statistical learning in autism has been previously discussed in the literature (Zwart et al, 2018, Kissine, 2021). According to these studies, autistic impairments in top-down attentional processes, such as social orienting — which are critical for bootstrapping language acquisition (Kuhl, 2007) — may result in a heightened dependence on bottom-up mechanisms, including implicit statistical learning.”

      References:

      Zwart, F.S., Vissers, C.T.W.M., Kessels, R.P.C. and Maes, J.H.R. (2018), Implicit learning seems to come naturally for children with autism, but not for children with specific language impairment: Evidence from behavioral and ERP data. Autism Research, 11: 1050-1061. https://doi.org/10.1002/aur.1954

      Kissine, M. (2021). Autism, constructionism, and nativism. Language 97(3), e139-e160. https://dx.doi.org/10.1353/lan.2021.0055.

      Kuhl, P.K. (2007), Is speech learning ‘gated’ by the social brain?. Developmental Science, 10: 110-120. https://doi.org/10.1111/j.1467-7687.2007.00572.x

      References:

      (1) Moreau, C. N., Joanisse, M. F., Mulgrew, J., & Batterink, L. J. (2022). No statistical learning advantage in children over adults: Evidence from behaviour and neural entrainment. Developmental Cognitive Neuroscience, 57, 101154. https://doi.org/10.1016/j.dcn.2022.101154

      (2) Boucher, J., Lewis, V., & Collis, G. M. (2000). Voice processing abilities in children with autism, children with specific language impairments, and young typically developing children. Journal of Child Psychology and Psychiatry, 41(7), 847-857. https://doi.org/10.1111/1469-7610.00672

      (3) Schelinski, S., Borowiak, K., & von Kriegstein, K. (2016). Temporal voice areas exist in autism spectrum disorder but are dysfunctional for voice identity recognition. Social Cognitive and Affective Neuroscience, 11(11), 1812-1822. https://doi.org/10.1093/scan/nsw089

      (4) Lin, I.-F., Yamada, T., Komine, Y., Kato, N., Kato, M., & Kashino, M. (2015). Vocal identity recognition in autism spectrum disorder. PLOS ONE, 10(6), e0129451.https://doi.org/10.1371/journal.pone.0129451

      Reviewer #2 (Public review):

      Summary:

      This article looks at differences in how the brain entrains to, or tracks, the rhythmic presentation of syllables and words in speech in infants at increased likelihood versus low likelihood for autism. The authors first sought to characterize how brain responses are modulated by learning the statistical probability of a given syllable following the one before it over the first two years of life. They then sought to identify at which stages of word learning infants with increased likelihood of autism showed difficulties, and whether those difficulties worsened over time. Finally, they sought to indicate whether infants' statistical learning and word learning abilities could predict later verbal skills. The authors found similar developmental trajectories of neural entrainment to syllables in infants at high and low likelihood for autism, but infants at high likelihood for autism had overall weaker syllable-level entrainment. Infants at high versus low likelihood for autism showed different developmental trajectories for word entrainment. Lower syllable entrainment in high-likelihood infants corresponded with poorer verbal outcomes, but word entrainment was not associated with verbal outcomes. Event-related potential responses to words and part words were positively associated with verbal outcomes, however, but only in low-likelihood infants.

      Strengths:

      Overall, the article provides rigorous statistical analysis of longitudinal EEG data to provide strong support for the claims that neural entrainment to syllable and word features of speech may be a useful marker for language development difficulties, particularly in infants at increased likelihood for neurodevelopmental disorders. The EEG data collection and preprocessing procedures are well within standards in the field. Readers should take care to note that authors indexed neural entrainment to speech using phase-locking values instead of spectral power.

      Weaknesses:

      While the statistical analyses are rigorous, a few of the components of the models are not clearly defined, and some corrections and thresholds for significance warrant further justification. Further, a few stimuli and participant details that could influence results are not specified. It is not clear whether all participants came from majority French-speaking families; differences in the amount of French language exposure (compared to other languages that may be spoken by a participant's family) could influence results. The standardized volume of the stimuli is also not included. As a result, readers should be encouraged to interpret that neural entrainment to speech features is likely a useful mechanism to explain differences in language development, while taking this interpretation with some caution.

      We thank the reviewer for these remarks.

      Regarding the amount of French exposure: while all participants were raised in primarily French-speaking environments (i.e., French as the dominant language at home and daycare), the parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. We did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism. The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      Regarding the volume of stimuli, they were played at 50cm distance with an intensity of 75dB. Both considerations have been included in the new version of the manuscript. In general, we moved most of the figures present in Supplementary material to figure supplements to improve readability.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor Comments:

      Figure 6: The figure caption is not complete (there is no description for the right half of panel B).

      We thank the reviewer for this observation, figure 6 caption has been completed.

      Reviewer #2 (Recommendations for the authors):

      Broadly speaking, I would recommend reducing the number of abbreviations in this article, and I would recommend that the authors take care as to where these abbreviations are being introduced. Many of the abbreviated terms are defined in the Materials and Methods section, which is presented after the abbreviations are used in the main results.

      We acknowledge that our extensive use of abbreviations compromises the readability of the manuscript. Consequently, we have removed the following abbreviations:

      - SL (replaced by statistical learning)

      - LC (replaced by latent component)

      - ASD (replaced by autism)

      - TP (replaced by transition probability)

      - MEG (replaced by magneto-encephalogram)

      - MSEL (replaced by Mullen Scale of Early Learnings)

      The remaining abbreviations are:

      HL (high likelihood for autism), LL (low likelihood for autism), EEG (electroencephalogram), PLS-c (partial least square correlation), ERP (event-related potential), RND (random), STR (structured), BSR (bootstrap ratio), AIC (Akaike Information Criterion), PLV (phase locking value), DQ (developmental quotient), APSI (Autism Parent Screen for Infants).

      Moreover, we carefully reviewed how abbreviations were introduced and identified that PLS-c, STR and RND were not defined prior to the Method section. This oversight has been corrected in the reviewed manuscript.

      I would also recommend that the authors be careful with the structuring of the Introduction, particularly with their research questions and hypotheses. The article initially makes clear that the research questions are focused on the developmental trajectory of statistical learning, the levels of word learning that may differentiate high-likelihood versus low-likelihood infants, and the stability of those differences, and associations between statistical learning and various levels of word learning with verbal outcomes. The use of acoustic variability across syllables, while a valuable methodological tool, is somewhat presented as an additional research question, but not clearly stated or tested as such.

      We acknowledge that the introduction (particularly the paragraph from lines 173 to 184) may have implied that speaker variability across syllables was one of our primary research aims. We clarify here that speaker variability was introduced as a mean to increase task difficulty, particularly for high-likelihood (HL) participants, with the aim of amplifying the effect sizes in our analyses.

      To address this, we have removed the theoretical discussion on speaker variability in autism and typical development (lines 173–184) and explicitly stated that speaker variability was not a research question in this study. Crucially, our experimental design did not include a control condition without speaker variability, and thus we could not test its specific effects on statistical learning across age trajectories and groups.

      Page 6, lines 173-176 (pages and lines refer to the reviewed uploaded manuscript): “It is worth noting, however, that our study was not designed to isolate or quantify the specific impact of speaker variability on statistical learning, as the experimental design did not include a baseline control condition omitting this acoustic variation.”

      The authors do a nice job in the Materials & Methods explaining PLS-c and defining the latent components and bootstrapped ratios that will be shared in the Results. An additional brief iteration defining these statistical elements is needed at the beginning of the Results section.

      We thank the reviewer for their appreciation of our Method section. We agree that an additional iteration in the result section would improve readability. We added the following paragraph at the very beginning of the Result section, briefly defining PLS-c and its main statistical output (latent components and bootstrap ratios):

      Pages 6-7, lines 193-202: “Briefly, PLS-c is a data-driven multivariate modelling approach designed to identify significant patterns of electrode clusters (from a brain data matrix containing electrophysiological measures, here PLV) and their associations with “behavioral” variables (from a behavioral design matrix, here age-related parameters). Patterns of brain x behavior associations are called latent components, and their statistical significance is evaluated using permutation testing (n=1000, Bonferroni correction for number of components tested, alpha=.006). Brain and behavioral variables respective contributions to any significant latent component are tested with bootstrapping (500 random samples and replacement), with bootstrap ratios (BSR) greater than 2.3 indicating a stable contribution (for details, see the Materials and Methods section).”

      (1) Page 18 Line 576. The authors need to clarify whether participants were required to be in primarily French-speaking environments and whether there was a minimum amount of French language exposure that participants were required to have if they were exposed to additional languages besides French in their everyday life.

      The reviewer raises a valid concern regarding participants’ language exposure. In this study, all participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare. The parental questionnaire at intake indicated that 45% of the sample was exposed to additional languages, reflecting Geneva’s highly multicultural demographics. However, we did not quantify the extent of this exposure, which could range from very occasional exposure to situations close to true bilingualism.

      The structural sensitivity hypothesis (Weiss et al., 2020) posits that additional language exposure may enhance detection of statistical structures in artificial language input, even when these structures differ from those in native languages. Yet, empirical support is mixed: Yim & Rudoy (2013) found no bilingualism effect in a paradigm close to ours (triplet segmentation via auditory statistical learning, n=112 children), whereas most studies reporting bilingual advantages for statistical learning involved tasks very distinct from ours, like artificial grammar and phonotactic rule learning, or multi-cue integration for segmentation (Weiss et al., 2020).

      To include these considerations, Limitations and Material and methods sections were modified as follows:

      Page 19, lines 604-607: “Second, although all participants were primarily exposed to French, we did not quantify additional language exposure, precluding any analysis of its potential moderator effects on statistical learning in our groups and age-trajectories. However, prior work has reported no effect of bilingualism on auditory triplet segmentation in children (Yim & Rudoy, 2013).”

      Page 20, lines 630-631: “All participants were raised in primarily French-speaking environments, with French as the dominant language at home and daycare.”

      References:

      Weiss DJ, Schwob N, Lebkuecher AL. Bilingualism and statistical learning: Lessons from studies using artificial languages. Bilingualism: Language and Cognition. 2020;23(1):92-97. doi:10.1017/S1366728919000579

      Yim D, Rudoy J. Implicit statistical learning and language skills in bilingual children. J Speech Lang Hear Res. 2013 Feb;56(1):310-22. doi: 10.1044/1092-4388(2012/11-0243). Epub 2012 Aug 15. PMID: 22896046.

      (2) Page 18 Line 588. Further, the authors should clarify whether the 7 infants in the HL group, due to early parental concerns were defined by the 18-21-month APSI scores or by parental report prior to study enrollment.

      These 7 infants were recruited based on early parental concerns prior to intake. The APSI score at 18-21 months is only reported to provide an illustration of the amount of early autistic signs that were present in these 7 infants, and to provide an estimation of their probability to develop autism later on based on Sacrey et al., 2018 longitudinal study on the APSI predictive value. We agree with the reviewer that our phrasing suggests that the APSI was used as an inclusion criterion. We rephrased the page 20 lines 642-646 as follows:

      “The 7 other HL infants presented with early parental concerns for autism, based on parental report prior to enrollment. Their Autism Parent Screen for Infants (APSI) total score at their 18-21 months visit was 15.6±6.4, [8-22] range – a score greater than 8 reflecting a 63% positive predictive value for autism in HL populations.”

      (3) Page 20 Line 641. The authors should specify the volume of the stimuli.

      The volume of stimuli was reported in the main text (page 22, lines 695-696) as follows:

      “Stimuli were played on a Bose® Companion 2 Series III at a 50cm distance with an intensity of 75dB.”

      (4) I'd prefer Figure 1 to be reorganized slightly - at present, the placement of the arrows explaining the analysis steps is not intuitive.

      We addressed the reviewer’s comments (4) and (5) together as they both refer to Figure 1B.

      (5) Page 23 lines 718-719. I think it would be helpful to explicitly define each of the interaction variables included in the behavioral design matrix. Further, this matrix should be labeled consistently in both Figure 1B and in the main text.

      We refined figure 1B and its corresponding main text (in Methods section) for clarity. The arrows are now simpler and more parsimonious, labels (e.g., participant i, visit n, behavior design matrix and its parameters) are now standardized between the figure and the main text, and the interaction terms at lines 718-719 are explicitly defined.

      (6) Page 23 lines 726-731: It would be helpful to know whether applying a Bonferroni correction in addition to completing permutation testing is standard when evaluating latent components derived from PLS-c. The authors should also cite justification for a bootstrap ratio cutoff of 2.3 for defining stability.

      In PLS-c analyses, multiple comparisons correction across latent components and bootstrap ratio (BSR) thresholding at 2.3 are commonly adopted practices.

      - Correction for multiple comparisons in PLS-c: PLS-c performs singular decomposition of the data into latent components equal in number to the variables included in the behavior design matrix (7-9 in our study, depending on the inclusion of Verbal outcome as an input variable). Each latent component’s statistical significance is assessed through permutation testing, generating a null distribution for its singular value (Krishnan et al., 2011). Given the multiple tests (one permutation test per latent component), Type I error inflation must be addressed. Recent PLS-c studies commonly applied Bonferroni correction (default procedure in the myPLS toolbox, used by Zoeller et al., 2017, and Delavari et al, 2021), though FDR correction has also been used (Lombardo et al, 2018).

      - Stability threshold for bootstrap and replacement: Within each latent component, saliences’ stability (brain/behavior parameter contributions to each latent component) are evaluated using bootstrapping (Krishnan et al., 2011). The bootstrap ratio (BSR) of each parameter, calculated as the saliency divided by its bootstrap-derived standard error, functions analogously to a z-score under normality assumptions. The BSR can then be used to assess the stability of the saliency (i.e., how stable is its contribution to the latent component). BSR thresholds in the literature typically range from 1.96 to 3.0. Krishnan et al (2011) state that when BSR are “larger than 2 the corresponding saliences are considered significantly stable”. Delavari et al (2021) and our study used a 2.3 thresholding, corresponding to a 99.0% bootstrap confidence interval not crossing the zero line – roughly equivalent to a two-tailed p<.001. Lombardo et al (2018) used a looser threshold of 1.96, corresponding to a 95% confidence interval not crossing the zero line (~two-tailed p<.05), while Zöller et al (2017) used a more stringent 3.0 thresholding (~p<.001, or 99.9% confidence interval not crossing the zero line).

      Thus, our application of Bonferroni correction for multiple comparisons and our 2.3 BSR threshold aligns with established conventions.

      We added following lines in the manuscript:

      Page 25 lines 784-785: “Bonferroni correction was applied to account for multiple comparisons across the 9 tested latent components in the PLS-c, yielding an adjusted alpha of .006 (Zoeller et al, 2017; Delavari et al, 2021).”

      Page 25 lines 789-792: “BSR are analogous to Z-scores and can be used to assess the stability of the saliency. We considered BSR > 2.3 as stable, corresponding to a 99.0% bootstrap confidence interval not crossing zero – roughly equivalent to a two-tailed p<.001 (Delavari et al., 2021; Krishnan et al., 2011).”

      References:

      Delavari F, Sandini C, Zöller D, Mancini V, Bortolin K, Schneider M, Van De Ville D, Eliez S. Dysmaturation Observed as Altered Hippocampal Functional Connectivity at Rest Is Associated With the Emergence of Positive Psychotic Symptoms in Patients With 22q11 Deletion Syndrome. Biol Psychiatry. 2021 Jul 1;90(1):58-68. doi: 10.1016/j.biopsych.2020.12.033. Epub 2021 Jan 18. PMID: 33771350.

      Lombardo, M.V., Pramparo, T., Gazestani, V. et al. Large-scale associations between the leukocyte transcriptome and BOLD responses to speech differ in autism early language outcome subtypes. Nat Neurosci 21, 1680–1688 (2018). https://doi.org/10.1038/s41593-018-0281-3

      Daniela Zöller, Marie Schaer, Elisa Scariati, Maria Carmela Padula, Stephan Eliez, Dimitri Van De Ville. Disentangling resting-state BOLD variability and PCC functional connectivity in 22q11.2 deletion syndrome. NeuroImage, Volume 149, 2017, Pages 85-97, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2017.01.064

      Anjali Krishnan, Lynne J. Williams, Anthony Randal McIntosh, Hervé Abdi, Partial Least Squares (PLS) methods for neuroimaging: A tutorial and review, NeuroImage, Volume 56, Issue 2, 2011, Pages 455-475, ISSN 1053-8119, https://doi.org/10.1016/j.neuroimage.2010.07.034

      (7) I have a few minor grammar/formatting recommendations for the authors as well:

      (a) Should the Geneva Autism Cohort be capitalized? At present, it is not.

      We agree with the reviewer’s suggestion, and we capitalized the Geneva Autism Cohort in the main text (page 18, line 571)

      (b) Page 24, line 750. Do the authors mean that the data was re-referenced to average?

      The preprocessed data is not average-referenced (see section Data pre-processing). Therefore, both for neural entrainment computation and ERPs, the data were average-referenced.

      (c) It would be nice to have a figure of the actual ERP for each condition and age group.

      We agree that PLS-c can be difficult to interpret without the raw actual ERPs on which it was modelled. We direct the reviewer to supplementary figure S6 at page 59, which displays the raw ERPs for each condition (part-word, word, and their subtraction) per age group. Supplementary figures S7-8 at pages 60-61 further illustrate topographical ERPs for each group (high and low likelihood for autism). We deemed these figures too extensive for the main text. Instead, the most relevant ERP topographies are presented in Figures 4-6 to facilitate PLS-c interpretation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript from the Levy lab, the authors investigate whether SETD6 regulates hepatic lipid accumulation through direct methylation of PPARγ. They show that SETD6 binds and monomethylates PPARγ at K170, and provide evidence that this modification enhances PPARγ occupancy at target promoters, promotes expression of lipid metabolism genes, as well as facilitates lipid droplet accumulation in HepG2 cells. The authors also find a positive feedback loop or circuit in which PPARγ activates SETD6 transcription in a methylation-dependent manner, thereby reinforcing this lipogenic program. Overall, the work presents a novel SETD6PPARγ regulatory axis linking lysine methylation to transcriptional control of lipid storage genes, with possible relevance to NAFLD-associated biology.

      In all, I find this to be an important paper that describes and advances a new regulatory pathway that has significance to human health and disease. It would also be of interest to a broad audience. That said, there are also some concerns that the authors should address, as outlined below.

      We are grateful to the reviewer for the positive feedback and appreciation of our work.

      Major concerns (pertains to rigor - highest priority)

      (1) Overall, the work presented is of high quality, and the data nicely support the conclusions; however, a few panels should be strengthened that have missing controls or information:

      (a) The co-IP panel in Figure 1B lacks a lane where HA SETD6 is expressed without PPARγ. This control is needed to verify that the SEDT6-HA signal depends on PPARγ.

      We thank the reviewer for this valuable suggestion. The overexpression co-immunoprecipitation experiment referred to by the reviewer has been moved to Supplementary Figure S1 in the revised manuscript. In this experiment, immunoprecipitation was performed using an anti-FLAG antibody to pull down FLAG-tagged PPARγ. In the absence of FLAG-PPARγ, the anti-FLAG immunoprecipitation does not recover a bait protein, and therefore HA-SETD6 is not expected to be specifically immunoprecipitated. Thus, an HA-SETD6-only condition would primarily serve as a negative control for the anti-FLAG pull-down rather than provide additional information regarding the specificity of the interaction.

      Importantly, in the revised manuscript we have substantially strengthened the evidence supporting the SETD6–PPARγ interaction by adding two independent complementary experiments. First, we included a reciprocal endogenous co-immunoprecipitation (new Figure 2B), demonstrating that endogenous PPARγ co-immunoprecipitates with endogenous SETD6. Second, we added an independent proximity ligation assay (PLA) (new Figure 2D), which further confirms the interaction between SETD6 and PPARγ in cells. Together with the in vitro binding assay presented in Figure 2A, these orthogonal approaches provide compelling evidence for the specificity of the SETD6–PPARγ interaction. Therefore, we believe that the requested HA-SETD6-only control would not provide additional mechanistic insight beyond the comprehensive validation now included in the revised manuscript.

      (b) In Figure 1C, the authors should show that the co-IP works in both directions (include IP for PPARγ/blot for SETD6). I am a bit confused also over the labeling with IP on the left and on top of the panel next to the beads label. More importantly, the data would be stronger if the authors took advantage of a deletion line to validate that the co-IP is specific to the presence of both.

      We thank the reviewer for this helpful suggestion. We have revised the manuscript to strengthen the evidence supporting the endogenous interaction between SETD6 and PPARγ. Specifically, we now include a reciprocal endogenous co-immunoprecipitation (new Figure 2B), demonstrating that endogenous SETD6 co-immunoprecipitates with endogenous PPARγ and, conversely, that endogenous PPARγ co-immunoprecipitates with endogenous SETD6. These reciprocal experiments independently validate the specificity of the interaction.

      In addition, we have revised the figure layout and labeling to more clearly distinguish the immunoprecipitating antibody from the bead control, thereby addressing the reviewer's concern regarding the presentation of the co-immunoprecipitation data.

      Although we did not perform the co-immunoprecipitation in a depletion/knockout background, we believe that the combination of reciprocal endogenous co-immunoprecipitation (New Figure 2B), the independent proximity ligation assay (New Figure 2D), and the direct in vitro binding assay (Figure 2A) provides multiple orthogonal lines of evidence supporting a specific interaction between SETD6 and PPARγ.

      (c) The same IP labeling issue exists for Figure 3B (label is on the same and on top).

      We have revised the labeling in Figure 3B to clearly distinguish the immunoprecipitating antibody from the bead control, thereby improving the clarity of the figure.

      (d) Antibody information (e.g., where the pan-methyl Ab comes from and at what dilutions they are used at) is missing.

      We thank the reviewer for pointing this out. We have now added the missing information regarding the pan-methyl antibody to the Materials and Methods section, including the supplier, catalogue number, and experimental conditions used. Specifically, the pan-methyl antibody used in this study was purchased from Abcam (ab23366) and was used at a 1:500 dilution for western blot analysis and 2 μg per reaction for immunoprecipitation experiments.

      Nice to have experiments (medium priority - strongly consider)

      (2) A missing gap is how K170me1 contributes to DNA binding and gene transcription. One possibility is that methylation enhances the DNA-binding activity of PPARγ. Given that the authors have all of the reagents, it would be possible to perform a gel shift assay (or other approach) with and without SETD6-mediated methylation. Is DNA binding affected/enhanced?

      We thank the reviewer for raising this important point. To investigate whether K170 methylation could directly affect PPARγ binding to DNA, we performed structural modeling based on the available co-crystal structure of PPARγ bound to DNA (PDB: 3DZU). As shown in the new Figure 6F, K170 is positioned near the DNA-binding region; however, the modeled K170me1 side chain is predicted to face away from the DNA interface and does not appear to sterically interfere with the PPARγ–DNA interaction. In addition, modeling of multiple K170me1 rotamers did not suggest any major disruption of the DNA-bound conformation.

      These observations suggest that K170 methylation is unlikely to directly alter the intrinsic DNA-binding affinity of PPARγ. In contrast, our ChIP-qPCR experiments demonstrate that K170 methylation positively regulates PPARγ occupancy at target promoters in cells. Together, these findings support a model in which K170 methylation promotes PPARγ chromatin association and transcriptional activity through mechanisms other than direct modulation of DNA binding, such as altered cofactor recruitment or protein–protein interactions.

      We agree with the reviewer that future biochemical approaches, including EMSA/gel shift assays or quantitative DNA-binding measurements, will be valuable to directly determine whether K170 methylation affects the intrinsic DNA-binding affinity of PPARγ. We have incorporated this new structural analysis and the corresponding discussion into the revised manuscript.

      (3) Along these lines, I wonder if there is another possibility: could SETD6-mediated methylation of PPARγ drive SETD6-PPARγ interaction? In other words, in the K170R, is SETD6 still even associated with PPARγ, and this interaction is required for promoter recruitment? Alternatively, would a catalytic dead version of SETD6 fail to associate with PPARγ? Currently, no experiments test the impact of an unmethylatable version of PPARγ or a catalytic dead version of SETD6 on SETD6-PPARγ interaction or SETD6 recruitment to promoters.

      We thank the reviewer for this insightful suggestion. To address whether SETD6 catalytic activity is required for its association with PPARγ, we performed an additional PLA experiment comparing SETD6 WT and the catalytic mutant SETD6 Y285A. As shown in the revised New Figure 2D, both SETD6 WT and SETD6 Y285A showed comparable proximity to PPARγ in cells, indicating that SETD6 catalytic activity is not required for the physical association between SETD6 and PPARγ.

      These findings support a model in which SETD6 first recognizes and binds PPARγ independently of its catalytic activity. Subsequent methylation of PPARγ at K170 is therefore likely to regulate the downstream functional consequences of this interaction, including enhanced chromatin occupancy and transcriptional activation, rather than the initial SETD6–PPARγ association itself.

      In addition, we generated using AlphaFold a structural model of the SETD6–PPARγ complex as a supportive visualization (New figure S2). Given the limited confidence of the prediction, we interpret this model cautiously and include it in the Supplementary Information rather than the main figures. We agree with the reviewer that future studies examining SETD6 recruitment to PPARγ target promoters and the effect of the PPARγ K170R mutant on SETD6–PPARγ association will further refine the molecular mechanism.

      Minor concerns (text and figure display)

      (4) The text has multiple typos and grammatical errors, and there are some issues with the figure display.

      We thank the reviewer for this comment. We carefully revised the manuscript to correct typographical and grammatical errors throughout the text and also addressed the figure display issues noted by the reviewer.

      Reviewer #2 (Public review):

      Summary:

      In this work, the authors investigated the regulation of the transcription factor PPARγ by the post-translational modification lysine methylation. The data demonstrate that the lysine methyltransferase SETD6 targets PPARγ for methylation using biochemical and cell-based assays. Methylation of PPARγ occurs in its DNA binding domain, and the authors demonstrate that loss of methylation limits PPARγ chromatin binding, particularly to lipid storage and metabolism gene promoters. As a physiological output, the authors demonstrate that deletion of SETD6 and loss of PPARγ methylation also disrupt lipid droplet accumulation in hepatocytes. In addition, the authors uncover a positive feedback loop in which SETD6 methylation of PPARγ also regulates its binding to the SETD6 promoter and expression of the gene.

      Strengths:

      One of the key strengths of this manuscript is the novelty of the findings in terms of identifying a new mode of regulation of PPARγ that modulates its chromatin association in cells and thereby regulates lipid metabolism genes. The authors nicely combine biochemical studies of SETD6 activity with cell-based assays investigating PPARγ and SETD6 function in regulating lipid storage. Data supporting this conclusion is largely convincing, and frequently, multiple assays are used to provide sufficient support to the conclusions. This work therefore expands regulatory modes of PPARγ and identifies a new target for SETD6, an enzyme that targets a number of other transcription factors. Furthermore, the regulatory loop that controls SETD6 expression via PPARγ methylation is likely important for understanding SETD6 function in different cell types that have high levels of lipid accumulation or regulation. The gene expression and lipid accumulation assays are useful for testing the physiological outcome of loss of SETD6 activity or PPARγ methylation directly.

      We thank the reviewer for his/her positive feedback on the manuscript.

      Weaknesses:

      The data presented in the manuscript are largely convincing in support of the authors' conclusions; however, there are some errors in the presentation of the figures and some issues in the text that would benefit from editing. Furthermore, there are some important questions not fully addressed in the results or discussion. 

      It would be great if the authors could speculate more on the diverse roles of SETD6 in methylated transcription factors and/or provide more context regarding the conditions that are likely to support methylation of PPARγ by SETD6. 

      We thank the reviewer for this important suggestion. In the revised Discussion, we expanded the manuscript to better place our findings within the broader context of SETD6-mediated regulation of transcription factors. Previous studies from our group and others demonstrated that SETD6 methylates multiple chromatin-associated transcriptional regulators, including RelA, E2F1, TWIST1, and BRD4, thereby modulating transcriptional selectivity, chromatin occupancy, and cofactor recruitment. We now discuss the possibility that SETD6 functions as a context-dependent signalling integrator that selectively regulates transcription factor activity through lysine methylation under distinct physiological conditions.

      In addition, we expanded the Discussion regarding potential cellular contexts that may favour PPARγ methylation by SETD6. Because PPARγ activity is strongly induced during lipid overload and fatty acid exposure, conditions associated with steatosis and metabolic stress may enhance the functional importance of SETD6-dependent methylation. We also discuss the possibility that chromatin accessibility, ligand-dependent activation of PPARγ, and metabolic signaling pathways may collectively influence the formation and stability of the SETD6–PPARγ complex.

      Also, while a potential cross-talk between methylation and phosphorylation is described in the discussion, it would be great to provide more structural insight into how this might regulate DNA binding of PPARγ and/or discuss whether there are other possibilities given the location of the target lysine in the DNA binding domain.

      We thank the reviewer for this valuable suggestion. To provide additional structural insight into the potential interplay between methylation and phosphorylation within the PPARγ DNA-binding domain, we performed structural modeling based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU). As shown in the new Figure S6, the modeled K170me1 side chain is predicted to face away from the DNA interface and does not introduce steric clashes with DNA, suggesting that K170 methylation is unlikely to directly alter the DNA-binding interface through steric effects. In contrast, phosphorylation of the neighboring residue T166 is predicted to introduce multiple intramolecular steric clashes within the DNA-binding domain. These structural changes could influence the local conformation or dynamics of the DNA-binding domain and thereby indirectly modulate PPARγ DNA binding or transcriptional activity.

      In addition, as suggested by the reviewer, we expanded the Discussion to consider alternative mechanisms by which K170 methylation may regulate PPARγ function. While our ChIP-qPCR experiments demonstrate that K170 methylation positively regulates PPARγ chromatin occupancy at target promoters, the structural modeling suggests that this effect is unlikely to arise from direct steric modulation of the DNA interface. Instead, K170 methylation may influence chromatin occupancy by regulating protein–protein interactions, cofactor recruitment, local conformational dynamics, or other chromatin-associated mechanisms. We have incorporated these new structural analyses and the expanded discussion into the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The KEGG panels are too low resolution to read.

      We thank the reviewer for this comment. We improved the resolution of the KEGG pathway enrichment panels and, as suggested by the reviewer (see below), we separated the upregulated and downregulated gene sets to improve clarity and readability. These changes are now reflected in the revised new Figures 5C, 5D, 6B and 6C.

      (2) Figure 3 panel D is hard to visualize; a better version or repeat experiment is needed.

      We thank the reviewer for this comment. To address this concern, we replaced the original Figure 3D with a new independent experiment that more clearly demonstrates the methylation of endogenous PPARγ by SETD6. We believe that the new data provide substantially stronger evidence and improve the clarity of the revised manuscript.

      (3) There is a typo in Figure 2A (line present in "S" of Signal).

      We thank the reviewer for pointing out this typo. The error in Figure 2A has now been corrected in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Overall, the experiments and analyses presented are sufficient to support the conclusions and interpretations of the work. However, there are some issues of presentation and writing that are worth addressing. These comments are listed below:

      (1) My only substantial recommendation is to improve the discussion to provide more context to the findings in terms of both the larger role of SETD6 in methylating transcription factors (some of whom also regulate its expression) and the potential modes through which PPARγ DNA binding activity could be regulated by methylation.

      We thank the reviewer for this insightful suggestion. In the revised manuscript, we substantially expanded the Discussion to better place our findings within the broader context of SETD6mediated regulation of transcription factors. We now discuss previous studies demonstrating that SETD6 methylates multiple chromatin-associated transcriptional regulators, including RelA, E2F1, TWIST1, and BRD4, thereby modulating chromatin occupancy, cofactor recruitment, and transcriptional selectivity. We further propose that SETD6 functions as a context-dependent signaling regulator that integrates distinct cellular pathways through the selective lysine methylation of transcription factors.

      In addition, we expanded the Discussion regarding the potential mechanisms by which PPARγ K170 methylation regulates transcriptional activity. We incorporated new structural modeling based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU), which suggests that K170 methylation is unlikely to directly alter the DNA-binding interface through steric effects, whereas phosphorylation of the neighboring residue T166 may induce intramolecular steric clashes within the DNA-binding domain. We also expanded the Discussion to consider alternative mechanisms by which K170 methylation may regulate PPARγ function, including modulation of chromatin occupancy, local conformational dynamics, protein–protein interactions, and recruitment of transcriptional cofactors or chromatin-associated proteins. Finally, we discuss that future biochemical studies will be important to determine whether K170 methylation also influences the intrinsic DNA-binding affinity of PPARγ.

      Can the authors incorporate any other published structural data to speculate on the role of methylation or describe more about how it is expected that the methylation-phosphorylation crosstalk modulates DNA binding? This type of discussion would better highlight the potential importance of this new modification on PPARγ.

      We thank the reviewer for this important suggestion. In the revised manuscript, we expanded the Discussion and incorporated additional structural analyses (new Figure 6F and new figure S6) based on the published PPARγ–DNA co-crystal structure (PDB: 3DZU). K170 is positioned within the DNA-binding domain, between the two zinc-finger motifs that mediate DNA recognition and stabilization on PPRE-containing DNA. Our structural modeling predicts that the K170me1 side chain is oriented away from the DNA interface and does not introduce steric clashes with DNA, suggesting that methylation is unlikely to directly alter the DNA-binding interface through steric effects. Nevertheless, lysine methylation can influence protein surface properties, protein– protein interactions, and recognition by regulatory binding partners, raising the possibility that K170 methylation modulates PPARγ chromatin occupancy or promoter selectivity through indirect mechanisms.

      In contrast, structural modeling predicts that phosphorylation of the neighboring residue T166 introduces multiple intramolecular steric clashes within the DNA-binding domain. These clashes could alter the local conformation or dynamics of the DNA-binding domain and thereby indirectly influence PPARγ–DNA interactions and transcriptional activity. Together, these observations suggest that K170 methylation and T166 phosphorylation may represent a regulatory crosstalk that fine-tunes PPARγ function through distinct structural mechanisms.

      We further expanded the Discussion to consider additional, non-mutually exclusive mechanisms by which K170 methylation may regulate PPARγ function, including modulation of chromatin occupancy, protein–protein interactions, cofactor recruitment, and stabilization of transcriptional complexes at target genes. While these possibilities require further mechanistic investigation, we agree with the reviewer that these structural considerations highlight the potential regulatory importance of this newly identified PPARγ modification.

      (2) In the gene expression experiments presented in Figure 5, it would be useful if the GO terms were described as enriched in either the up- or down-regulated gene sets. From the way it is presented, it is not clear if specific categories of genes are found enriched in those upregulated or downregulated upon KO of SETD6. This is also true for the experiments presented in Figure 6 regarding the PPARγ mutant.

      We thank the reviewer for this important suggestion. In the revised manuscript, we separated the pathway enrichment analyses into upregulated and downregulated gene sets for both the SETD6 knockout RNA-sequencing experiments (Figure 5) and the PPARγ WT versus K170R mutant analysis (Figure 6). This revision provides improved clarity regarding which biological pathways are positively or negatively associated with SETD6 depletion or disruption of PPARγ K170 methylation. The updated KEGG enrichment analyses are now presented in the revised new Figures 5C, 5D, 6B and 6C.

      In addition, if there is a significant overlap in genes misregulated in both mutants, this could be shown in the figure.

      We thank the reviewer for this suggestion. We compared the differentially expressed genes identified in the SETD6 KO cells and the PPARγ K170R mutant cells to evaluate the extent of overlap between the two datasets. However, we did not observe a substantial or statistically significant overlap in misregulated genes under the thresholds used in our analysis. Therefore, we decided not to include this comparison in the revised figure. Nevertheless, both datasets consistently showed enrichment for pathways associated with lipid metabolism and PPAR signaling, supporting a functional connection between SETD6 and PPARγ-mediated transcriptional regulation.

      (3) There are some typos and grammatical errors throughout the work, and it should be carefully edited. One example is the following heading: PPARγ K170 methylation by SETD6 regulates mediates lipid droplets formation.

      We thank the reviewer for this comment. The manuscript was carefully revised to correct typographical and grammatical errors throughout the text. In particular, the heading mentioned by the reviewer was corrected in the revised manuscript.

      Minor errors in the figures:

      (1) Formatting issue in the labeling for the x-axis of Figure 1E.

      We thank the reviewer for pointing out this formatting issue. The labeling of the x-axis in Figure 1E has been corrected in the revised manuscript.

      (2) Figure 2B - HA-SETD6 should be labeled as minus for the first lane.

      We thank the reviewer for pointing this out. We corrected the labeling in Figure 2B (now Figure S1), and the first lane is now properly indicated as negative for HA-SETD6.

      (3) Figure 2D - Labeling needs improvement. Is this FLAG-PPARγ? "NC" was not defined in the legend. If negative control, what type?

      We thank the reviewer for this comment. We improved the labeling and figure legend of Figure 2D for clarity. Specifically, we now clearly indicate that the experiment was performed using Flag-PPARγ, and we defined “NC” in the legend as the negative control condition. In addition, during the revision process we noticed that the original PLA experiment was performed in HeLa cells rather than HepG2 cells, as previously indicated. This has now been corrected throughout the revised manuscript.

      (4) Figure 3 - The CRSIPR control should be described somewhere in the legend or the methods.

      We thank the reviewer for this comment. We clarified the description of the CRISPR control cells in the Materials and Methods section. Specifically, we now explicitly state that the CRISPR control (CT) cells were generated using the empty lentiCRISPR vector without SETD6-targeting sgRNAs.

      (5) Figure 5B and 6A - The legend is not labeled nor defined in the text- fold-change, log2 fold change, z score?

      We thank the reviewer for pointing this out. We revised the figure legends for Figures 5B and 6A to explicitly define the heatmap scale and normalization method. Specifically, we now indicate that the heatmaps represent normalized gene expression values displayed as Z-scores.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Please include an uptake control (early time point) or time‑course to distinguish phagocytosis from intracellular killing.

      We agree that distinguishing uptake from intracellular killing is important for interpreting bactericidal assays. Due to a recent transition, we are unable to conduct additional early‑time‑point assays. To address this transparently, we have revised the manuscript to clarify that our measurements represent overall bacterial load reduction, reflecting the combined effects of uptake and killing.

      The normalization as ‘fold killing’ is non‑standard; please report absolute CFU (log scale).

      We have retrieved the raw data and re‑expressed all bactericidal activity measurements as absolute CFU. All relevant figures, legends, and text have been updated accordingly.

      Authors report quantification of cytokine concentrations, yet no information is provided regarding how these measurements were performed.

      Cytokine concentrations were quantified by ELISA. We have now added details regarding assay kits, sample preparation, and detection parameters to the Methods section.

      While the choice of IL‑1β and IL‑6 is straightforward, the focus on IL‑18 requires explicit justification.

      IL‑18 is a macrophage‑associated pro‑inflammatory cytokine with established links to inflammasome activation and trained immunity pathways, providing a clear justification for its inclusion.

      The methodology used to quantify immune cell populations presented in Figure 2 is not described.

      We have added a detailed description of the flow cytometry methodology, including staining and gating strategies.

      Immune cell quantification would be expected in the context of the challenge experiment as well.

      While we agree that such data would be valuable, additional mouse experiments cannot be performed because the animals used in the challenge model are not currently available during the transition period. If feasible, we are exploring ex vivo flow cytometry data from TWIK2 mutant versus wild‑type macrophages.

      AMs are not considered recruited immune cells; this should be corrected.

      We have corrected this in the figure legend and throughout the manuscript.

      The authors report n = 5 for the survival curves in the figure legend, whereas n = 7 is stated in the Methods section.

      We have corrected the sample size to ensure consistency between the figure legend and Methods.

      ATAC‑seq peaks are referred to as ‘genes’ and ‘differentially expressed genes’.

      We have corrected the terminology in the manuscript. ATAC‑seq identifies differentially accessible chromatin regions, which are then annotated to the nearest downstream gene.

      In Figure 7, trained WT and Nlrp3-/- mice display similar levels of bacterial clearance. How should this result be interpreted?

      A portion of this phenotype is via the normalization of phagocytosis to ‘fold killing’. Presentation of the raw CFU data shows a trend towards reduced bacterial clearance in Nlrp3-/- mice.

      Reviewer #2 (Public review):

      Sample numbers for experiments 1, 2, and 6 are not provided.

      We have added explicit n values for all experiments and verified their accuracy against the original records.

      The Discussion would benefit from a clear summary of study caveats.

      We agree and have added a dedicated paragraph outlining key Caveats as described.

      Specific identities of DEGs are not provided; only pathway enrichment is shown.

      We have now included a supplementary table listing the differentially expressed genes identified in our analysis.

      Controls for subcellular fractionation and dye microscopy should be included.

      Controls for subcellular fractionation have added in figure 3B and controls for dye microscopy have now been added in supplementary figure 1.

      The text states that protease inhibitors diminish ATP‑induced training effects, but the figure does not show significance.

      We have re‑examined the data and updated the figure to include statistical testing where appropriate.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      This study elucidates the molecular linkage between the mobilization of damaged rDNA from the nucleolus to its periphery and the subsequent repair process by HDR. The authors demonstrate that the nucleolar adaptor protein Treacle mediates rDNA mobilization, and the MDC1-RNF8-RNF168 pathway coordinates the recruitment of the BRCA1-PALB2-BRCA2 complex and RAD51 loading. This stepwise regulation appears to prevent aberrant recombination events between rDNA repeats. This work provides compelling evidence for the recruitment of the Treacle-TOPBP1-NBS1 complex to rDNA DSBs and demonstrates the critical role of MDC1 in the rDNA damage response. There are some issues with the over-interpretation of results as described subsequently. Some aspects could be strengthened, for example, a potential role of the RAP80-Abraxas axis, the origin of the repair synthesis (HDR vs. NHEJ), and a direct comparison of the RNF8 and RNF168 recruitment in the absence or presence of MDC1.

      We thank the reviewer for the positive assessment of our work and for the constructive suggestions. We agree that certain aspects of the manuscript required clarification, in particular the potential contribution of the RAP80–ABRAXAS pathway, the interpretation of the repair synthesis assay, and the role of RNF168 recruitment. We have addressed these points experimentally where feasible and have revised the manuscript accordingly to avoid overinterpretation.

      Reviewer #1 (Recommendations for the authors):

      Major comments

      (1) In Figures 4C, 4D, and S4B-D, BRCA1 and RAD51, recruitment to nucleolar caps is partially reduced upon RNF168 depletion. Despite this, the authors broadly conclude that recruitment mainly depends on the MDC1-RNF8-RNF168 pathway. Since the RAP80-Abraxas pathway may also contribute, as briefly mentioned in the Discussion, siRNA knockdown of Abraxas would help clarify the relative roles of these two pathways.

      We thank the reviewer for this important suggestion. To directly address the potential contribution of the RAP80–ABRAXAS pathway, we depleted RAP80 by siRNA and analysed BRCA1 and RAD51 recruitment to nucleolar caps following I-PpoI-induced rDNA damage.

      Strikingly, RAP80 depletion strongly impaired the formation of both BRCA1 and RAD51 nucleolar caps in two independent cell lines (U2OS and RPE1) (new Figure 5). These findings demonstrate that the RAP80–ABRAXAS pathway plays a critical role in BRCA1 recruitment at nucleolar caps.

      Together with our observation that RNF168 depletion only partially reduces BRCA1 and RAD51 recruitment, these results indicate that both RNF168-dependent and RAP80–ABRAXAS-dependent pathways contribute to HDR factor recruitment downstream of RNF8-mediated chromatin ubiquitylation.

      We have revised the model Figure (Figure 9) and the Results and Discussion sections accordingly to reflect this dual-pathway model and to avoid overemphasising the contribution of RNF168.

      (2) In Figure 7C, the EdU-γH2AX PLA assay detects DNA synthesis at rDNA breaks, but it remains unclear whether this signal reflects HDR- or NHEJ-mediated repair. Since MDC1 functions upstream of the DSB repair pathway choice, the observed reduction in PLA signal upon MDC1 depletion does not necessarily reflect impaired HDR alone. Synchronizing cells in G2 or using cell cycle markers would help clarify the repair context and strengthen the interpretation.

      We thank the reviewer for raising this important point. We agree that the EdU–gH2AX PLA assay does not exclusively report on HDR-mediated DNA synthesis and may also capture other forms of repair-associated DNA synthesis.

      In the revised manuscript, we have therefore tempered our interpretation and now describe this assay more cautiously as a readout of DNA synthesis at sites of rDNA damage, rather than as a direct measure of HDR activity.

      Importantly, our conclusion that MDC1 promotes HDR factor recruitment at nucleolar caps is based primarily on the reduced accumulation of BRCA1, PALB2, and RAD51, which are well-established markers of HDR. The PLA assay is now presented as supportive evidence for ongoing DNA synthesis at these sites rather than as a definitive indicator of HDR.

      We agree that further experiments, such as cell cycle synchronization or the use of phase-specific markers, would help to more precisely define the repair context, and we have included this point in the Discussion.

      (3) The authors propose that MDC1 is essential for RNF8-RNF168 recruitment, specifically at nucleolar rDNA breaks. A side-by-side comparison of RNF8 or RNF168 localization in the presence and absence of MDC1, with IR-treated conditions, would provide important validation of this model. Including representative images in Figure S2C would further support the claim.

      We agree with the reviewer that a direct analysis of RNF8 and RNF168 recruitment in the presence and absence of MDC1 would provide valuable mechanistic insight. We therefore attempted to address this experimentally.

      However, despite testing multiple antibodies, we were unable to obtain specific and reproducible signals for RNF8 and RNF168 at nucleolar caps, precluding a reliable analysis of its recruitment under these conditions.

      Given this technical constraint, we have revised the manuscript to avoid overinterpretation regarding direct RNF168 recruitment and instead focus on functional readouts of downstream ubiquitylation-dependent signalling, such as BRCA1 and RAD51 accumulation.

      We note that the requirement for MDC1 in BRCA1 and RAD51 recruitment at nucleolar caps is consistent with a role of MDC1 upstream of RNF8-dependent chromatin ubiquitylation, in line with its established function at IR-induced DSBs.

      Minor comments:

      (1) The legend for Figure 8 should more clearly explain the proposed mechanism and include concise titles or descriptions for each sub-panel.

      We agree with the reviewer that the model should be described in the Figure legend. We have thus updated the model to accommodate the new data and wrote a legend that concisely explains the proposed model. We do not think that titles for each sub-panel are required. Instead, we separately referred to the sub-panels in the legend.

      (2) Typos:

      (a) Page 8: PRE1 MDC1, as "RPE1 MDC1;

      (b) S3 Figure legend: Dhermacon";

      (c) Page 29: "80.103"-please clarify or correct.

      We thank the reviewer for pointing out these errors. These have been corrected in the revised manuscript

      Reviewer #2 (Public review):

      Summary:

      DNA double-strand breaks (DSB) in repeated DNA pose a challenge for repair by homologous recombination (HR) due to the potential of generating chromosomal aberrations, especially involving repeats on different chromosomes. This conceptual caveat led to a long-held notion that HR is not active in repeated DNA, which was disproven in groundbreaking work by Chiolo showing in Drosophila that DSBs in pericentromeric repeats are mobilized to the nuclear periphery for repair by HR. A similar mechanism operates in mouse cells, as shown by the Gautier laboratory, but the mobilization goes to the nucleolar periphery, called nucleolar caps. In this manuscript, the authors reexamine the role of MDC1 in the mobilization of DSBs in rDNA in human cells. Previous work has shown that MDC1 is replaced by Treacle, the gene associated with Treacher Collins syndrome 1, in its role as the main adaptor of the DNA damage response, and these results are confirmed here. The novelty of this contribution lies in the discovery that MDC1 is required downstream in the recruitment of BRCA1 and RAD51 to nucleolar DSBs that were mobilized to the nucleolar cap. Using multiple MCD knockout models and DSBs induced by the nuclease PpoI, which cleaves at nuclear sites as well as in the 28S rDNA, convincingly documents this role of MDC1 and shows that it acts upstream of the RNF8-RNF168 ubiquitylation axis. Using a proxy assay of co-localization of EdU incorporation at DSBs (gammaH2AX), evidence is provided that MDC1 is required for HR in rDNA. MDC1 was not required for RAD51 recruitment to IR-induced foci, but it is unclear whether this is related to the different DSB chemistry (enzymatic versus IR) or to the localization of the DSB (rDNA versus unique sequence genome).

      Strengths:

      (1) The manuscript is well-written, and the experimental evidence is nicely presented.

      (2) Multiple MDC1 knockout models are used to validate the results.

      (3) Convincing back-complementation data clarify the relationship between MDC1 and RNF8.

      Weaknesses:

      (1) The recruitment of BRCA2 was not directly demonstrated. This caveat could be recognized, as IF for BRCA2 is challenging.

      (2) PpoI also induces DSBs in the non-rDNA genome. These DSBs would be an ideal control to establish nucleolar specificity of the events described and clarify whether the difference between IR and PpoI is the chemical structure of the DSB or the location of the DSB.

      We thank the reviewer for the positive and insightful evaluation of our work. We appreciate the recognition of the conceptual advance and the robustness of our experimental approaches. We have carefully considered the reviewer’s suggestions and have revised the manuscript to clarify interpretation where appropriate, particularly regarding BRCA2 recruitment and the specificity of I-PpoI-induced DNA damage. Where possible, we have also added new analyses to strengthen the conclusions.

      Reviewer #2 (Recommendations for the authors):

      (1) The claim that the BRCA1-PALB2-BRCA2 is recruited (abstract, end of results section, discussion page 15) should be qualified as BRCA2 recruitment was not directly demonstrated.

      We thank the reviewer for this important point. We agree that BRCA2 recruitment was not directly demonstrated in our study, as reliable immunofluorescence detection of BRCA2 remains technically challenging.

      We have therefore revised the manuscript throughout (Abstract, Results, and Discussion) to avoid overstatement and now refer more precisely to the recruitment of BRCA1, PALB2, and RAD51, rather than implying direct recruitment of a BRCA1–PALB2–BRCA2 complex.

      We note that BRCA2 function is supported indirectly by the observed RAD51 loading, which depends on BRCA2 activity. However, we have clarified this point to ensure that our conclusions remain fully supported by the presented data.

      (2) The temporal sequence established in Figure 1, 1hr BRCA1 and 2 hrs PALB2, argues against recruitment of a stable BRCA1-PALB2-(BRCA2) complex. This should be acknowledged.

      We thank the reviewer for this insightful observation. We agree that the temporal separation between BRCA1 accumulation (1 h) and PALB2/RAD51 recruitment (2 h) argues against the recruitment of a pre-assembled, stable BRCA1–PALB2–BRCA2 complex.

      We have revised the manuscript to reflect this interpretation and now describe the recruitment of HDR factors as a sequential process rather than as the assembly of a pre-formed complex. This is consistent with current models in which BRCA1 promotes subsequent PALB2 and BRCA2 recruitment, ultimately leading to RAD51 loading.

      (3) The model predicts that MDC1-KO cells are proficient for transcriptional repression after nucleolar DSB induction. Has this been tested?

      We did not specifically test this in the current work, but previous results published by our group revealed that siRNA-mediated depletion of MDC1 in human cells had a minimal effect on rDNA transcriptional inhibition after DNA damage (Larsen et al., 2024).

      (4) The nuclear PpoI DSBs could be analyzed as a specificity control, and clarify whether the difference between IRIF and PpoI DSBs relates to the DSB chemistry or location.

      We thank the reviewer for this important point. We agree that I-PpoI induces DNA breaks both within rDNA repeats and at additional genomic loci.

      To address this, we have now analysed the formation of gH2AX-positive nucleolar caps and non-nucleolar gH2AX foci over time following I-PpoI expression (new Figure 1–figure supplement 2). We find that nucleolar caps form rapidly and are prominent at early time points, whereas gH2AX foci accumulate more gradually.

      These results indicate that nucleolar caps and non-nucleolar DNA damage responses can be distinguished both spatially and temporally, and support the use of nucleolar caps as a specific readout for rDNA damage in our study.

      In addition, we note that RAD51 recruitment to IR-induced foci is not affected by MDC1 loss, suggesting that the requirement for MDC1 in RAD51 loading is specific to nucleolar rDNA breaks rather than reflecting differences in DSB chemistry alone. We have clarified this point in the Discussion.

      Additional points:

      (5) Page 4 top: Shieldin.

      Corrected.

      (6) The general reader will be interested to learn about the connection of the Treacle function with Treacher Collins syndrome. Maybe a paragraph could be added to discuss this?

      We thank the reviewer for this suggestion. We agree that the relationship between Treacle and Treacher Collins syndrome may be of interest to a broad readership. Since the developmental pathology of Treacher Collins syndrome is currently thought to arise primarily from impaired ribosome biogenesis and nucleolar dysfunction rather than defective nucleolar DNA damage signalling, we felt that an extensive discussion would be beyond the scope of the present study. We have, however, added a brief statement introducing Treacle as the product of the TCOF1 gene mutated in Treacher Collins syndrome and noting that whether its DNA damage response function contributes to disease pathology remains an open question.

      (7) Figure 7: A short explanation could be added as to why hypoxia conditions were chosen for the p53-deficient cell lines.

      We thank the reviewer for pointing this out. We have added a brief explanation in the figure legend to clarify that hypoxia conditions were used to stabilise replication stress and enhance detection of DNA repair intermediates in p53-deficient cells.

      (8) A short statement on whether the repair of nuclear DBS is affected by Treacle could be added.

      We thank the reviewer for this interesting point. While our study focuses on nucleolar DNA damage, we did not observe evidence that Treacle is required for the repair of non-nucleolar DSBs. We have added a brief statement in the Discussion to clarify that Treacle appears to function specifically in the nucleolar DNA damage response.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors describe a role of sumoylation at K81 in p66Shc which affects endothelial dysfunction. This explores a new mechanism for understanding the role of PTMs in cellular processes.

      Strengths:

      The experiments are well planned and the results are well represented.

      Vascular tonality experiments were carried out nicely, given the amount of time and effort one needs to put in to get clean results from these experiments.

      Weaknesses:

      (1) The production of ROS has been measured in a very superficial way.

      The term "ROS" confers a plethora of chemical species which exerts different physiological effects on different cells and situations.

      Mitochondria through one of the source, but not the only source of ROS production. Only measuring ROS with mitosox do not reflect the cellular condition of ROS in a specific condition. I would suggest authors consider doing IF of oxidative stress specific markers, carbonyl group and also, maybe, Amplex red for determining average oxidative stress and ros production in the cells.

      As suggested, we employed an additional ROS-sensitive probe, H<sub>2</sub>DCFDA, which revealed an overall increase in intracellular ROS levels upon SUMO2 overexpression; this effect was reversed by knockdown of p66Shc. In addition, we performed the Amplex Red assay on conditioned media to assess extracellular ROS release. SUMO2 overexpression did not alter ROS levels detected in the media, whereas a paradoxical increase was observed following p66Shc knockdown.

      Author response image 1.

      Amplex Red assay performed in HUVECs with and without knockdown of p66Shc expressing SUMO2 (Ad-SUMO2) or a control virus (Ad-LacZ).

      Amplex Red predominantly detects hydrogen peroxide (H<sub>2</sub>O<sub>2</sub>), which may originate from NADPH oxidases or be generated through the dismutation of superoxide. In contrast to superoxide, H₂O₂ is relatively stable, membrane-permeable, and well recognized as a second messenger in endothelial signaling, where it modulates kinase and phosphatase activity and supports physiological vascular functions. Thus, the elevated H<sub>2</sub>O<sub>2</sub> detected following p66Shc knockdown may represent a signaling-competent redox state rather than a pathological increase in oxidative stress. Moreover, the SUMO2-p66Shc-mediated increase in mitochondrial ROS may be efficiently buffered by cellular antioxidant defense mechanisms, thereby limiting detectable extracellular ROS release.

      (2) 8-OHG signal seems very confusing in Figure 7E. 8-ohg is supposed to be mainly in the nucleus and to some extent in mitochondria. The signal is very diffused in the images. I would suggest a higher magnification and better resolution images for 8-ohg. Also, the VWF signal is pretty weak whereas it should be strong given the staining is in aorta. Authors should redo the experiments.

      We have provided a better image for figure 7E. We repeated the staining with another antibody for vWF which showed stronger signal for vWF.

      (3) PCA analysis is quite not clear. Why is there a convergence among the plots? Authors should explain. Also, I would suggest that the authors do the analysis done in Figure 8B again with R based packages. IPA, though being user-friendly, mostly does not yield meaningful results and the statistics carried out is not accurate. Authors should redo the analysis in R or Python whichever is suitable for them.

      We thank the reviewer for their valuable feedback and insightful suggestions.

      Regarding the PCA analysis, the observed convergence among the data points reflects the underlying biological similarity between the samples within each group. Given the relatively low abundance and limited number of quantified peptides in our dataset, the variance captured by PCA is modest, leading to partial overlap between groups. This convergence is therefore likely due to shared biological characteristics and inherent sample variability, rather than technical issues.

      In response to the reviewer’s suggestion on pathway analysis, we have re-performed the analysis using R-based approaches. Specifically, we utilized established pipelines PROGENy-based signaling pathway activity inference. These methods provide statistically robust and reproducible results. The updated analyses and corresponding figures have been included in the revised manuscript, replacing the previous IPA-based results (Fig. 8). We believe these additions strengthen the interpretation of signaling pathway alterations in our dataset.

      (4) The MS analysis part seems pretty vague in methods. Please rewrite.

      We have revised the methodology for MS.

      Reviewer #2 (Public review):

      Summary:

      The article builds on the earlier work that both p66Shc and SUMOylation are essential nitric oxide (NO) based development of endothelial vasculature (PMID: 10580504; 28760777 and 35187108). The current manuscript brings forward a finding of how SUMO2ylation of p66Shc mediated ROS production which is essential for endothelial cells. They further identify that lysine 81 of p66Shc is the residue which is conjugated to SUMO2 and is crucial for mitochondrial localization. They further show that K81 SUMO2ylation is essential for S36 phosphorylation.

      Strengths:

      Convincingly shows that p66Shc is SUMO2ylated on lysine 81 in cells and also shows that the phosphorylation (serine 36) reduces upon loss of this critical SUMOylation site.

      Weaknesses:

      All the experiments performed here are in overexpression background therefore, it would be crucial to show that p66Shc is SUMO2ylated at physiological levels.

      As detecting endogenous SUMO2-p66Shc is technically challenging considering the almost 92% homology between SUMO2 and SUMO3 and the absence of a specific antibody to detect p66Shc, we generated a custom-made antibody which can detect SUMO2-p66Shc (YenZym, CA). Using this antibody, we performed immunoprecipitation which showed endogenous SUMO2-p66Shc at a molecular weight higher than p66Shc suggesting the SUMO2 modification of p66Shc at physiological level.

      Reviewer #3 (Public review):

      Summary:

      The authors set out to determine how SUMO2 impairs endothelial function through direct modification of the protein p66Shc. p66Shc is known to promote reactive oxygen species production, and here the authors demonstrate that SUMO2 modifies p66Shc at lysine-81, resulting in increased phosphorylation, mitochondrial translocation. These are prosed to mediate the detrimental effects of SUMO2 in a mouse model of hyperlipidemia.

      Strengths:

      A major strength of this work is the multi-pronged approach combining biochemical assays, proteomic analyses, and a genetically modified mouse model expressing a SUMOylation resistant mutant of p66Shc. These experiments comprehensively illustrate that lysine-81 SUMOylation of p66Shc is necessary for the observed endothelial dysfunction in hyperlipidemic conditions.

      Weaknesses:

      One notable weakness is that the link between the observed cellular changes and the ultimate in vivo phenotype remains only partially explored. While the authors successfully show that p66ShcK81R knockin mice are protected from endothelial dysfunction in a hyperlipidemic context, additional experiments characterizing the broader tissue-specific roles, or examining further endothelial assays in vivo, would strengthen the mechanistic conclusions. It would also be beneficial to see more direct evaluations of p66Shc subcellular localization in the protective knockin mice to complement the proteomic findings.

      We agree with the reviewer’s suggestion. However, due to limited resources, we could not pursue additional studies.

      Despite these gaps, the data broadly support the authors' main conclusions. The authors lay out a plausible mechanistic pathway for how hyperlipidemia and increased global SUMOylation can converge on the oxidative stress pathway to provoke vascular dysfunction.

      The likely impact of this work on the field is noteworthy. Beyond clarifying how a single post-translational modification event can influence the pathophysiology of endothelial cells, the study provides a model for investigating broader roles of SUMO2 in other cardiovascular conditions and highlights the importance of identifying additional SUMOylation sites and their downstream impact.

      In conclusion, by demonstrating the direct SUMOylation of p66Shc at lysine-81 and linking that modification to endothelial dysfunction in a hyperlipidemic mouse model, this paper offers valuable insights into how broadly acting post-translational modifiers can evoke specific pathological effects.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Please rearrange the figures. It is very hard to follow.

      We rearranged some of the figures for a better presentation.

      Reviewer #2 (Recommendations for the authors):

      **Please note that I am not an expert of mouse work therefore I will not be able to comment on the mouse work (Figure 7).

      The following suggested changes major concerns will strengthen the work, and the minor concerns will increase the paper's accessibility to eLife's broad scientific community.

      Major concerns.

      (1) All the work done here is based on overexpression studies, therefore it will be very useful to show that p66Shc gets SUMO2ylated under physiological conditions using SUMO-Trap beads or using tandem SUMO-Interacting Motifs (as shown in Silva-Ferrada et al., 2013: https://doi.org/10.1038/srep01690).

      We have demonstrated endogenous SUMO2 conjugation of p66Shc under physiological conditions by immunoprecipitation using a custom-generated antibody specific for SUMO2-p66Shc. This approach allows detection of SUMO2ylated p66Shc without reliance on overexpression systems, thereby confirming that SUMO2 modification of p66Shc occurs endogenously.

      (2) The authors claim that K81 is the only site of SUMO2ylation and based on the HUVEC experiments (Fig 3C, D) it appears that p66Shc K81R mutant still gets SUMO2ylated indicating that there are other residues which can get SUMO2ylated. Do you see the other sites getting SUMO2ylated in your mass spec data?

      Our mass spectrometry analysis did not identify SUMO2 modification on lysine residues other than K81. However, we acknowledge that SUMOylation detected in vitro on recombinant protein may differ from SUMOylation occurring in a cellular context, where protein conformation, interacting partners, and local enzyme availability can influence modification patterns. These differences may account for the appearance of multiple SUMOylated p66Shc species in cell lysates despite the absence of additional SUMOylation sites in the mass spectrometry dataset. Our emphasis on K81 is based on its localization within the CH2 domain of p66Shc, a region unique to the p66 isoform. Previous studies have demonstrated that post-translational modifications within the CH2 domain, such as phosphorylation, are critical for activation of the oxidative and pro-apoptotic functions of p66Shc. Accordingly, we focused our mechanistic analyses on SUMO2ylation at K81. Nevertheless, we do not exclude the possibility that additional lysine residues within the PTB or CH1 domains of p66Shc may also undergo SUMOylation in cells. Importantly, our functional data support the conclusion that SUMO2ylation at K81 is a key regulatory modification driving the oxidative activity of p66Shc.

      (3) In Figure 4 the authors claim that SUMO2ylation at K81 is a prerequisite for S36 phosphorylation, however the phosphorylation changes are marginal (Figure 4 A, E) and why do they not see the higher molecular weight bands in the western blots.

      We appreciate the critique and agree with the reviewer that our data do not provide direct evidence that K81 SUMO2ylation and S36 phosphorylation occur simultaneously on the same p66Shc molecule. Rather, our findings suggest that SUMO2ylation at K81 may facilitate or enhance S36 phosphorylation, but is not an absolute requirement for this modification. Accordingly, we have revised the wording in the main text to avoid implying a strict prerequisite relationship and to more accurately reflect the magnitude of the observed changes in S36 phosphorylation.

      (4) The authors further claim that the SUMO2ylation at K81 effects S36 phosphorylation which has consequences for mitochondrial localization. However, K81R still translocate to mitochondria which is indicative of SUMO2ylation is important but not necessary (Fig 5A, C).

      We agree with the reviewer’s observation and have revised the main text accordingly, as described in our prior response, to clarify that K81 SUMO2ylation facilitates, but is not essential for, S36 phosphorylation and mitochondrial translocation.

      (5) To know if the K81 SUMO2ylation had any effect on phosphorylation it would be interesting to see how the phosphomimic (S36D) mutant along with/without K18R will behave in mitochondrial localization experiments.

      We agree that examining the double mutant (S36D/K81R) would provide valuable mechanistic insight into the relationship between K81 SUMO2ylation and S36 phosphorylation in regulating mitochondrial localization. Although we are currently unable to perform these experiments due to constraints in manpower and resources, we have acknowledged this as a limitation in the Discussion and plan to pursue these studies in future work to further validate the mechanism.

      Minor Concerns

      (a) In Figure 1A the authors claim that SUMO2ylation of p66Shc promotes ROS production however they do not include any well-known ROS induced proteins in the western blots such as KEAP1 or NRF2 or any other marker.

      We agree that NRF2 is a well-established transcriptional regulator of antioxidant responses. However, the primary aim of Figure 1A was to directly examine the effect of SUMO2ylation on p66Shc-mediated ROS production. While NRF2 and other ROS-responsive proteins reflect downstream adaptive responses, they do not provide a direct measure of ROS generation. Therefore, we focused on measuring SUMO2-induced changes in cellular ROS levels. To further strengthen this conclusion, we have performed additional complementary ROS assays, which are now included in the revised manuscript.

      (b) In Figure 2D, the concentration of Anacardic acid used is low and therefore the SUMO2ylation is not completely inhibited. It would be advisable if the authors go to concentration of complete inhibition.

      We acknowledge that some residual SUMO2ylation is visible in the immunoblots following anacardic acid treatment. However, the assay demonstrates a clear, dose-dependent reduction in SUMO2ylation levels, which is sufficient to interpret the effect of SUMO2 inhibition on p66Shc function. Increasing the concentration further could introduce off-target effects, and the current assay provides a reliable and physiologically relevant readout.

      (c) The figure panels need uniform fonts and also the size/resolution of the western blots needs to be improved

      We thank the reviewer for this suggestion. All figures have been updated to ensure uniform fonts, and the western blot images have been improved for size and resolution in the revised manuscript.

      (d) Supplemental figures are very low resolution.

      High-resolution images have been provided.

      (e) Please introduce abbreviations before using them (like LDLr and ND).

      The abbreviations have been explained.

      (f) Figure 6A seems like data that belongs in the SI rather than the main text.

      We kept it in main figure as the knock-in mouse was generated for this study.

      (a) Figure 7B needs a WT HFD control.

      It is a valid suggestion. However, it is known in the literature (we have prior experience as well) that wild-type mice do not exhibit a drastic increase in serum lipid level with high-fat diet feeding.

      (b) The differences in Figure 7E are not profound and immediately obvious. Is there a better way to show the data, potentially with quantification?

      We agree, the difference is not that huge, we have provided the qualification as Fig. 7F.

      (h) Figure panels 6E-H could use some labels to differentiate the data from WT vs p66ShcK81R mutant mice within the figure panel.

      We mentioned that on the top and have inserted a vertical line to separate the groups.

      (i) Lines 165-168, 182-185, 214-218 and 282-286: These sentences can be rewritten for more clarity.

      We have modified these sentence for clarity.

      (j) Line 512 should read, "Two-way ANOVA, ***P<0.001."

      Thank you, we have corrected the mistake.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We sincerely appreciate you and the reviewers for investing time and effort in evaluating our manuscript. After carefully reading the comments and suggestions, we found they are insightful, constructive, and critical for improving the quality of our work. Based on these valuable recommendations, we have substantially revised the manuscript as summarized below.

      Abstract: Inappropriate or ambiguous statements have been revised to improve clarity.

      Introduction: (1) The study purpose and hypotheses have been re-organized in a clearer and more concise way. (2) A mechanistic rationale for sequence-specific degradation has been provided and the use of PMA treatment has been explained. (3) The terminologies related to extracellular DNA and 16S rRNA gene amplicons have been clarified.

      Materials and Methods: 1) More detailed description of the microcosm experiment has been added. 2) The design and rationale of GAPDH F-tagged primers and the use of fusion primers for Illumina library preparation have been clarified. 3) More details about PCR amplification, DNA purification, and pooling strategies have been added. 4) We have corrected and standardized primer naming throughout the manuscript; 5) More details about bioinformatic workflow have been added. 6) We have defined statistical parameters and multiple testing corrections. 7) All abbreviations have been defined and standardized.

      Results: 1) The terminology for PMA-treated DNA has been revised and it has been clarified interpretation as “PMA-treated prokaryotic community” rather than “living community”. 2) The figures and legends have been updated for clarity, and the explicit explanation of “ASV I” and “ASV II” in pairwise comparisons have been added. 3) the figures (e.g., Figs. 2–5, S2–S8) have been reorganized to better reflect results; 4) Inappropriate statements or misleading interpretations have been removed.

      Discussion: A detailed section on technical limitations have been added. The limiatons added mainly include: 1) PCR amplification bias and recommendations for spike-in standards or multi-primer approaches; 2) differential DNA extraction efficiency due to variable cell lysis; and 3) limitations of using 16S rRNA amplicons as proxies for natural extracellular DNA and the limitations of PMA treatment efficiency in soil matrices;

      eLife Assessment

      This valuable study introduces an innovative experimental design to address a crucial and timely issue in microbial ecology: the potential bias in soil microbial community analyses caused by extracellular DNA degradation. While the evidence showing variable degradation rates of extracellular DNA is convincing, additional conceptual, methodological, and statistical clarifications could reinforce the claims and the study's contribution to the field. This research will appeal to microbial ecologists and researchers interested in using molecular techniques to evaluate microbial community structure.

      We sincerely appreciate the editors for the careful assessment of our work and for recognizing the value of addressing extracellular DNA degradation in soil microbial community analyses. We also greatly appreciate the reviewers’ constructive feedbacks concerning the need for additional conceptual, methodological, and statistical clarifications. We agree that further refinement in these areas will strengthen our claims and enhance the study’s contribution to the field. Based on these insightful suggestions, we have carefully revised the manuscript to provide clearer conceptual framework, more detailed methodological descriptions, and more rigorous statistical analyses. We believe these revisions have substantially improved the clarity and robustness of our work. More details about the revisions have been provided in the following responses.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript investigates the degradation dynamics of extracellular DNA in soils and its impact on estimates of microbial abundance and diversity. By combining a broad geographic sampling design with a primer-labeling strategy, qPCR quantification, amplicon sequencing, and PMA treatment, the authors aim to disentangle total versus intracellular DNA signals and explore sequence-specific degradation patterns. The topic is relevant, particularly given the increasing awareness of relic DNA as a confounding factor in microbial ecology. The experimental design is ambitious and potentially impactful. However, several conceptual inconsistencies, methodological ambiguities, and statistical limitations currently weaken the robustness of the conclusions. These issues need to be addressed.

      We sincerely appreciate the reviewer for the constructive assessment of our work. We also appreciate the reviewer’s critical insights regarding the conceptual inconsistencies, methodological ambiguities, and statistical limitations that currently weaken the robustness of the conclusions. We agree with the reviewer that addressing these issues is essential to strengthen our work. Based on these valuable comments, we have carefully revised the manuscript to clarify the conceptual framework. Additionally, we have provided more detailed methodological descriptions, and enhance the statistical rigor of our analyses. We believe these revisions have substantially improved the clarity, consistency, and overall robustness of our conclusions.

      Strengths:

      The manuscript addresses a timely and important question in microbial ecology, particularly given the growing recognition that relic DNA can bias interpretations of community composition derived from amplicon sequencing. The study is ambitious in scope, incorporating a broad geographic sampling design across multiple soil types, which enhances the generalizability of the findings. The use of a controlled microcosm experiment combined with a primer-labeling strategy to track extracellular DNA dynamics is conceptually innovative and provides a structured framework to investigate degradation processes.

      In addition, the integration of multiple approaches, including qPCR for absolute quantification, high-throughput sequencing for community profiling, and PMA treatment to differentiate extracellular from intracellular DNA, represents a comprehensive attempt to disentangle complex sources of bias in soil microbiome analyses. The effort to link degradation dynamics with environmental variables and to explore sequence-level patterns further demonstrates the authors' intent to move beyond descriptive analyses toward a mechanistic understanding.

      We sincerely thank the reviewer for the positive and encouraging comments of our work.

      Weaknesses:

      Several conceptual and methodological issues currently limit confidence in the study's conclusions. Key terms such as "sequence-specific degradation" are not clearly defined or supported by a mechanistic or structural hypothesis, making it difficult to interpret the biological meaning of the results. In addition, the bioinformatic workflow presents inconsistencies, particularly the use of ASVs followed by clustering at 97% similarity, which undermines the resolution required to support sequence-level inferences. Statistical analyses are also insufficiently described, including unclear definitions of "T values," a lack of detail on pairing structure, and no indication of multiple testing correction.

      Furthermore, important methodological details are missing or unclear, including primer design (e.g., GAPDH tag vs ACTF), Illumina library preparation (e.g., adapter and indexing strategy), and validation of PMA treatment efficiency. The interpretation of PMA-treated samples as representing "living communities" is likely overstated, given the known limitations of the method in soil systems. Finally, typographical errors, inconsistent terminology, and unclear phrasing throughout the manuscript reduce readability and further complicate interpretation.

      We sincerely appreciate the reviewer’s thorough and critical evaluation of the manuscript’s weaknesses. The issues raised regarding conceptual clarity, bioinformatic consistency, statistical rigor, methodological transparency, and the interpretation of PMA treatment have been fully acknowledged. We also recognized that typographical errors, inconsistent terminologies, and unclear phrasing largely reduced readability. In response to these valuable comments, the manuscript has been carefully revised as follows. (1) The clearer definition of the term “sequence-specific degradation” has been provided. (2) The bioinformatic workflow was streamlined to ensure consistency. (3) The descriptions of statistical analyses were substantially expanded, including explicit definitions of “t values,” detailed clarification of the pairing structure, and the application of appropriate multiple testing corrections. (4) Missing details regarding primer design, Illumina library preparation, and PMA treatment validation have been added to the Method section. (5) Interpretations of PMA-treated samples have been revised to more accurately reflect methodological limitations in soil systems. (6) The manuscript has been thoroughly proofread to correct typographical errors, standardize terminology, and enhance overall clarity. These revisions are believed to substantially address the concerns raised and significantly strengthen the manuscript. A point-by-point response to the specific comments is provided below.

      Reviewer #2 (Public review):

      Summary:

      This manuscript describes the results of an interesting study examining the rate of degradation of extracellular DNA in soil ecosystems using a clever experimental approach. 16S ribosomal RNA genes were amplified from soil samples, and then purified PCR amplicons, containing a 5' linker sequence on the forward primer, were introduced to soils and monitored over time using real-time quantitative PCR and NGS amplicon sequencing. The study was able to measure rates of overall extracellular DNA degradation, but also sequence-specific degradation rates. I like the idea and execution of the study, and the results are interesting. The manuscript needs some help to improve the overall readability. Please see general and editorial comments below.

      We sincerely thank the reviewer for the positive and encouraging assessment of our study. We have carefully revised the manuscript to enhance clarity, streamline the presentation, and refine the language throughout. We believe these improvements have made the manuscript more readable and easier to follow. We are also grateful for the general and editorial comments provided, which have been addressed as outlined below.

      Strengths:

      Innovative experimental design that is well deployed across a large number of soil types, revealing interesting variability in extracellular DNA degradation.

      We sincerely thank the reviewer for the positive and encouraging assessment of our work.

      Weaknesses:

      (1) The manuscript needs another review to improve the readability of the document.

      We thank the reviewer for this helpful suggestion. We fully agree that improving readability is essential for effectively communicating our findings. Based on the comment, we have carefully revised the manuscript to enhance clarity and readability. We have streamlined sentence structures, standardized terminology, corrected typographical errors, and improved the logical organization of the text. We believe these revisions have substantially improved the overall readability of the manuscript.

      (2) The authors have used 16S genes to look at sequence-specific degradation. But 16S rRNA genes are actually pretty well conserved, and there isn't as much genetic variation across this gene among organisms as there is for other genes. It might be more relevant to look at metagenomic DNA degradation from high AT, high GC organisms, etc. This would be more generalizable than 16S genes.

      We thank the reviewer for this insightful comment. We agree with the reviewer that 16S rRNA genes are relatively conserved compared to functional genes or whole metagenomic DNA, and that studying degradation of more variable sequences (e.g., high‑AT, high‑GC regions, or metagenomic DNA) would provide greater generalizability. However, we would like to clarify the rationale for using 16S rRNA gene amplicons in the present study. First, the 16S rRNA gene remains the most widely used phylogenetic marker in soil microbial ecology (Knight et al., 2018). Demonstrating sequence‑specific degradation with this well‑established marker directly informs a large body of existing research that relies on 16S RNA gene‑based community analyses. Second, despite its conserved nature, the targeted fragment in this study is belong to the highly varied region (V4) of 16S rRNA gene. Accordingly, we indeed observed significant sequence‑specific variation in degradation rates among different 16S rRNA gene amplicon sequence variants (ASVs) (Fig. 2c, 3a). This indicates that even within a conserved marker gene, sequence‑dependent degradation biases exist and can affect diversity estimates. Third, our study was designed as a proof‑of‑concept to establish a methodological framework for quantifying both overall and sequence‑specific degradation rates. Using a single, well‑characterized marker allowed us to develop and validate the primer‑labeling and qPCR/sequencing workflow without the additional complexity of metagenomic DNA (e.g., variable fragment lengths and complex mineral associations). In the revised manuscript, we have added the following sentence to the Discussion section to address the concerns from the reviewer.

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015)”

      (3) Consideration of differential cell lysis during soil DNA extraction needs to be considered as well.

      We thank the reviewer for raising this important technical consideration. We agree that differential cell lysis during soil DNA extraction is a well‑recognized source of bias in microbial community analysis. Different microbial taxa (e.g., Gram‑positive vs. Gram‑negative bacteria, spores, or fungi) vary in their cell wall structure and susceptibility to lysis, which can lead to under‑representation of certain groups and over‑representation of others in the extracted DNA. This bias affects both total DNA extracts and PMA‑treated fractions, potentially influencing our estimates of the relative contributions of intact‑cell derived DNA versus extracellular DNA. However, currently, eliminating these biases are still challenging, and thus we have added the following sentence to the Discussion to address this concern.

      L305-311

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015).”

      (4) It is not clear why the authors didn't put GAPDH linkers on the reverse primer as well. This would have given an easier amplicon to amplify (no degeneracies at all).

      The decision to place the GAPDH linker only on the forward primer (515F) was intentional to balance the need for tracking exogenous extracellular DNA with amplification efficiency, sequencing quality, and cost-effectiveness. Adding linkers to both primers would increase the total amplicon length, potentially reducing amplification efficiency, especially in complex soil samples with degraded or low-quality DNA. More importantly, the reverse primer used in our study is a degenerate primer designed to target the 16S rRNA gene across diverse bacterial taxa, and extending it with an additional GAPDH linker could introduce further complexity, decrease amplification efficiency, and increase primer-dimer formation. Additionally, single-end labeling allows the usage of standard 16S rRNA reverse primers with existing barcodes, whereas dual-end labeling would require synthesis of new barcode-labeled primers, increasing both cost and time. Our preliminary experiments confirmed that single-end labeling produced reproducible amplification curves (~85% efficiency) and high-quality sequencing reads, which were sufficient for quantifying degradation rates. We have added a clarification in the Methods section to explain this rationale.

      L365-371

      “The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments

      (1) Inconsistency between ASV inference and 97% sequence recruitment

      The bioinformatic pipeline presents a major conceptual inconsistency. ASVs are inferred using UNOISE3, but reads are subsequently mapped to ASVs at a 97% similarity threshold, effectively reintroducing OTU-level clustering. Given that the manuscript's central claim is sequence-specific degradation, this step undermines the single-nucleotide resolution that ASVs provide and may obscure biologically meaningful differences. The authors should either reanalyze the data using a consistent ASV framework (exact matching), or explicitly treat the analysis as OTU-like and moderate claims of sequence specificity.

      We appreciate the reviewer's critical evaluation of our bioinformatic pipeline. We also apologized for our unclear statements in our original manuscript. We understand the concern that mapping reads to ASVs at 97% similarity might appear to reintroduce OTU-level clustering. Actually, we used the default pipeline provided by the authors of USEARCH with otutab command to generate the ASV table.

      Following the logic and recommendations of the USEARCH/UNOISE3 developer (Robert Edgar), this approach is a standard procedure for robust noise management rather than a conceptual inconsistency. First, in our pipeline, ASVs (ZOTUs) are first inferred using the UNOISE3 algorithm, which effectively identifies “true” biological sequences at single-nucleotide resolution. Secondly, according to the USEARCH manual, while the ASVs themselves represent exact biological sequences, the raw reads inevitably contain stochastic sequencing errors. Using “exact matching” for recruitment would discard a significant portion of the data that originates from a specific ASV but carries minor random errors. Mapping at 97% identity is the recommended method to recruit these noisy reads back to their correct biological origin (the ASV centroid). Meanwhile, during the recruitment process, reads are not randomly assigned to any ASV within the 97% identity radius. Instead, the algorithm follows a “highest similarity first” principle. For instance, if a specific read exhibits 98% similarity to ASV1 and 97% similarity to ASV2, it is strictly assigned to ASV1. A read is only discarded if its highest similarity to any ASV falls below the 97% threshold. Unlike traditional OTU clustering (where sequences are clustered together based on similarity from the start), our approach maintains the ASV as a fixed biological reference. The quantification of degradation rates is performed on these high-resolution centroids. Thus, our claims regarding sequence-specific degradation remain valid, as the underlying biological variation is defined by the ASVs. To avoid any possible confusion, we have rewritten the relevant paragraph in the Methods section (subsection 4.6) as follows.

      L446-457

      “ASVs were generated using the UNOISE3 non‑clustering denoising algorithm (Edgar, 2016), which infers 100% exact sequence variants by distinguishing biological sequences from PCR/sequencing errors. ASVs with total sequence counts fewer than 9 across all samples were removed to reduce noise. To quantify the abundance of each ASV, an ASV table was generated by mapping the quality‑filtered raw reads back to the ASV set using the otutab command. A 97% similarity threshold was applied for this recruitment to accommodate stochastic sequencing noise while maintaining biological resolution. Crucially, the mapping followed a best-hit priority rule, where each read was assigned to the ASV with the highest per cent identity within the 97% radius. This approach ensures that reads derived from the same biological template are accurately counted toward their respective ASV, preventing the underestimation of abundances that would occur with exact matching while strictly preserving the single-nucleotide resolution of the ASV framework.”

      (2) Undefined "ASV I" and "ASV II" groups

      The manuscript refers to "ASV I" and "ASV II" groups in pairwise comparisons of degradation rates (e.g., Fig. 3), but these groups are not defined anywhere in the text.

      It is unclear whether these represent: predefined biological categories, arbitrary pairwise ASV comparisons, or groupings based on taxonomy, abundance, or degradation rate.

      In addition, the pairing structure underlying these comparisons is not described. While a paired t-test is mentioned, it is unclear how ASVs were paired (e.g., within sites, across samples, or across time points).

      The current terminology ("groups") is potentially misleading and suggests biological structure where none may exist. The authors should explicitly define these terms, clarify the pairing scheme, and revise terminology if these are simply pairwise comparisons.

      We thank the reviewer for this keen observation. We completely agree that the terms "ASV I" and "ASV II" were poorly defined and potentially misleading.

      We would like to clarify that "ASV I" and "ASV II" were not intended to represent predefined biological categories (such as groupings based on taxonomy, abundance, or degradation rates). Instead, they were merely used as a labeling convention to indicate the directionality of pairwise comparisons within the heatmap matrix. Specifically, "ASV I" referred to the ASVs represented in the rows, while "ASV II" referred to those in the columns. In the original Fig. 3, blue indicated that the degradation rate of the row ASV was significantly lower than that of the column ASV, and red indicated the opposite. To avoid any confusion, we have removed the “ASV I/II” terminology throughout the manuscript and figures, replacing them with “Row ASVs” and “Column ASVs”. To address this issue, we have revised the Figure 3 legend to include a more explicit explanation:

      L820-824

      “In the heatmap, each cell represents a pairwise comparison between two ASVs. Blue indicates that the degradation rate of the ASVs listed in the row (row ASVs) is significantly lower than that of the ASVs listed in the column (column ASVs); red indicates that the row ASVs has a significantly higher degradation rate than the column ASV. A positive t value indicates that the row ASVs degrades significantly faster than the column ASVs; a negative t value indicates the opposite.”

      We thank the reviewer for raising the important issue regarding the definition of “t values” in our statistical analysis. We apologize for the lack of clarity in the original manuscript. To clarify, the T values presented in Figure 3a represent the test statistics (t-values) from paired t-tests comparing the degradation rate constants of two ASVs across the 30 study sites. The T-value was obtained from a paired t-test between two ASVs across the same samples. The t-value indicates the magnitude and direction of the difference between the two ASVs’ degradation rates relative to the variability across sites. A positive t-value (colored red in the heatmap) indicates the row ASVs degrades significantly faster than the column ASVs; a negative t-value (colored blue) indicates the opposite.

      L544-548

      “As for the analysis, we performed paired t‑tests across all the study sites. Thus, the degradation rates were essentially compared within each site, with both values originating from a same soil sample under identical incubation conditions. A positive t value indicates that the first ASV has a significantly higher degradation rate than the second one, and a negative t value indicates the opposite. The p values were adjusted for multiple comparisons using the FDR method.”

      (3) Lack of definition and justification of "T values"

      The manuscript reports "T values" for comparisons between ASVs but does not clearly define how these values are calculated. Although a paired t-test is mentioned, it remains unclear how the pairing was constructed, whether assumptions (normality, independence) were evaluated, and whether corrections for multiple comparisons were applied. Given the large number of ASVs, failure to control for multiple testing could inflate false positives. More broadly, the use of a simple paired t-test may not be appropriate given the hierarchical and compositional structure of the data.

      We sincerely thank the reviewer for pointing out the need to clarify the definition and justification of the t-values presented in our manuscript. Each t-value represents the test statistic from a paired t-test comparing the degradation rates of two ASVs across the same set of samples. The paired t-test assumes that the differences between paired observations are approximately normally distributed and that the pairs are independent across columns. We have evaluated the normality of differences using standard diagnostic plots and verified that the assumption is reasonably satisfied given the sample size. We performed a correction for multiple comparisons using the False Discovery Rate (FDR) procedure to control for potential false positives. We have revised the Methods section to clearly define t-values.

      (4) Conceptual validity of "sequence-specific degradation"

      The manuscript repeatedly refers to "sequence-specific degradation" of extracellular DNA; however, this concept is not clearly defined nor supported by a biological or structural hypothesis. It is unclear what "sequence-specific" refers to (e.g., nucleotide composition, GC content, secondary structure, taxonomic identity), whether differences are expected in conserved versus variable regions of the 16S rRNA gene, or what mechanistic basis would explain differential degradation among sequences. Given that the analysis is based on short 16S V4 amplicons, and no structural or biochemical framework is provided, it is difficult to interpret whether the observed differences truly reflect intrinsic sequence-dependent degradation or are instead driven by methodological or statistical artifacts (e.g., abundance effects, amplification bias).

      I believe the authors should explicitly define what is meant by "sequence-specific degradation," provide a biologically grounded hypothesis (e.g., structural accessibility, GC content, stem-loop stability), and align their interpretation with the resolution and limitations of the data.

      We thank the reviewer for this critical conceptual comment. We apologize that “sequence‑specific degradation” was not clearly defined and lacked a biological or structural hypothesis. To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation. This adjustment ensures that the hypotheses are directly grounded in the theoretical framework (e.g., GC content, thermodynamic stability, and secondary structures) presented in the paragraph.

      We now define “sequence‑specific degradation” as statistically significant differences in first‑order degradation rate constants among distinct ASVs, mainly arising from intrinsic DNA properties (base composition, secondary structure, and restriction sites) or differential mineral adsorption.

      L99-104

      “Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      We also expanded the mechanistic discussion to include GC content and secondary structure.

      L230-235

      “We also examined whether GC content could explain the observed sequence‑specific patterns, but no significant correlation was found (Fig. S4), suggesting that simple base composition is not the primary driver in this study. However, this does not exclude the possibility that higher‑order structural features (e.g., hairpin loops) or sequence‑specific nuclease recognition motifs contribute to differential degradation (Wang et al., 2007). This should be tested in future studies using synthetic DNA constructs with controlled structural elements.”

      We acknowledge that inferring sequence‑specific degradation from combined relative abundance and qPCR data is subject to potential methodological artifacts, including compositional effects, PCR amplification bias, and abundance‑dependent detection limits. However, we have taken several stringent steps to minimize these concerns. Specifically, we restricted our analysis to ASVs that were present in more than 90% of the study sites and for which the degradation curve fits yielded R<sup>2</sup> > 0.5, ensuring that only robustly detected and reliably modeled sequences were retained. Because our analysis tracks the ratio of each ASV across a time series, any sequence-specific PCR amplification bias remains constant for that particular sequence. By focusing on the rate of change rather than absolute read counts, such systematic biases are mathematically canceled out during the calculation of degradation kinetics.

      (5) Conceptual ambiguity in "GAPDH F-labeled 16S rRNA genes"

      The manuscript repeatedly refers to "GAPDH F-labeled 16S rRNA genes," which is confusing and may be misinterpreted as targeting GAPDH rather than 16S. It should be clearly stated that GAPDH refers to glyceraldehyde-3-phosphate dehydrogenase, and a GAPDH-derived sequence is used as a synthetic tag appended to a 16S primer. Additionally, the divergence of this tag from microbial sequences should be justified to ensure specificity. There is also an inconsistency in primer naming (e.g., "GAPDH F" vs "ACTF" in the figures), which should be corrected.

      We sincerely thank the reviewer for this important comment. We agree that the phrase “GAPDH F‑labeled 16S rRNA genes” could be confusing, as it may be misinterpreted as targeting the GAPDH gene rather than the 16S rRNA gene. We have revised the manuscript to avoid this ambiguity and to provide clear justification for the use of the GAPDH tag. GAPDH (glyceraldehyde‑3‑phosphate dehydrogenase) is a human housekeeping gene. Its forward primer sequence (GAPDH F: 5′‑CAT TGG CAA TGA GCG GTT C‑3′) was used as a synthetic tag appended to the 16S primer because (i) no homologous sequences exist in soil DNA (confirmed by PCR), and (ii) its melting temperature is compatible with the reverse primer. This tag allows specific tracking of exogenous DNA without interference from native soil sequences.

      Throughout the manuscript, ambiguous phrases such as “GAPDH F‑labeled 16S rRNA genes” have been replaced with more precise terms, “GAPDH F‑tagged 16S rRNA gene amplicon fragments” clarifying that the tag is an appendage and not the amplification target.

      We have checked the entire manuscript and confirm that “ACTF” does not appear anywhere. To avoid confusion, the primer is now consistently referred to as “GAPDH F” in all figures, legends, and text.

      L360-371

      “GAPDH is a primer for a human housekeeping gene and it has no homologous sequences in soils. Subsequently, GAPDH was selected as the label primer based on two criteria. First, this primer was selected to avoid interference from the original soil sequences (Huang et al., 2014; Yang et al., 2021; Arvizu-Hernandez et al., 2025), and no detectable PCR amplification was observed for the primer set GAPDH F-806R across all the soil DNA samples included in this study. Second, the melting temperature (Tm) value of GAPDH F approximately matched that of 806R. The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      (6) Limitations of using PCR amplicons as proxies for extracellular DNA

      The study uses PCR-generated amplicons to simulate extracellular DNA. While useful for controlled comparisons, these fragments may not reflect the physicochemical diversity of natural extracellular DNA (e.g., adsorption to minerals, fragment size variability, protection within aggregates). This limitation should be explicitly acknowledged, and conclusions should be framed accordingly.

      We appreciate the reviewer’s constructive feedback. We fully acknowledge that using PCR-generated amplicons to simulate extracellular DNA (eDNA) has inherent limitations in capturing the full physicochemical diversity of naturally occurring eDNA in soils. Specifically, we agree that PCR fragments may not replicate features such as highly variable fragment size distributions, associations with complex cellular components (e.g., vesicles or protein complexes), or long-term physical sequestration within soil micro-aggregates. Despite of these limitations, the use of uniform primer-tagged PCR amplicons was a deliberate choice to enable precise tracking of exogenous DNA degradation kinetics while eliminating background interference from endogenous soil eDNA. This design is a prerequisite for the high-resolution kinetic modeling of sequence-specific decay. Furthermore, in our bioinformatic pipeline, the 97% mapping threshold was specifically applied to minimize the influence of stochastic sequencing and PCR errors on abundance quantification. In the revised manuscript, these potential limitations have been addressed.

      L295-299

      “First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020).”

      (7) Interpretation of sequence-specific degradation

      Sequence-specific degradation rates are inferred from combining relative abundance data with total qPCR estimates. This approach is sensitive to compositional effects, amplification biases, and abundance-dependent detection limits. It remains unclear whether observed differences reflect true sequence-specific degradation or methodological artifacts. This limitation should be discussed more explicitly.

      We thank the reviewer for highlighting this critical methodological point. In our study, sequence-specific degradation rates were estimated by combining ASV-relative abundances with total qPCR-derived 16S rRNA gene copy numbers. We acknowledge that this approach may be influenced by compositional effects, PCR amplification biases, and abundance-dependent detection limits. However, the degradation rate constant (k) in our study, represents the rate of change for a specific sequence over time. Since PCR amplification biases are generally sequence-specific and consistent across samples processed under identical conditions, these systematic errors are mathematically canceled out when calculating the relative change (slope) for the same ASV across a time series. Second, all qPCR measurements were performed with three technical triplicates with standard curves to ensure quantitative reliability. Third, relative abundances were converted to absolute abundances using total qPCR estimates, allowing cross-taxa comparisons that reduce compositional bias. This approach is widely recognized in microbial ecology as a robust method. To address this concern, we have added some explanations in the revised manuscript.

      L84-86

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequence variants (ASVs) under identical soil and incubation conditions.”

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      L311-314

      “While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017)..”

      (8) Overinterpretation of PMA-treated samples as "living communities"

      The manuscript interprets PMA-treated DNA as representing intracellular or "living" microbial communities. While PMA is useful, this interpretation should be treated with caution in soils. PMA efficiency can be affected by soil matrix complexity, DNA adsorption to particles, incomplete light penetration, and permeability of compromised cells. Importantly, no validation of PMA efficiency is presented.

      We thank the reviewer for this important caution. We agree that interpreting PMA‑treated DNA as representing “living” or “intracellular” communities is an overstatement in soil systems. In the revised manuscript, we no longer describe PMA-treated DNA as a direct proxy for the “living community,” but instead refer to it as the “PMA-treated prokaryotic community”.

      Although we did not directly validate PMA efficiency in this study, we used a standardized PMA protocol that has been widely applied in microbial ecology, and our goal was to obtain a comparative estimate of the influence of extracellular DNA on community analysis across soils under a consistent methodological framework. Based on previous studies (Carini et al., 2016; Du et al., 2025), which found that in similar soil types, PMA treatment can significantly reduce the interference of extracellular DNA and alter the community structure, this indirectly proves the effectiveness of this technique.

      Nevertheless, we agree that future studies should include explicit validation controls, such as live/dead cell mixtures, heat-killed controls, or soil-specific PMA efficiency tests, to better quantify method performance across diverse soil matrices. We have added a dedicated paragraph in the "Methodological Considerations and Limitations" section to discuss how soil-specific properties (e.g., turbidity, adsorption capacity) might lead to incomplete exclusion of extracellular DNA, thereby advising a more cautious interpretation of the "viable" community data.

      L496-500

      “To inhibit amplification of eDNA, soils were incubated with propidium monoazide (PMA), as described previously (Carini et al., 2016). Upon photoactivation, eDNA can form covalent bonds through cross-linking, leading to the inhibition of its PCR amplification. In contrast, microbes with intact cell membranes exclude PMA, and their DNA is not cross-linked with PMA, and remains amenable to PCR amplification.”

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      Minor Comments

      (1) Line 79: Provide examples of how extracellular DNA contributes to nutrient cycling (e.g., P, N sources) and signal transduction (e.g., horizontal gene transfer).

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have added specific examples to clarify how extracellular DNA contributes to nutrient cycling and signal transduction. Specifically, we now note that extracellular DNA can serve as a source of phosphorus and nitrogen following enzymatic degradation, thereby contributing to soil nutrient turnover. We also clarify that extracellular DNA plays an important role in horizontal gene transfer, acting as a genetic reservoir that can be taken up by competent microorganisms and thereby facilitating the spread of functional traits such as antibiotic resistance. These examples have been added to improve the clarity and biological context of this statement.

      L48-52

      “EDNA serves as a critical vector for horizontal gene transfer (HGT), facilitating the uptake of genetic material by competent microorganisms and promoting the spread of functional traits such as antibiotic resistance (Liu et al., 2024). In addition, eDNA participates in soil biogeochemical cycling because its enzymatic degradation releases bioavailable nutrients, particularly phosphorus and nitrogen, which can be reused by soil microorganisms (Ye et al., 2022).”

      (2) Line 79: Replace "for an extended period of time" with a more precise or referenced timescale.

      We agree with the reviewer. We have replaced the vague phrase with a precise timescale. Extracellular DNA can persist in soils for months to years.

      (3) Line 94: Clarify what is meant by "high-level structure" (e.g., secondary structure, environmental association).

      We thank the reviewer for pointing out this ambiguity. In the original manuscript, the phrase “high-level structure” was not sufficiently precise. In the revised version, we have clarified that this refers primarily to higher-order structural properties of DNA molecules, such as secondary structure, local conformational features, and sequence-dependent interactions with minerals or organic matter in soil. These characteristics may influence the accessibility of extracellular DNA to nucleases and thus affect degradation rates. We have revised the text accordingly to improve clarity and precision.

      L84-104

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequences (ASVs) under identical soil and incubation conditions. The potential variations in sequence-specific eDNA degradation rates can be attributed to several factors. First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C: N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021). Third, the persistence of soil DNA is often associated with its adsorption and protection by minerals and humus in soils (Cai et al., 2006b; Vuillemin et al., 2017; McKinney and Dungan, 2020). Thus, sequence-dependent differences in the physicochemical behavior of DNA molecules, including their affinity for soil minerals and organic matter, may also contribute to variation in degradation rates among sequences (Levy-Booth et al., 2007; Morrissey et al., 2015). Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (4) Line 111: The hypothesis is not clearly linked to the rationale. If sequence-specific degradation is expected, clarify whether it relates to conserved vs variable regions or structural features (e.g., stems vs loops).

      We thank the reviewer for this helpful comment. We agree that the original manuscript did not clearly link the hypothesis regarding sequence-specific degradation to its mechanistic rationale. In the revised manuscript, we have clarified that the expectation of sequence-specific degradation is not simply based on conserved vs variable regions of the 16S rRNA gene, but rather on the potential for sequence differences to influence intrinsic physicochemical properties, including base composition, local conformational features, potential secondary structures, and motif-dependent nuclease susceptibility. These factors may alter DNA accessibility to extracellular nucleases, providing a mechanistic basis for sequence-specific degradation. This clarification is now reflected in the Introduction and linked to the formal hypothesis statement.

      To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses (H1–H3) immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation.

      L84-104

      “In this study, “sequence‑specific degradation” refers to statistically significant differences in first‑order degradation rate constants (k, day<sup>⁻¹</sup>) among distinct 16S rRNA gene amplicon sequences (ASVs) under identical soil and incubation conditions. The potential variations in sequence-specific eDNA degradation rates can be attributed to several factors. First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C:N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021). Third, the persistence of soil DNA is often associated with its adsorption and protection by minerals and humus in soils (Cai et al., 2006b; Vuillemin et al., 2017; McKinney and Dungan, 2020). Thus, sequence-dependent differences in the physicochemical behavior of DNA molecules, including their affinity for soil minerals and organic matter, may also contribute to variation in degradation rates among sequences (Levy-Booth et al., 2007; Morrissey et al., 2015). Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (5) Lines 310-311: Clearly indicate which portion of the primers corresponds to the modified (GAPDH-derived) sequence. Provide full annotated primer sequences.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we now clearly indicate which portion of the forward primer corresponds to the GAPDH-derived synthetic tag and which portion corresponds to the 16S rRNA gene primer sequence. We have also provided the full annotated primer sequences in the Methods section to avoid ambiguity.

      Specifically, the modified forward primer is now described as:

      GAPDH-F-515F: 5′-CAT TGG CAA TGA GCG GTT C-GTG CCA GCM GCC GCG GTA A-3′,

      where CAT TGG CAA TGA GCG GTT C is the GAPDH-derived synthetic tag and GTG CCA GCM GCC GCG GTA A is the 16S rRNA gene forward primer sequence (515F).

      The reverse primer is:

      806R: 5′-GGA CTA CHV GGG TWT CTA AT-3′.

      L354-359

      “Briefly, exogenous eDNA was prepared by PCR amplification using a modified forward primer consisting of a GAPDH F tag fused to the 16S rRNA gene primer 515F, together with the reverse primer 806R. The full primer sequences were as follows: GAPDH-F-515F: 5'-CAT TGG CAA TGA GCG GTT C-GTG CCA GCM GCC GCG GTA A-3', in which CAT TGG CAA TGA GCG GTT C represents the GAPDH F tag and GTG CCA GCM GCC GCG GTA A represents the 16S rRNA gene forward primer sequence (515F); and 806R: 5'-GGA CTA CHV GGG TWT CTA AT-3'.”

      (6) Lines 310-311: Explicitly define GAPDH and justify its use as a synthetic tag.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we now explicitly define GAPDH as glyceraldehyde-3-phosphate dehydrogenase, a human housekeeping gene. Specifically, the GAPDH-derived sequence was selected for two reasons. First, it is highly divergent from known soil microbial 16S rRNA gene sequences and did not produce detectable amplification when tested with soil DNA using the GAPDH tagged 806R primer pair, indicating that it would not interfere with endogenous soil DNA signals. Second, its melting temperature was compatible with that of the reverse primer, which allowed stable amplification of the tagged 16S amplicons under our PCR conditions.

      L360-371

      “GAPDH is a primer for a human housekeeping gene and it has no homologous sequences in soils. Subsequently, GAPDH was selected as the label primer based on two criteria. First, this primer was selected to avoid interference from the original soil sequences (Huang et al., 2014; Yang et al., 2021; Arvizu-Hernandez et al., 2025), and no detectable PCR amplification was observed for the primer set GAPDH F-806R across all the soil DNA samples included in this study. Second, the melting temperature (Tm) value of GAPDH F approximately matched that of 806R. The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing”

      (7) Line 346: Start a new paragraph to clearly separate this as a distinct experiment.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have started a new paragraph.

      (8) Line 346: Specify the number of samples analyzed for consistency.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have now explicitly specified the number of samples.

      A total of 120 samples were analyzed in this moisture gradient experiment: 2 ecosystems (Kaiyuan and Dashanbao) × 5 moisture levels (10%, 25%, 50%, 75%, and 100% of water holding capacity) × 6 incubation time points (0, 1, 3, 6, 12, and 24 days) × 2 replicates. Just two technical replicates were performed for this validation experiment, as the primary aim was to assess the trend of moisture effects rather than statistical inference across replicates.

      L408-410

      “This complementary experiment included two sites, five moisture levels, six incubation time points, and two replicates per treatment combination, resulting in a total of 120 soil samples.”

      (9) Lines 351-352: Replace "harvested" with "collected."

      We have replaced “harvested” with “collected” as suggested.

      (10) Line 367: Clarify how Illumina adapters and indices were added (e.g., two-step PCR, fusion primers).

      We thank the reviewer for this helpful suggestion. We have revised the Methods section to clarify how Illumina adapters and indices were incorporated. We used a pooled amplicon library preparation strategy. Individual samples were first amplified with primers containing sample-specific barcode sequences. The barcoded amplicons from multiple samples were then pooled and used for library preparation with the ALFA-SEQ DNA Library Prep Kit. Universal Illumina-compatible adapters were first ligated to the pooled amplicons. After bead-based purification, an indexing PCR was performed using the index primer mix, which introduced the complete P5/P7 sequences and a library-level Illumina index into the library molecules. Thus, sample demultiplexing was based on the sample-specific barcodes introduced during amplicon PCR, whereas the Illumina index was used to identify the pooled sequencing library. We have clarified this procedure in the revised manuscript.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).

      (11) Provide more detail on chimera removal, filtering thresholds, and normalization choices.

      We thank the reviewer for this helpful suggestion. The raw paired-end reads were first merged, and primer sequences were removed using the search_pcr2 script in USEARCH. Reads with more than two primer mismatches were discarded. Quality filtering was then performed using fastq_filter, and sequences with quality scores below 20 were removed. Redundant reads were collapsed using fastx_uniques. Amplicon sequence variants (ASVs) were generated using the UNOISE3 denoising algorithm, which also performs built-in chimaera filtering during ASV inference. In addition, ASVs with total sequence counts fewer than 9 were excluded to reduce the influence of low-frequency noise.

      L443-460

      “Briefly, paired-end reads were merged using USEARCH, and primer sequences (GAPDH-F-515F and 806R) were removed using the search_pcr2 script. Reads with more than two primer mismatches were discarded. Quality filtering was performed using the fastq_filter script, and sequences with quality scores below 20 were removed. Redundant sequences were dereplicated using the fastx_uniques script. ASVs were generated using the UNOISE3 non‑clustering denoising algorithm (Edgar, 2016), which infers 100% exact sequence variants by distinguishing biological sequences from PCR/sequencing errors. ASVs with total sequence counts fewer than 9 across all samples were removed to reduce noise. To quantify the abundance of each ASV, an ASV table was generated by mapping the quality‑filtered raw reads back to the ASV set using the otutab command. A 97% similarity threshold was applied for this recruitment to accommodate stochastic sequencing noise while maintaining biological resolution. Crucially, the mapping followed a best-hit priority rule, where each read was assigned to the ASV with the highest per cent identity within the 97% radius. This approach ensures that reads derived from the same biological template are accurately counted toward their respective ASV, preventing the underestimation of abundances that would occur with exact matching while strictly preserving the single-nucleotide resolution of the ASV framework. Taxonomic annotation of the ASVs was performed in QIIME2 with the Silva v138 database. A total of 89322 prokaryotic ASVs were obtained. To standardize sequencing depth across samples, the read number of each sample was rarefied to 53251 using the rarefy function in the vegan package in R.”

      (12) Line 412: Rephrase to refer to 16S amplicon addition rather than 16S rRNA genes (along the whole text), as only the V4 region is analyzed.

      We thank the reviewer for this helpful suggestion. we have rephrased references to “16S rRNA genes” to “16S rRNA gene amplicon fragments”

      (13) Ensure consistent primer naming throughout (e.g., GAPDH F vs ACTF).

      We have checked the entire manuscript and confirm that only “GAPDH F” is used as the label primer.

      (14) Finally, the manuscript would benefit from careful language editing. Several typographical errors, grammatical inconsistencies, and unclear phrases are present throughout. Examples include:

      Misspellings such as "diffrence" (e.g., figure legends) and inconsistent capitalization. Inconsistent terminology (e.g., "genes," "amplicons," and "fragments" used interchangeably without clarification). Redundant or awkward phrasing (e.g., repeated use of "extracellular 16S rRNA genes"). Occasional subject-verb agreement issues and missing articles.

      We sincerely apologize for the language issues. The manuscript has now undergone a thorough language editing process by a native English‑speaking colleague.

      Recommendation

      Major revision: The manuscript addresses an important problem and presents a promising approach. However, key issues related to conceptual clarity, bioinformatic consistency, statistical rigor, and interpretation of PMA-based results must be resolved. With substantial revision and clarification, the study has the potential to make a meaningful contribution to the field.

      We sincerely thank the reviewer for the thorough, constructive, and critical evaluation of our manuscript. We greatly appreciate the recognition that our study addresses an important problem and presents a promising approach. We also acknowledge the key issues raised regarding conceptual clarity, bioinformatic consistency, statistical rigor, and interpretation of PMA‑based results. We have taken these comments very seriously and have substantially revised the manuscript accordingly, more details about the revisions are described in the following point-by-point responses.

      Reviewer #2 (Recommendations for the authors):

      Editorial comments:

      (1) Title: I recommend removing "across China" from the title. In many ways, the study has nothing to do specifically with China, and you limit the broad applicability of the study. The same work could have been done with soils from Africa, for example. Also, it might be ok to remove 16S rRNA as well. The 16S rRNA genes are a proxy for rates of extracellular DNA degradation, but the study isn't exactly about 16S either.

      We thank the reviewer for this thoughtful suggestion regarding the title. We have revised the title to “The overall and sequence-specific degradation of soil extracellular DNA fragments: rates and influential factors.”

      (3) L44-45: "...such as real-time PCR, high-throughput amplicon sequencing, and metagenomic analysis...".

      We thank the reviewer for this suggestion. We have revised the order according to the suggestions of the reviewer.

      L44-45

      “The investigation of soil microbial abundance and diversity heavily relies on DNA-based technologies, such as real-time PCR, high-throughput amplicon sequencing, and metagenomic analysis.”

      (4) L48: remove "they".

      We agree with the reviewer and have removed the extraneous “they”.

      (5) L51: "noise factor"; "...persistence can lead to...".

      We have revised the sentence as suggested.

      (6) L53: remove theoretical.

      We have removed “theoretical”.

      (7) L58: remove "the".

      We have removed "the".

      (8) L86: Is restriction digestion of DNA a likely extracellular process in soil?

      We thank the reviewer for this thoughtful comment. We agree that the original wording may have overstated the likelihood of classical restriction digestion as a dominant extracellular process in soils. Our intention was not to suggest that intracellular restriction enzyme systems operate directly in the soil matrix in the same manner as they do within living cells. Rather, we aimed to indicate more generally that sequence-dependent nuclease susceptibility could contribute to differential degradation among extracellular DNA fragments.

      L87-95

      “First, sequence-dependent degradation can arise from differences in nucleotide composition, particularly GC content. This influences the thermodynamic stability and base-stacking interactions of the DNA duplex, thereby altering its accessibility to extracellular nucleases (Marrone and Ballantyne, 2008; Wolpe and Guertin, 2022). Second, local conformational features and the formation of potential secondary structures, such as stem-loops or hairpins, can create steric hindrance that protects the phosphodiester backbone. Differences in base composition also alter the elemental stoichiometry (e.g., C: N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021).”

      (9) L94-99: The authors might also consider the different nitrogen content of different bases; this might also affect sequence-specific selection of DNA for degradation.

      We thank the reviewer for this insightful suggestion. We agree that differences in the elemental composition of DNA bases, including nitrogen content, may provide an additional mechanistic explanation for sequence-dependent degradation. In the revised manuscript, we have incorporated this point into the Introduction.

      L92-95

      “Differences in base composition also alter the elemental stoichiometry (e.g., C:N ratio) of DNA molecules, potentially affecting microbial preference for recycling specific sequences as nutrient sources (Cai et al., 2006a; Buitrago et al., 2021).”

      (10) L110-112: These are not really written in hypothesis form. Also, what about a hypothesis about degradation rates and soil type/temperature/moisture?

      We thank the reviewer for this constructive critique. We have rewritten the hypotheses. To improve the logical flow of the manuscript, we have restructured the Introduction by moving the three central hypotheses immediately following the discussion of the biochemical mechanisms underlying sequence-specific degradation. This adjustment ensures that the hypotheses are directly grounded in the theoretical framework.

      L99-104

      “Consequently, we proposed three central hypotheses. (1) The degradation rates of eDNA amplicon fragments were expected to be highly sequence‑specific. (2) The rates and patterns of eDNA fragments degradation would be influenced by environmental factors such as temperature and moisture content. (3) The sequence‑specific degradation of extracellular 16S rRNA gene amplicon fragments would significantly influence estimates of soil prokaryotic abundance and diversity.”

      (11) L116: "GAPDH F-labeled 16S rRNA gene amplicon fragments....".

      We thank the reviewer for this helpful suggestion. we have rephrased references to “16S rRNA genes” to “16S rRNA gene amplicon fragments”

      (12) L117: "rapidly".

      We agree with the reviewer and have revised.

      (13) L118-120: "After a 48-day incubation period, 0.2 to 3.1% of the initial spike GADPH F-labeled 16S rRNA gene amplicon fragments ...".

      We agree with the reviewer and have revised.

      (14) L125: Spell out SEM in first usage.

      We thank the reviewer for this suggestion. In the revised manuscript, we have spelled out SEM as Structural equation modeling.

      (15) L128: I don't like the idea of putting this Figure in supplemental materials.

      We thank the reviewer for this suggestion. We have moved Figure S2 (moisture gradient microcosm experiment) to the main text as Figure 1f.

      (16) L154: The term "intracellular prokaryotic abundance" is not the right term. This makes one think of an intracellular parasite. I think you want something like: "Approximately 40% of sequences in total soil DNA extraction NGS amplicon libraries were derived from intact cells, while the remaining represented extracellular DNA. Conversely, greater than 80% of observed richness was derived from intact cells." (Please check that I stated this correctly.) I would also suggest some statistics or ranges here.

      We thank the reviewer for this important terminological clarification. We agree that the term “intracellular prokaryotic abundance” is misleading, as it could imply intracellular parasites. In the revised manuscript, we have replaced this with a clearer description and We have also added the across‑site ranges to provide statistical context.

      L163-166

      “The PMA treatment revealed that intact cells accounted for approximately 40% (range: 9–73%) of the total 16S rRNA gene copies. In contrast, over 80% (range: 27–97%) of the observed ASV richness was associated with sequences originating from intact cells (Fig. 4a and b).”

      (17) L168: "...a significant NEGATIVE correlation was observed...".

      We agree with the reviewer and have revised.

      (18) L169: "However, no significant relationship was observed...".

      We agree with the reviewer and have revised.

      (19) L194-195: What about pH and temperature?

      We thank the reviewer for this comment. We agree that pH and temperature are important environmental factors that can influence microbial DNA degradation and community composition. However, our results (Fig. 1c) indicate that soil moisture is the most dominant factor affecting extracellular DNA degradation. Therefore, in the revised manuscript, we have focused the explanation primarily on soil moisture, while acknowledging that pH and temperature may also be important influencing factors.

      L208-211

      “Third, environmental factors, including soil moisture, pH, and temperature, can predominantly govern enzymatic reaction rates (He et al., 2024; Shah et al., 2024). Indeed, strong positive correlations were observed between moisture content and eDNA degradation rates in both the survey and microcosm experiments (Fig. 1d-f).”

      (20) L199: "findings".

      We have revised as suggested.

      (21) L227-229: This sounds more like results.

      We thank the reviewer for this comment. We agree that the original first sentence in L227–229 reads more like results. Our intention was to introduce the discussion by linking extracellular DNA to potential impacts on prokaryotic community analysis, rather than to present specific findings at this point. We have reorganized this section as follows.

      L246-248

      “Accordingly, we further explored how DNA may influence prokaryotic community analyses using PMA treatment, and significant disparities were observed between the profiles of the total and PMA-treated soil prokaryotic communities (Fig. 4).”

      (22) L230: Need to also consider differential cell lysis during DNA extraction.

      We thank the reviewer for this important comment. We agree that differential cell lysis during DNA extraction could influence the observed community profiles, as microbial taxa differ in cell wall composition and resistance to mechanical or chemical lysis. In the revised manuscript, we explicitly acknowledge this limitation in the relevant section. We also clarify that a standardized DNA extraction protocol (DNeasy PowerSoil kit) was used to efficiently lyse a broad range of microbial taxa, but some taxon-specific lysis bias may remain. Future studies could combine multiple lysis methods or spike-in controls to quantify and correct for potential extraction bias.

      L262-265

      “However, as DNA extraction efficiency may differ between intact cells and eDNA, the actual differences between total and living prokaryotic abundance could be smaller than those observed in this study. Similarly, the overestimated prokaryotic richness may arise from historically accumulated microbial taxonomic information stored in eDNA pools (Deshpande and Fahrenfeld, 2023; Wang et al., 2024).”

      L309-311

      “Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015).”

      (23) L232: Need to also consider that PMA treatment is not perfect and can be affected by substrate, the ability of light to access DNA for crosslinking, etc.

      We thank the reviewer for this important reminder. We agree that PMA treatment is not perfect and that its efficiency can be affected by soil matrix properties (e.g., organic matter, clay minerals) and the ability of light to penetrate the sample for DNA crosslinking. In the revised manuscript, we have explicitly acknowledged these limitations in the discussion.

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (24) L240: Can extracellular DNA have an ecological role?

      We thank the reviewer for this thoughtful question. Yes, extracellular DNA (eDNA) does have important ecological roles beyond being a potential bias in molecular analyses. In the revised manuscript, we have added statements to highlight that extracellular DNA can serve as a nutrient source (e.g., nitrogen and phosphorus) for microbes and may also contribute to horizontal gene transfer. This emphasizes that extracellular DNA may actively influence microbial community structure and function, in addition to its role in potentially inflating observed abundance and richness.

      L267-275

      “We observed a significant correlation between eDNA degradation rates and the overall structure of the prokaryotic community, but this relationship was absent in PMA-treated communities (Fig. 5b). This discrepancy highlights the divergent ecological roles of extracellular and intracellular DNA. Analyses of the total community integrate intracellular DNA from metabolically active cells with eDNA which primarily originates from historical microbial residues (Lennon et al., 2018). EDNA incorporates signals that likely reflect the legacy effects of past environmental conditions (Wang et al., 2021). In contrast, the PMA-treated community reflects transient microbial activity driven by current selective pressures. Additionally, eDNA can serve as a nutrient source and facilitate horizontal gene transfer, which may further shape its interactions with contemporary microbial communities (Levy-Booth et al., 2007).”

      (25) L256: Why would microorganisms selectively degrade one DNA sequence vs another? This seems to be likely to be stochastic in terms of which sequences are taken up by microorganisms. However, different DNA sequences might hydrolyze differently or be otherwise damaged, and that could lead to differential degradation of a viable amplicon. It might be interesting to incorporate long pieces of DNA with different internal primer sites and use quantitative PCR to determine how sequences are degrading.

      We thank the reviewer for this important mechanistic insight. We agree that the observed correlation between degradation rate and sequence abundance does not necessarily imply active microbial preference. It could equally reflect stochastic encounter rates or intrinsic chemical differences (e.g., AT‑rich regions hydrolyzing faster). We have revised the corresponding paragraph in the Discussion.

      L279-292

      “This finding suggests that abundant eDNA degrades at a faster rate compared to rare eDNA. As mentioned earlier, this could be explained by several mechanisms. First, as soil eDNA is subject to enzymatic degradation and microbial recycling, abundant DNA sequences may be more likely to be encountered and degraded by extracellular nucleases simply due to their higher copy numbers (Levy-Booth et al., 2007; Nagler et al., 2018). Similarly, if microbes preferentially take up DNA as a nutrient source, they may degrade abundant sequences more frequently as a stochastic consequence of higher encounter rates (Finkel and Kolter, 2001). However, we also found that the relationships between the sequence-specific degradation rates and the effect sizes of extracellular 16S rRNA gene amplicon fragments varied across the study sites (Fig. S1g). The sequence-specific effect sizes of extracellular 16S rRNA gene amplicon fragments are mainly determined by both their production and degradation rates (Pietramellara et al., 2009; Sirois and Buckley, 2019). These inconsistent correlations emphasize the critical role played by the production rates of extracellular 16S rRNA genes in influencing the analysis of prokaryotic communities. Therefore, future studies should systematically determine both the production and degradation rates of eDNA.”

      (26) L282-283: This belongs in the discussion.

      We agree with the reviewer and have revised accordingly.

      (27) L289: "as well as measurements of total organic carbon".

      We agree with the reviewer and have revised accordingly.

      (28) L338: Any water content for these soils?

      We thank the reviewer for this comment. The water contents of soils from all study sites are reported in Supplementary Table 2.

      (29) L349-350: You mean that you measured the total soil extracted DNA and then added 1% as labeled 16S?

      Yes, for each soil sample, we extracted total soil DNA and quantified its concentration (ng DNA per gram of soil). We then added exogenous GAPDH‑tagged 16S amplicon fragments at an amount equal to 1% of this total DNA concentration. This concentration was chosen to mimic a realistic pulse of extracellular DNA input without overwhelming the endogenous DNA pool. We apologize for any confusion caused by the imprecise wording in the original manuscript.

      L392-398

      “The microcosm experiment was conducted using 30 g of soil for each sample. After pre-incubation at 20℃ for one week, each soil was thoroughly mixed with the GAPDH F‑tagged 16S rRNA gene amplicon fragments and incubated further at 20℃ (Fig. S8). The amount of exogenous GAPDH F‑tagged 16S rRNA gene amplicon fragments added to each soil sample was equivalent to 1% of the total DNA concentration naturally present in that soil, as determined fluorometrically prior to the experiment. This concentration was chosen to approximate natural eDNA fluxes resulting from microbial lysis, ensuring experimental relevance to in situ conditions (Table S2).”

      (30) L354: Remember that soil recovery from intact cells is going to be lower than for extracellular DNA. So, you are probably overestimating the contribution of extracellular DNA to the total DNA in the system.

      We thank the reviewer for this comment. We agree that DNA recovery from intact cells is generally lower than from extracellular DNA due to differential cell lysis efficiencies. Consequently, the contribution of extracellular DNA to total soil DNA may be somewhat overestimated in our study. We have clarified this limitation in the revised manuscript.

      L262-265

      “However, as DNA extraction efficiency may differ between intact cells and eDNA, the actual differences between total and living prokaryotic abundance could be smaller than those observed in this study. Similarly, the overestimated prokaryotic richness may arise from historically accumulated microbial taxonomic information stored in eDNA pools.”

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (31) L362: Amplification efficiency is pretty low. I think you would have been better served with GAPDH on both ends, and that would have given you a much higher efficiency qPCR.

      We thank the reviewer for this comment. The actual qPCR amplification efficiency in our assay was approximately 85%, which, although slightly below the ideal range, was still acceptable and produced reproducible amplification curves and reliable quantification for degradation-rate calculations.

      We acknowledge that the amplification efficiency in our qPCR experiments using a GAPDH F-labeled 16S primer on one end was suboptimal. The current design used a single GAPDH tag at the forward primer to avoid potential amplification bias or primer-dimer formation that could arise from extending the degenerate reverse primer. In addition, dual-end labeling would have required synthesis of new barcode-labeled tagged primers, increasing both cost and experimental complexity. Thanks again for the constructive comments, which provided us with the direction for future experiment optimization.

      L365-371

      “The GAPDH was incorporated only into the forward primer for several reasons. Methodologically, adding a long linker to the degenerate reverse primer (806R) could reduce amplification efficiency or introduce bias. Economically, single-end labeling allowed us to use the standard reverse primer already carrying sample-specific barcodes, avoiding the costly synthesis of a full set of dual-labeled barcoded primers. This design minimized the risk of secondary structure and primer-dimer artifacts while maintaining sufficient specificity and compatibility with downstream qPCR and sequencing.”

      (32) L367: Not enough detail on how barcoded libraries were made. UDIs?

      We thank the reviewer for this helpful comment. We have now clarified the library preparation and indexing strategy in the revised Methods section. This amplicon diversity sequencing used a pooled-library strategy. Individual samples were first distinguished by sample-specific inline barcodes introduced during the amplicon PCR step. After amplification, barcoded PCR products from multiple samples were pooled and subjected to library construction using the ALFA-SEQ DNA Library Prep Kit. Universal Illumina-compatible adapters were ligated to the pooled amplicons, followed by an indexing PCR that introduced the complete P5/P7 sequences and a library-level Illumina index. Thus, the Illumina index was used to identify the pooled sequencing library, whereas sample demultiplexing was performed according to the sample-specific inline barcodes. We have revised the Methods section to make this procedure explicit.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).”

      (33) L368: Why was the # of cycles increased?

      Thank you for your question. In the original manuscript (L368), we stated that the number of PCR cycles was increased to 35. This was mainly because the exogenously added GAPDH F‑labeled 16S rRNA genes had a relatively low initial abundance in the soil and gradually degraded during the microcosm incubation, with their copy numbers becoming particularly low at the last time points (see Fig. 1a). To ensure sufficient PCR product for high‑throughput sequencing from samples at all time points (especially those with low abundance at later stages), we appropriately increased the cycle number to 35.

      L429-431

      “The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing.”

      (34) L372: Were sequencing adapters ligated onto the pool?

      We thank the reviewer for this question. Yes, in this amplicon diversity sequencing workflow, sequencing adapters were ligated onto the pooled amplicon products. Briefly, individual samples were first amplified with sample-specific barcode sequences, allowing each sample to be distinguished after sequencing. The barcoded PCR products from multiple samples were then pooled for library construction. Universal Illumina-compatible adapters were ligated to this pooled amplicon library using the ALFA-SEQ DNA Library Prep Kit. After adapter ligation and purification, an indexing PCR was performed to introduce the complete P5/P7 sequences and a library-level Illumina index. We have clarified this pooled-library construction workflow in the revised Methods section.

      L424-440

      “The community profiles of the GAPDH F-tagged 16S rRNA gene amplicon fragments were determined using high-throughput amplicon sequencing. Briefly, GAPDH F-tagged 16S rRNA gene amplicon fragments from the microcosm soils were first amplified from individual samples using GAPDH F and barcode-labeled 806R primers. The reverse primer 806R carried a 12-bp sample-specific barcode, whereas the GAPDH F primer did not contain a barcode. Therefore, each sample was assigned a unique barcode during PCR, which allowed sample demultiplexing after sequencing. The PCR reaction system and thermal cycling conditions were similar to those described above, except that the number of amplification cycles was increased to 35 to obtain sufficient amplicon products for sequencing. The barcoded PCR products from individual samples were purified using a GeneJET Gel Extraction Kit (Thermo Scientific, Lithuania), quantified, and then pooled in equimolar amounts for subsequent library construction. Sequencing libraries were prepared from the pooled barcoded amplicons using the ALFA-SEQ DNA Library Prep Kit according to the manufacturer’s protocol. Universal Illumina-compatible adapters were first ligated to the pooled amplicon products, followed by bead-based purification. An indexing PCR was then performed using the index primer mix, which introduced the complete P5/P7 flow-cell binding sequences and a library-level Illumina index into the pooled library molecules. The indexed library was purified, quantified, and subjected to paired-end sequencing on the NovaSeq platform at MAGIGENE Co., Ltd. (Guangzhou, China).”

      (35) L378: "Amplicon sequence variants".

      We agree with the reviewer and have revised accordingly.

      (36) L380: Why were ASVs with fewer than 9 reads removed?

      We thank the reviewer for this question. The threshold of removing ASVs with fewer than 9 total reads across all samples was applied to reduce noise from sequencing errors and PCR artifacts. Our justification is supported by both the default parameters of the UNOISE3 algorithm and common practice in amplicon sequencing analysis.

      The USEARCH manual specifies that the -minsize parameter in the unoise3 command defaults to 8. This means that unique sequences occurring fewer than 8 times are discarded by the algorithm during ASV inference, as they are unlikely to represent true biological variants. Our threshold of 9 is slightly more conservative than the default (9 > 8), ensuring that only ASVs with a minimal level of abundance are retained. This choice is directly aligned with the algorithm’s intrinsic noise‑filtering logic.

      (37) L402: Please don't forget to discuss that PCR bias can contribute to uncertainty in the abundance of each taxon.

      Thank you for this important reminder. We agree that PCR bias (e.g., primer‑template mismatches, GC content differences, and variable amplification efficiency) can contribute to uncertainty in the abundance estimates of each taxon. Following your suggestion, we have now added a paragraph in the Discussion section to address this issue. We state that sequence‑specific degradation rates and PCR bias may jointly affect the accuracy of taxon abundance estimates, and future studies should incorporate internal standards or multiplex PCR strategies to correct for such biases. Thank you for your careful review.

      L294-305

      “Despite the high-resolution insights afforded by our methodology, several limitations should be considered. First, utilizing PCR-amplified 16S rRNA gene fragments as proxies oversimplifies the structural and sequence complexity of natural soil eDNA pools. In natural environments, eDNA varies widely in fragment length and conformation, and exhibits complex interactions with mineral surfaces, all of which fundamentally affect degradation dynamics (Levy-Booth et al., 2007; McKinney and Dungan, 2020). Additionally, the highly conserved nature of the 16S rRNA gene means that the nucleotide variability explored here (e.g., GC content gradients) does not fully capture the genomic heterogeneity of entire metagenomes (Knight et al., 2018). Consequently, our reported degradation rates indicate the decay potential of highly accessible linear eDNA rather than a universal rate for all soil DNA fractions. Future studies incorporating diverse metagenomic DNA, especially those with extreme AT or GC contents, are essential for building a more generalizable predictive framework for eDNA persistence (Morrissey et al., 2015).”

      (38) L414: Suggest: "To inhibit amplification of extracellular DNA, soils were incubated with propidium monoazide (PMA), as described previously (REF). Briefly, soil (X grams) was mixed with PMA in a total volume of Y (ml).

      We thank the reviewer for this suggestion. We have revised the Methods section to provide a clearer description of PMA treatment, specifying the soil amount (0.50 g) and the total volume (0.5 mL).

      L496-497

      “To inhibit amplification of eDNA, soils were incubated with PMA, as described previously (Carini et al., 2016).”

      L505-506

      “In this study, 0.50 g of soil was mixed with PMA in a total volume of 0.5 mL (40 µM PMA in phosphate‑buffered saline, PBS), while the control soil samples were mixed with PBS without PMA.”

      (39) L416: In contrast, microbes with intact cell membranes exclude PMA, and their DNA is not cross-linked with PMA, and remains amenable to PCR amplification.

      We agree and have revised.

      (40) L418-420: wording/sentence is strange and needs work.

      Thank you for pointing this out. We have reviewed the sentence at L418‑420 and agree that the wording is awkward. Moreover, the content only listed the advantages of the PMA method without acknowledging its limitations, making the statement less balanced. Therefore, in the revised manuscript, we have deleted this sentence. The limitations of the PMA method have been addressed in the Discussion section.

      L502-503

      “Currently, PMA treatment is a widely used to suppress PCR amplification of eDNA (Xue et al., 2023; Canini et al., 2024).”

      L305-314

      “Second, methodological biases inherent in quantifying the intracellular community must be acknowledged (Du et al., 2025). Although PMA treatment is widely used to exclude eDNA, its efficiency in complex soil matrices can be compromised by limited light penetration in turbid suspensions and competitive adsorption to soil particles (Nocker et al., 2007; Carini et al., 2016; Heise et al., 2016). Compounding this issue, downstream DNA recovery is subject to differential cell lysis, as taxa with robust cell walls (e.g., Gram-positive bacteria) may resist extraction (Frostegård et al., 1999; Albertsen et al., 2015). While our standardized bead-beating protocol and calculation of degradation rate constants (k) minimize systematic biases, future studies should integrate complementary viability markers (e.g., RNA-based analyses or protein synthesis activity probes) and multi-extraction comparisons to robustly validate these ecological patterns (Emerson et al., 2017).”

      (41) L421-422: PMA treatment is a widely used method for inhibiting the enzymatic processing of extracellular DNA (Xue, Canini).

      We agree and have revised.

      (42) L425: include volume of PBA.

      We thank the reviewer for this comment. We have revised the Methods section to include the volume of PMA used

      L505-506

      “In this study, 0.50 g of soil was mixed with PMA in a total volume of 0.5 mL (40 µM PMA in phosphate‑buffered saline, PBS).”

      (43) L429-430: Don't use the word precipitates- use "pellets".

      We agree and have revised.

      (44) L433: "The abundance of 16S rRNA genes was determined using quantitative PCR employing a LightCycler...".

      We agree and have revised.

      (45) L445-: Section 4.9 - needs citations for PERMANOVA, NMDS, SEM, etc.

      Thank you for your suggestion. We have added the necessary citations for PERMANOVA, NMDS, SEM, and other methods in Section 4.9.

      L531-539

      Prokaryotic community structure differences among the study sites and incubation time points were examined through non-metric multidimensional scaling analysis (NMDS), permutation multivariate analysis of variance (PERMANOVA), and Permutational Analysis of Multivariate Dispersion (PERMDISP) (Kruskal, 1964; Anderson, 2001). Random forest modeling was conducted to assess the importance of environmental and soil variables in predicting the overall degradation rates of extracellular 16S rRNA gene amplicon fragments. Structural equation modeling (SEM) was employed to further evaluate the direct and indirect effects of soil moisture, soil pH, MAP, and prokaryotic abundance on the overall degradation rates of extracellular 16S rRNA gene amplicon fragments (Grace, 2006).

      (46) L698: A few comments. It would be nice to know how many different 16S sequences were tracked for differential degradation and shown in the figure.

      We thank the reviewer for this helpful comment. We would like to clarify that Fig. 1A does not track the degradation of individual 16S rRNA gene amplicon sequences, but instead shows the overall degradation dynamics of the total added exogenous DNA pool. The data points are derived from total 16S gene copy numbers measured via qPCR at each incubation time point. Consequently, this quantification inherently includes all sequences present within the added pool. The multiple lines visualized in the figure represent the collective degradation trajectories of the entire DNA pool across different study sites

      To address sequence-level changes, we further analyzed the richness and composition of the GAPDH F-tagged 16S rRNA gene amplicon fragments, which are presented in Fig. 2A and related analyses.

      (47) L699: Better to use "16S rRNA gene amplicon fragment abundance" as the term.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we have replaced the original wording with “16S rRNA gene amplicon fragment abundance” where appropriate.

      (48) Y-axis for Figures 1A and 2A should be GAPDH-labeled, not ACTB-labeled.

      We apologize for this mistake. We have corrected this error in the revised manuscript.

      (49) For Figure 1b: Why not use box plots and ANOVA for different soil types?

      Thank you for your valuable suggestion. In the original Figure 1b, we used a bar plot to display the degradation rate constants across the 30 study sites. This choice was intended to emphasize the continuous variation among sites and their gradient relationships with environmental factors (e.g., soil moisture, MAP), which were then used in random forest and structural equation modeling. The bar plot better illustrates the spatial continuum of degradation rates rather than treating ecosystem types as discrete categories.

      Nevertheless, we fully agree that a boxplot grouped by ecosystem type (grassland, forest, cropland, desert) would help readers quickly grasp the overall differences among land‑use types. In the revised manuscript, we have added a boxplot grouped by ecosystem type and performed one‑way ANOVA followed by Tukey HSD post‑hoc tests (Fig. 1.). The results show that degradation rate constants differ significantly among ecosystem types (P < 0.05).

      L124-128

      “The degradation rate constants of the spiked extracellular 16S rRNA gene amplicon fragments displayed considerable variability among the study sites, ranging from 0.05 to 0.16 day<sup>-1</sup> (Fig. 1b). Furthermore, we found that degradation rate constants differed significantly among ecosystem types (Fig. 1c, P < 0.05). Specifically, cropland and forest soils exhibited significantly higher degradation rates than grassland soils (P < 0.05).”

      (50) For Figure 2: Where are PERMANOVA and PERMDISP values for the figure?

      We thank the reviewer for this comment. In the revised manuscript, we have added the PERMANOVA and PERMDISP values corresponding to Figure 2 in the figure legend and Results section (Fig. 2).

      (51) I found Figure 2b to be hard to see. The 48-day circles are almost invisible. Difficult to know what the authors are trying to show here, since there is so much variability associated with soil type.

      We thank the reviewer for this comment. Figure 2b is intended to illustrate the temporal changes in microbial community structure during the incubation. The different colored circles represent samples at different time points (1, 3, 6, 12, 24, and 48 days), showing how communities shift over time. We apologize that in the original Figure 2b, the 48‑day samples were nearly invisible and that the high variability among soil types obscured the intended message. In the revised manuscript, we have added a black border around every data point, which greatly enhances the visibility of the 48‑day samples (and all time points). We now use distinct shapes to represent different ecosystem types (grassland, forest, cropland, desert) in the NMDS ordination, and added PERMANOVA results in both the Results section and the figure legend (Fig. R2b).

      (52) Figure 4A: Y-axis need a label like "16S rRNA gene abundance".

      We agree and have revised.

      (53) Figure 4B: I'd like to see a Shannon index too, not just richness.

      Thank you for your suggestion. We agree that the Shannon index, which integrates both richness and evenness, provides a valuable complement to richness alone. In the revised manuscript, we added an analysis of the Shannon index to compare α‑diversity between total DNA (PMA‑untreated) and intact cell DNA (PMA‑treated) samples (Fig.4).

      (54) Figure 4D: Would be good to have lines linking the intact cell vs total abundance. Also, what about a box plot of Bray-Curtis (or similar) dissimilarity between intact cell and total microbial analysis across the dataset?

      Thank you for your suggestions. Regarding the addition of connecting lines in Figure 4D, after careful consideration we decided not to add them for the following reason: the total and PMA-treated communities from the same site are already coded with the same color (different colors for different sites), which effectively indicates the pairing. Adding lines would greatly reduce readability due to dense overlapping lines, especially given the number of sites. Therefore, we kept the original color‑based pairing design.

      To address your second suggestion, we have added a bar plot showing the distribution of Bray‑Curtis dissimilarities between total (PMA‑untreated) and intact cell (PMA‑treated) communities across all study samples (Fig. R3d).

      (55) Figure 5B: What do correlations with p > 0.05 show? I would remove these from the image.

      We thank the reviewer for this suggestion. We agree that correlations with p > 0.05 do not represent statistically significant relationships and may cause confusion. In the revised manuscript, we have removed these non-significant correlations from Figure 5B.

      (56) Figure 6: "Incubations of 0, 3, 6, 12, 24, and 48 days".

      We agree with the reviewer and have revised as suggested.

      References

      Albertsen, M., Karst, S.M., Ziegler, A.S., Kirkegaard, R.H., Nielsen, P.H., 2015. Back to basics–the influence of DNA extraction and primer choice on phylogenetic analysis of activated sludge communities. PLoS One 10, e0132783.

      Anderson, M.J., 2001. A new method for non‐parametric multivariate analysis of variance. Austral Ecology 26, 32-46.

      Arvizu-Hernandez, E., Ocadiz-Delgado, R., Gariglio, P., 2025. E7HPV16 Oncogene and 17beta-Estradiol Stress Promote Oncogenic microRNA Expression Patterns, Cell Proliferation and Cervical Intraepithelial Neoplasia 1. Cell Biochemistry and Function 43, e70065.

      Buitrago, D., Labrador, M., Arcon, J.P., Lema, R., Flores, O., Esteve-Codina, A., Blanc, J., Villegas, N., Bellido, D., Gut, M., 2021. Impact of DNA methylation on 3D genome structure. Nature Communications 12, 3243.

      Cai, P., Huang, Q., Zhang, X., Chen, H., 2006a. Adsorption of DNA on clay minerals and various colloidal particles from an Alfisol. Soil Biology and Biochemistry 38, 471-476.

      Cai, P., Huang, Q.Y., Zhang, X.W., 2006b. Interactions of DNA with clay minerals and soil colloidal particles and protection against degradation by DNase. Environmental Science & Technology 40, 2971-2976.

      Carini, P., Marsden, P.J., Leff, J.W., Morgan, E.E., Strickland, M.S., Fierer, N., 2016. Relic DNA is abundant in soil and obscures estimates of soil microbial diversity. Nature Microbiology 2, 1-6.

      Deshpande, A.S., Fahrenfeld, N.L., 2023. Influence of DNA from non-viable sources on the riverine water and biofilm microbiome, resistome, mobilome, and resistance gene host assignments. Journal of Hazardous materials 446, 130743.

      Du, Y., Wang, Z., Liu, K., Chai, G., Chi, Y., Li, T., Duan, Y., Xia, T., Liu, D., Che, R., 2025. The performance of different methods in characterizing soil live prokaryotic diversity and abundance is highly variable. iMetaOmics, e70011.

      Emerson, J.B., Adams, R.I., Román, C.M.B., Brooks, B., Coil, D.A., Dahlhausen, K., Ganz, H.H., Hartmann, E.M., Hsu, T., Justice, N.B., 2017. Schrödinger’s microbes: tools for distinguishing the living from the dead in microbial ecosystems. Microbiome 5, 86.

      Finkel, S.E., Kolter, R., 2001. DNA as a nutrient: novel role for bacterial competence gene homologs. Journal of Bacteriology 183, 6288-6293.

      Frostegård, Å., Courtois, S., Ramisse, V., Clerc, S., Bernillon, D., Le Gall, F., Jeannin, P., Nesme, X., Simonet, P., 1999. Quantification of bias related to the extraction of DNA directly from soils. Applied and Environmental Microbiology 65, 5409-5420.

      Grace, J.B., 2006. Structural equation modeling and natural systems. Cambridge University Press.

      He, P., Li, L.-J., Dai, S.-S., Guo, X.-L., Nie, M., Yang, X., Kuzyakov, Y., 2024. Straw addition and low soil moisture decreased temperature sensitivity and activation energy of soil organic matter. Geoderma 442, 116802.

      Heise, J., Nega, M., Alawi, M., Wagner, D., 2016. Propidium monoazide treatment to distinguish between live and dead methanogens in pure cultures and environmental samples. Journal of Microbiological Methods 121, 11-23.

      Huang, C., Xie, D.C., Cui, J.J., Li, Q., Gao, Y., Xie, K.P., 2014. FOXM1c Promotes Pancreatic Cancer Epithelial-to-Mesenchymal Transition and Metastasis via Upregulation of Expression of the Urokinase Plasminogen Activator System. Clinical Cancer Research 20, 1477-1488.

      Knight, R., Vrbanac, A., Taylor, B.C., Aksenov, A., Callewaert, C., Debelius, J., Gonzalez, A., Kosciolek, T., McCall, L.-I., McDonald, D., 2018. Best practices for analysing microbiomes. Nature Reviews Microbiology 16, 410-422.

      Kruskal, J.B., 1964. Nonmetric multidimensional scaling: a numerical method. Psychometrika 29, 115-129.

      Lennon, J.T., Muscarella, M.E., Placella, S.A., Lehmkuhl, B.K., 2018. How, when, and where relic DNA affects microbial diversity. mbio 9, e00637-00618.

      Levy-Booth, D.J., Campbell, R.G., Gulden, R.H., Hart, M.M., Powell, J.R., Klironomos, J.N., Pauls, K.P., Swanton, C.J., Trevors, J.T., Dunfield, K.E., 2007. Cycling of extracellular DNA in the soil environment. Soil Biology and Biochemistry 39, 2977-2991.

      Liu, Q.H., Yuan, L., Li, Z.H., Leung, K.M.Y., Sheng, G.P., 2024. Natural organic matter enhances natural transformation of extracellular antibiotic resistance genes in sunlit water. Environmental Science & Technology 58, 17990-17998.

      Marrone, A., Ballantyne, J., 2008. Sequence Specificity of BAL 31 Nuclease for ssDNA Revealed by Synthetic Oligomer Substrates Containing Homopolymeric Guanine Tracts. PLoS One 3, e3595.

      McKinney, C.W., Dungan, R.S., 2020. Influence of environmental conditions on extracellular and intracellular antibiotic resistance genes in manure-amended soil: A microcosm study. Soil Science Society of America Journal 84, 747-759.

      Morrissey, E.M., McHugh, T.A., Preteska, L., Hayer, M., Dijkstra, P., Hungate, B.A., Schwartz, E., 2015. Dynamics of extracellular DNA decomposition and bacterial community composition in soil. Soil Biology and Biochemistry 86, 42-49.

      Nagler, M., Insam, H., Pietramellara, G., Ascher-Jenull, J., 2018. Extracellular DNA in natural environments: features, relevance and applications. Applied Microbiology and Biotechnology 102, 6343-6356.

      Nocker, A., Sossa-Fernandez, P., Burr, M.D., Camper, A.K., 2007. Use of propidium monoazide for live/dead distinction in microbial ecology. Applied and Environmental Microbiology 73, 5111-5117.

      Pietramellara, G., Ascher, J., Borgogni, F., Ceccherini, M., Guerri, G., Nannipieri, P., 2009. Extracellular DNA in soil and sediment: fate and ecological relevance. Biology and Fertility of Soils 45, 219-235.

      Shah, A., Huang, J., Han, T., Khan, M.N., Tadesse, K.A., Daba, N.A., Khan, S., Ullah, S., Sardar, M.F., Fahad, S., 2024. Impact of soil moisture regimes on greenhouse gas emissions, soil microbial biomass, and enzymatic activity in long-term fertilized paddy soil. Environmental Sciences Europe 36, 120.

      Sirois, S.H., Buckley, D.H., 2019. Factors governing extracellular DNA degradation dynamics in soil. Environmental Microbiology Reports 11, 173-184.

      Vuillemin, A., Horn, F., Alawi, M., Henny, C., Wagner, D., Crowe, S.A., Kallmeyer, J., 2017. Preservation and significance of extracellular DNA in ferruginous sediments from Lake Towuti, Indonesia. Frontiers in Microbiology 8, 1440.

      Wang, X., Ganzert, L., Bartholomaus, A., Amen, R., Yang, S., Guzman, C.M., Matus, F., Albornoz, M.F., Aburto, F., Oses-Pedraza, R., Friedl, T., Wagner, D., 2024. The effects of climate and soil depth on living and dead bacterial communities along a longitudinal gradient in Chile. The Science of the total environment 945, 173846.

      Wang, Y.-T., Yang, W.-J., Li, C.-L., Doudeva, L.G., Yuan, H.S., 2007. Structural basis for sequence-dependent DNA cleavage by nonspecific endonucleases. Nucleic Acids Research 35, 584-594.

      Wang, Y., Yan, Y., Thompson, K.N., Bae, S., Accorsi, E.K., Zhang, Y., Shen, J., Vlamakis, H., Hartmann, E.M., Huttenhower, C., 2021. Whole microbial community viability is not quantitatively reflected by propidium monoazide sequencing approach. Microbiome 9, 1-13.

      Wolpe, J., Guertin, M., 2022. Regional and Single Nucleotide Correction of Sequence Bias in Chromatin Accessibility Data. The FASEB Journal 36.

      Yang, A., Liu, X., Liu, P., Feng, Y.Z., Liu, H.B., Gao, S., Huo, L.M., Han, X.Y., Wang, J.R., Kong, W., 2021. LncRNA UCA1 promotes development of gastric cancer via the miR-145/MYO6 axis. Cellular & Molecular Biology Letters 26, 33.

      Ye, M., Zhang, Z., Sun, M., Shi, Y., 2022. Dynamics, gene transfer, and ecological function of intracellular and extracellular DNA in environmental microbiome. iMeta 1, e34.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This useful study uses creative scalp EEG decoding methods to attempt to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. However, the evidence supporting the conclusions is incomplete due to concerns with the experimental design and methodology. This paper would be of interest to researchers studying cognitive control and adaptive behavior, if the concerns raised in the reviews can be addressed satisfactorily.

      We thank the editors and the reviewers for their positive assessment and constructive feedback of our work, which led us to think more deeply about the conceptual and methodological aspects of this project and further strengthen the manuscript. Based on the comments, we included more control analyses and revised the manuscript accordingly. Please see below our responses to each comment raised in the reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. Two types of learned associations are characterized, one being associations between stimulus features and control states (SC), and the other being stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates, which are used to test properties of their dynamics.

      The conclusion is reached that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing SC and SR correlates are: (1) not entirely overlapping in cross-decoding; (2) simultaneously observed on average over trials in overlapping time bins; (3) independently correlate with RT; and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Thank you for acknowledging the strength in EEG decoding and design. We have addressed all your concerns raised below point by point.

      Weaknesses:

      I still have my concern from the first round that the decoders are overfit to temporally structured noise. As I wrote before, the SC and SR classes are highly confounded with phase (chunk of session). I do not see how the control analyses conducted in the revision adequately deal with this issue.

      In the figures, there are several hints that these decoders are biased. Unfortunately, the figures are also constructed in such a way that hides or diminishes the salience of the clues of bias. This bias and lack of transparency discourage trust in the methods and results.

      I have two main suggestions:

      (1) Run a new experiment with a design that properly supports this question.

      I don't make this suggestion lightly, and I understand that it may not be feasible to implement given constraints; but I feel that this suggestion is warranted. The desired inferences rely on successful identification of SC and SR representations. Solidly identifying SC and SR representations necessitates an experimental design wherein these variables are sufficiently orthogonalized, within-subject, from temporally structured noise. The experimental design reported in this paper unfortunately does not meet this bar, in my opinion (and the opinion of a colleague I solicited).

      An adequate design would have enough phases to properly support "cross-phase" cross-validation. Deconfounding temporal noise is a basic requirement for decoding analyses of EEG and fMRI data (see e.g., leave-one-run-out CV that is effectively necessary in fMRI; in my experience, EEG is not much different, when the decoded classes are blocked in time, as here). In a journal with a typical acceptance-based review process, this would be grounds for rejection.

      Please note that this issue of decoder bias would seem to weaken the rest of the downstream analyses that are based on the decoded values. For instance, if the decoders are biased, in the within-trial correlation analysis, how can we be sure that co-fluctuations along certain dimensions within their projected values are driven by signal or noise? A similar issue clouds the LMM decoding-RT correlations.

      We appreciate the reviewer’s concern with the potential confound of temporally structured noise (TSN) in the EEG data. As we understand it, TSN refers to a process that the noise structure drifts over time. It follows that noise structure should be more similar for temporally closer trials and that the TSN’s bias on decoding accuracy is stronger for test trials that are closer to the training data. In the previous round of revision, we conducted a control analysis that reduced the influence of TSN by maximizing the temporal distance between training and test data (the distance between the centers of the training and test data of the same SC/SR manipulation is about 400 trials given the experimental design) and showed comparable decoding accuracy with the main results. As the reviewer finds this analysis unconvincing, we reason that the reviewer believes that the TSN has a long-term effect, such that it remains relatively stable over time and can be picked up by trials temporally distant from the training data. With this assumption and the assumption that this effect may not be linear, constructing a theoretically unbiased decoder requires perfectly counter-balanced training data (i.e., for every training trial of class A that is X trials away from the test data, there must be a training trial of all other classes that is exactly X trials away from the test data). As we were unable to achieve such a perfect design, we chose not to run an additional experiment. Instead, we focused on testing whether and how much TSN systematically biased the reported decoding accuracy.

      Please note that the existence of TSN in the EEG data is not sufficient to rule that the decoding results are biased. As TSN is stronger for trials closer to each other, the idea that auto-correlation biases decoding results would predict a distance effect, such that if a test trial is closer to a training trial of the same trial type, the higher similarity in TSN between the training and test data would more strongly inflate the decoding accuracy of the test trial, resulting in a negative correlation between distance between a test trial and its closest training trial of the same type and the test trial’s decoding accuracy. To test this predicted negative correlation, in each fold and each repetition of the cross-validation reported in the SC-SC and SR-SR decoders in Fig. 4, we calculated the distance (mean=5.84 trials, SD=2.05, 5th percentile =2.87, 95th percentile=9.45, one trial = 2.4-2.6s) between each test trial and its closest training trial of the same trial type. This distance was used as the predictor to predict decoding accuracy in a linear regression. Note that even if the relation between distance and decoding accuracy is non-linear, the linear relation will be negative because the relation is monotonic (similarity in noise structure decreases monotonically with temporal distance between trials). Similarly, because the effect is monotonic, if a long-range effect exists, it should also exist in short-range and be picked up by the distance range in this analysis. The regression coefficient is averaged across cross-validation folds and repetitions for each subject to match how the decoding accuracy was reported in the main text. Finally, the averaged regression coefficient was tested against 0 using a one-sample t-test. This analysis was conducted at each time point (from -250ms to 1500ms) separately. As shown in the figure below, no time point exhibited the negative correlation as predicted by the auto-correlation account. An alternative explanation is that this result indicates that TSN remains stable over time. If this is the case, TSN will be shared by all trials and will be unable to bias decoding results. Together with the control analysis introduced previously, this new control analysis supports the notion that the decoding results are not inflated by TSN in the EEG data. We included all the control analyses in the revised manuscript (page 13-14). Please note that this analysis is specific for the present dataset and we strongly agree with the reviewer that TSN is a key confounding factor in EEG analysis in general and should be carefully addressed.

      Lastly, we understand the concern with the early onset of above-chance decoding accuracy. Here, we provide an explanation: because of the blocked design (i.e., participants performed hundreds of trials with the same SC/SR associations), it is possible the participants learned the associations and used them to guide proactive cognitive control. As proactive cognitive control is anticipatory and sustained (Braver, 2012; Khan et al., 2025), it may be able to be decoded early on a trial, or even before trial onset. In the revised manuscript, we discussed this account along with the TSN issue as a limitation of the current project and directions for future research (page 24).

      (2) Increase transparency in the reporting of results throughout main text.

      Please do not truncate stimulus-aligned timecourses at time=0. Displaying the baseline period is very useful to identify bias, that is, to verify that stimulus-dependent conditions cannot be decoded pre-stimulus. Bias is most expected to be revealed in the baseline interval when the data are NOT baseline-corrected, which is why I previously asked to see the results omitting baseline correction. (But also note that if the decoders are biased, baseline-correcting would not remove this bias; instead, it would spread it across the rest of the epoch, while the baseline interval would, on average, be centered at zero.)

      Please use a more standard p-value correction threshold, rather than Bonferroni-corrected p<0.001. This threshold is unusually conservative for this type of study. And yet, despite this conservativeness, stimulus-evoked information can be decoded from nearly every time bin, including at t=0. This does not encourage trust in the accuracy of these p-values. Instead, I suggest using permutation-based cluster correction, with corrected p<0.05. This is much more standard and would therefore allow for better comparison to many other studies.

      I don't think these things should be done as control analyses, tucked away in the supplemental materials, but instead should be done as a part of the figures in the main text -- including decoding, RSA, cross-trial correlations, and RT correlations.

      Thank you for your suggestions. we have added the baseline period from 200 to 0 ms prior to the stimulus onset in all the stimulus-locked analyses and tested the significance with cluster-based permutation test (cluster-forming threshold p < 0.001, cluster-level p < 0.05, (Collins & Frank, 2018)) in all the analyses including decoding, RSA, cross-trial correlations and RT correlations. The results showed similar patterns, and they are all reported in the main text (please see all the figures and page 30-32 in the main text).

      Other issues:

      Regarding the analysis of the within-trial correlation of RSA betas, and "Cai 2019" bias:<br /> The correction that authors perform in the revision -- estimating the correlation within the baseline time interval and subtracting this estimate from subsequent timepoints -- assumes that the "Cai 2019" bias is stationary. This is a fairly strong assumption, however, as this bias depends not only on the design matrix, but also on the structure of the noise (see the Cai paper), which can be non-stationary. No data were provided in support of stationarity. It seems safer and potentially more realistic to assume non-stationarity.

      This analysis was included in the supplemental material. However, given that the correlation analysis presented in the Results is subject to the "Cai 2019" bias, it would seem to be more appropriate to replace that analysis, rather than supplement it.

      Regardless, this seems to be a moot issue, given that the underlying decoders seem to be overfit to temporally structured noise (see point above regarding weakening of downstream analyses based on decoder bias).

      Thank you for this important point. We now replaced the previous control analysis with a new one that does not assume stationary noise structure (page 19 in the revised manuscript). In Cai et al (2019), the source of confound is the covariance between observations. Specifically, as the observations in fMRI data are the BOLD signal at different time points, TSN can introduce covariance between nearby observations, which further biases the observed correlation between experimental conditions/trial types. In our case, the observations are decoding accuracy for different trial types. Thus, bias in the correlation may come from covariance between trial types. In this study, potential covariance between trial types includes the constrain that the decoding accuracy of all trial types adds up to 1 for a given trial (although we transformed the accuracy into logits prior to RSA, so the constrain may not hold), and the blocked design (as discussed above). Thus, to establish a baseline level of correlation between SC and SR representation strength, we took a similar shuffling approach as in Cai et al (2019) and randomly shuffled the trial types within each block. The reason to shuffle within each block is to preserve the covariance structure in the blocked design. We then repeated the same analysis using the shuffled data. The results of 10 shuffled analysis were averaged to form a baseline. Please note that (1) this control analysis was performed separately at each time point, hence removing the assumption of stationary noise structure, (2) this analysis also included as noise any covariance introduced by the proactive cognitive control guided by the learned SC and SR associations (see response to comment 1), thus it is more stringent than intended and (3) this control analysis started from decoding and was intended to provide a baseline for all downstream analysis. As shown in figures 2A, 3A, 7C and 8C, the reviewer was correct that the bias was not stationary, as the baseline of correlation coefficient varies over time. Additionally, the SC-SR representation strength correlation remained significantly above baseline between ~100 and ~ 450 ms following stimulus onset and between -180 and + 50 ms relative to response, suggesting that the noise structure (even when including potential proactive cognitive control) cannot fully explain the observed the SC-SR representation strength correlation. Considering the fact that this control analysis treated proactive control as a source of confound, this result does not necessarily contradict the absence of distance effect reported above.

      Outliers and t-values:

      More outliers with beta coefficients could be because the original SD estimates from the t-values are influenced more by extreme values. When you use a threshold on the median absolute deviation instead of mean +/-SD, do you still get more outliers with beta coefficients vs t-values?

      Thank you for your suggestion. We calculated the proportion of outliers with a threshold of median absolute deviation (defined as values beyond median ± 5 median absolute deviation) for each subject. The outliers remained less frequent for t-values than for beta coefficients (t-values: mean = 1.08%, SD = 0.12%; beta-values: mean = 4.45%, SD = 0.28%). Based on these results and to maintain consistent with previous studies employing the methods (Cellier et al., 2022; Kikumoto & Mayr, 2020; Kikumoto et al., 2022a; Kikumoto et al., 2022b; Rangel et al., 2023), we still decided to stay with t-values.

      Random slopes:

      Were random slopes (by subject) for all within-subject variables included in the LMMs? If not, please include them, and report this in the Methods.

      Thank you for your suggestion. The model failed to converge with random slopes of all variables. Thus, we chose not to add random slopes in the LMM. But we have added the random effects structure in the methods (see page 34).

      Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has a sold design and the analyses are appropriate for the research questions.

      Strength:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking the strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      We thank you for acknowledging our work on design and brain-behavior analysis. We have addressed all your concerns raised below point by point.

      Weakness:

      (1a) The distinction between Phase 2 and Phase 1&3 behavioral results, specifically the opposite effect of MC/MI in congruent trials raises some concerns with regard to the effectiveness of the ISPC manipulation. Why do RTs and error rates under MC congruent condition in Phase 2 seem to be worse than MI congruent?

      Thank you for raising these issues. In Phase 1, one color set (red and blue) was assigned to the MC condition, whereas another color set (yellow and green) was assigned to the MI condition. In Phase 2, these assignments were flipped, and they were flipped back again in Phase 3. Thus, the MC condition consisted of red and blue in Phases 1 and 3 but yellow and green in Phase 2, whereas the MI condition consisted of yellow and green in Phases 1 and 3 but red and blue in Phase 2 (Fig. 1b in the manuscript). This manipulation leads to seemingly opposite patterns between Phases 1 & 3 and Phase 2.

      However, when considering specific colors, the pattern is consistent across phases. In Phase 2, RTs and error rates for yellow and green (MC congruent) were worse than those for red and blue (MI congruent), which mirrors the pattern observed in Phases 1 and 3, where RTs and error rates for yellow and green (MI congruent) were worse than those for red and blue (MC congruent)

      We interpreted the results in Phase 2 as reflecting a typical ISPC effect, which is defined as a smaller conflict effect in the MI condition (MI incongruent – MI congruent) compared with the MC condition (MC incongruent – MC congruent). To our knowledge, the ISPC paradigm does not impose a specific prediction regarding the relative difference between MC-congruent and MI-congruent conditions.

      (1b) Could there be other factors at play here, e.g. order effect?

      We agree that order effect could play a role, such that memory from Phase 1 may influence the pattern in phase 2. For example, in phase 1, yellow and green were assigned to the MI condition, and participants therefore have associated these colors with a high control state (SC) and incongruent responses (SR). These prior associations could interfere with the newly learned mappings in Phase 2, where yellow and green were reassigned to the MC congruent condition (i.e., low control state and congruent responses). As a result, memory from Phase 1 may have weakened the expected MC in phase 2. A similar effect could also apply to the MI condition. Consequently, the same condition does not show parallel performance between phase 1 and phase 2, which may lead to different patterns in the difference between MC congruent and MI congruent conditions in phase 2.

      (1c) How does this potentially affect the neural analyses where trials from different phases were combined?

      Thank you for the question. As we mentioned above, the order effect could slow down the newly learned associations. However, we still found the ISPC effect in each phase, suggesting that all kinds of both SC and SR associations were formed and could be applied to the decoding and the following analyses cross phases. Relatedly, there might be confounded with temporal structured noise (TSN) when the neural analyses on decoding were combined the trials from different phases. However, we have performed the control decoding analyses and distance effect tests and confirmed that our decoding results were not driven by TSN (Please see comment #1 of R1).

      (1d) the manuscript does not mention whether there is counterbalancing for the color groups across participants, so far as I can tell.

      Thank you for the reminder. We have balanced the color groups by randomly dividing the participants into two groups and assigning different color sets to each group. The related interpretations have been included in task overview of the revised manuscript (page 6), which reads:

      “The color groups were counterbalanced across participants by red and blue as the color set of MC in one group while as the color set of MI in another group in the phase 1.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I commend the authors for addressing and clarifying my previous questions. One new comment regarding the newly added Figure 9: the response-locked behavioral correlation is much weaker compared to the stimulus-locked one, even never reaching the significance level. I think this difference should be discussed instead of simply glossing over it.

      Thank you for your suggestion. We discussed the difference in the discussion of revised manuscript (page 25), which reads:

      “Note that we found the negative prediction of the strength of SC and SR to RTs did not reach statistical significance with response-locked analysis as stimulus-locked analysis. It is possible that SC and SR representations have occurred before the stage of response processing, which is usually aligned with stimulus onset (Jiang et al., 2020a; Kang & Yu-Chin, 2024; Khan et al., 2025)”

      References

      Braver, T. S. (2012). The variable nature of cognitive control: a dual mechanisms framework. Trends Cogn Sci, 16(2), 106-113. doi:10.1016/j.tics.2011.12.010

      Cellier, D., Petersen, I. T., & Hwang, K. (2022). Dynamics of Hierarchical Task Representations. J Neurosci, 42(38), 7276-7284. doi:10.1523/JNEUROSCI.0233-22.2022

      Collins, A. G., & Frank, M. J. (2018). Within- and across-trial dynamics of human EEG reveal cooperative interplay between reinforcement learning and working memory. Proceedings of the National Academy of Sciences, 115(10), 2502-2507. doi:10.1073/pnas.1720963115

      Khan, A. U., Hoy, C. W., Anderson, K. L., Piai, V., King-Stephens, D., Laxer, K. D., . . . Bentley, J. N. (2025). Neural dynamics of proactive and reactive cognitive control in medial and lateral prefrontal cortex. iScience, 28(9), 113375. doi:10.1016/j.isci.2025.113375

      Kikumoto, A., & Mayr, U. (2020). Conjunctive representations that integrate stimuli, responses, and rules are critical for action selection. Proc Natl Acad Sci 117(19), 10603-10608. doi:10.1073/pnas.1922166117

      Kikumoto, A., Mayr, U., & Badre, D. (2022a). The role of conjunctive representations in prioritizing and selecting planned actions. Elife, 11. doi:10.7554/eLife.80153

      Kikumoto, A., Sameshima, T., & Mayr, U. (2022b). The Role of Conjunctive Representations in Stopping Actions. Psychol Sci, 33(2), 325-338. doi:10.1177/09567976211034505

      Rangel, B. O., Hazeltine, E., & Wessel, J. R. (2023). Lingering Neural Representations of Past Task Features Adversely Affect Future Behavior. J Neurosci, 43(2), 282-292. doi:10.1523/JNEUROSCI.0464-22.2022

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study of Drosophila mating behaviors has offered a powerful entry point for understanding how complex innate behaviors are instantiated in the brain. The effectiveness of this behavioral model stems from how readily quantifiable many components of the courtship ritual are, facilitating the fine-scale correlations between the behaviors and the circuits that underpin their implementation. Detailed quantification, however, can be both time-consuming and error-prone, particularly when scored manually. Song et al. have sought to address this challenge by developing DrosoMating, software that facilitates the automated and high-throughput quantification of 6 common metrics of courtship and mating behaviors. Compared to a human observer, DrosoMating matches courtship scoring with high fidelity. Further, the authors demonstrate that the software effectively detects previously described variations in courtship resulting from genetic background or social conditioning. Finally, they validate its utility in assaying the consequences of neural manipulations by silencing Kenyon cells involved in memory formation in the context of courtship conditioning.

      Strengths:

      (1) The authors demonstrate that for three key courtship/mating metrics, DrosoMating performs virtually indistinguishably from a human observer, with differences consistently within 10 seconds and no statistically significant differences detected. This demonstrates the software's usefulness as a tool for reducing bias and scoring time for analyses involving these metrics.

      (2) The authors validate the tool across multiple genetic backgrounds and experimental manipulations to confirm its ability to detect known influences on male mating behavior.

      (3) The authors present a simple, modular chamber design that is integrated with DrosoMating and allows for high-throughput experimentation, capable of simultaneously analyzing up to 144 fly pairs across all chambers.

      Weaknesses:

      (1) DrosoMating appears to be an effective tool for the high-throughput quantification of key courtship and mating metrics, but a number of similar tools for automated analysis already exist. FlyTracker (CalTech), for instance, is a widely used software that offers a similar machine vision approach to quantifying a variety of courtship metrics. It would be valuable to understand how DrosoMating compares to such approaches and what specific advantages it might offer in terms of accuracy, ease of use, and sensitivity to experimental conditions.

      (2) The courtship behaviors of Drosophila males represent a series of complex behaviors that unfold dynamically in response to female signals (Coen et al., 2014; Ning et al., 2022; Roemschied et al., 2023). While metrics like courtship latency, courtship index, and copulation duration are useful summary statistics, they compress the complexity of actions that occur throughout the mating ritual. The manuscript would be strengthened by a discussion of the potential for DrosoMating to capture more of the moment-to-moment behaviors that constitute courtship. Even without modifying the software, it would be useful to see how the data can be used in combination with machine learning classifiers like JAABA to better segment the behavioral composition of courtship and mating across genotypes and experimental manipulations. Such integration could substantially expand the utility of this tool for the broader Drosophila neuroscience community.

      (3) While testing the software's capacity to function across strains is useful, it does not address the "universality" of this method. Cross-species studies of mating behavior diversity are becoming increasingly common, and it would be beneficial to know if this tool can maintain its accuracy in Drosophila species with a greater range of morphological and behavioral variation. Demonstrating the software's performance across species would strengthen claims about its broader applicability.

      Reviewer #2 (Public review):

      This paper introduces "DrosoMating," an integrated hardware and software solution for automating the analysis of male Drosophila courtship. The authors aim to provide a low-cost, accessible alternative to expensive ethological rigs by utilizing a custom acrylic chamber and smartphone-based recording. The system focuses on quantifying key temporal metrics-Courtship Index (CI), Copulation Latency (CL), and Mating Duration (MD)-and is applied to behavioral paradigms involving memory mutants (orb2, rut).

      The development of open-source behavioral tools is a significant contribution to neuroethology, and the authors successfully demonstrate a system that simplifies the setup for large-scale screens. A major strength of the work is the specific focus on automating Copulation Latency and Mating Duration, metrics that are often labor-intensive to score manually.

      However, there are several limitations in the current analysis and validation that affect the strength of the conclusions:

      First, the statistical rigor requires substantial improvement. The analysis of multi-group experiments (e.g., comparing four distinct strains or factorial designs with genotype and training) currently relies on multiple independent Student's t-tests. This approach is statistically invalid for these experimental designs as it inflates the family-wise Type I error rate. To support the claims of strain-specific differences or learning deficits, the data must be analyzed using Analysis of Variance (ANOVA) to properly account for multiple comparisons and to explicitly test for interaction effects between genotype and training conditions.

      Second, the biological validation using $w^{1118}$ and $y^1$ mutants entails a potential confound. The authors attribute the low Courtship Index in these strains to courtship-specific deficits. However, both strains are known to exhibit general locomotor sluggishness (due to visual or pigmentation/behavioral defects). Since "following" behavior is likely a component of the Courtship Index, a reduction in this metric could reflect a general motor deficit rather than a specific lack of reproductive motivation. Without controlling for general locomotion, the interpretation of these behavioral phenotypes remains ambiguous.

      Third, the benchmarking of the system is currently limited to comparisons against manual scoring. Given that the field has largely adopted sophisticated open-source tracking tools (e.g., Ctrax, FlyTracker, JAABA), the utility of DrosoMating would be better contextualized by comparing its performance - in terms of accuracy, speed, or identity maintenance - against these existing automated standards, rather than solely against human observation.

      Finally, the visual presentation of the data hinders the assessment of the system's temporal precision. While the system is designed to capture time-resolved metrics, the results are presented primarily as aggregate bar plots. The absence of behavioral ethograms or raster plots makes it difficult to verify the software's ability to accurately detect specific transitions, such as the exact onset of copulation.

      We sincerely thank the reviewers for their constructive and detailed feedback, which has substantially improved the clarity and rigor of our work. Below is a summary of the major revisions.

      (1) Comparison with existing tools. We added a new main figure (Figure 5) and Table 1 systematically benchmarking DrosoMating against Ctrax and FlyTracker on identical low-quality, high-throughput mating videos. Both conventional tools failed under our recording conditions: Ctrax showed severe segmentation instability and fragmented trajectories, while FlyTracker frequently crashed during feature computation. These results demonstrate that DrosoMating's state-detection approach bypasses the pose-tracking limitations that impair established pipelines. We also revised the Discussion to clarify that DrosoMating is specialized for robust extraction of mating timing metrics from low-quality videos, not a replacement for general-purpose tracking tools.

      (2) Statistical analysis. Following Reviewer #2's recommendation, we re-analyzed Figure 3 using one-way ANOVA with Tukey's HSD for strain comparisons, and Figure 4 using two-way ANOVA with Sidak's post-hoc test for learning assays, including the critical Genotype by Training interaction term.

      (3) New supplementary data. We added Figure S4 showing basal locomotor velocity of single-housed males across all four strains to decouple motor defects from courtship deficits in w1118 and y1 mutants. We also added Figure S3 with behavioral ethograms for individual flies to visually demonstrate the system's temporal resolution and detection accuracy.

      (4) Additional improvements. We standardized statistical reporting across all figure legends, explicitly defined the segmentation threshold parameter s in the Methods, consistently used "Mating Duration (MD)" throughout, standardized MB247-GAL4 labeling, clarified that occluded frames are retained for mating-state detection via merged-contour analysis, and corrected minor errors including duplicate references and unnecessary quotation marks.

      We believe these revisions substantially strengthen the manuscript and clearly position DrosoMating within the existing ecosystem of behavioral analysis tools.

      We remain grateful for the valuable feedback from both reviewers and the editorial team, and we hope the revised version meets the standards for publication in eLife.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It's difficult to assess the utility of this tool in relation to the variety of alternative methods available for the automated scoring of courtship behavior, some of which offer more granular behavioral data than what DrosoMating has been presented to produce. A direct comparison to other approaches (e.g., FlyTracker, DeepLabCut, SLEAP) would be highly valuable for understanding what use cases DrosoMating is best suited for. Ideally, this analysis would include (i) a comparison of the accuracy of scoring courtship metrics and constituent behaviors, (ii) an assessment of the sensitivity of the various software to different experimental conditions, such as lighting and alternative chamber designs, and (iii) testing across Drosophila species, which could help highlight the particular strengths of DrosoMating over more established methods. At a minimum, a table comparing key features, requirements, and capabilities would help position DrosoMating within the existing ecosystem of tools.

      We sincerely appreciate the reviewer’s insightful comments on benchmarking DrosoMating against existing tools for automated courtship behavior analysis. We fully agree that direct comparisons are essential to clarify the unique advantages and ideal use cases of DrosoMating. We have extensively revised the manuscript by adding comparative experiments, a new figure(Figure 5) and comparative table (Table1), and expanded descriptions, as detailed below.

      (1) Direct comparison with conventional tracking tools

      We systematically tested Ctrax and FlyTracker on the same low-quality, high-throughput mating videos used for DrosoMating. The corresponding results are presented in Figure 5.

      Ctrax showed severe segmentation instability, including over-segmentation, under-segmentation, and fragmented trajectories, and failed to stably detect two flies.

      FlyTracker was able to generate background models and showed partially improved segmentation after threshold adjustment, but frequently crashed during feature computation and could not complete the end-to-end analysis pipeline.

      These results confirm that tracking-based tools are vulnerable to low contrast, chamber artifacts, and prolonged male–female overlap, whereas DrosoMating bypasses pose tracking and directly detects mating states.

      The revised manuscript text is as follows (line 265-342):

      "DrosoMating is more compatible with low-quality mating videos than conventional tracking-based pipelines.

      To evaluate whether conventional fly-tracking pipelines could be used as alternative tools for extracting mating-duration metrics, we tested Ctrax and FlyTracker on representative low-quality single-chamber mating videos recorded under our standard high-throughput conditions (Fig. 5). Ctrax is designed to estimate the position and orientation of multiple walking flies while maintaining individual identities over time (Branson et al., 2009) (https://ctrax.sourceforge.net/), whereas FlyTracker aims to track fly pose, including position, orientation, body size, wing and leg positions, and to generate trajectory- and feature-based outputs for downstream behavior analysis (Eyjolfsdottir et al., 2014). Because these programs were not readily compatible with our full high-throughput behavioral recording setup, we first cropped the original videos and tested single-chamber videos containing one male and one female fly.

      Using Ctrax, we observed that target detection was highly variable even within single-chamber videos. In representative frames, Ctrax could occasionally identify the two flies correctly (Fig. 5A). However, imperfect segmentation was frequently observed. In some frames, parts of the fly body were detected as additional targets, resulting in over-segmentation and an apparent increase in the number of detected flies (Fig. 5B, C). Conversely, when the male and female were close to each other or physically overlapped, the two animals were sometimes detected as a single target (Fig. 5D). These examples indicate that Ctrax detection was sensitive to the low contrast and overlapping fly bodies present in our mating videos. We then adjusted the Ctrax detection threshold to improve segmentation quality. Although threshold optimization improved detection in some frames, abnormal detections remained evident across randomly sampled frames (Fig. 5E). For example, some frames still showed incorrect target numbers, and a severe segmentation failure was observed in the lower-right example, where the detected objects did not correspond to two clearly separable flies. Thus, even after parameter optimization, Ctrax did not consistently maintain the expected two-target detection state in single-chamber mating videos.

      The instability of Ctrax detection was also reflected in the trajectory output. In the first 500 frames, the generated trajectories were fragmented into multiple colored track segments rather than two continuous trajectories corresponding to the male and female (Fig. 5F). In addition, some trajectories extended outside the chamber boundary, indicating tracking errors and identity instability. Consistently, frame-by-frame quantification of detected target number showed that the detected object count did not remain stable at the expected value of two flies per chamber (Fig. 5G). Because Ctrax failed to maintain stable two-fly detection and continuous trajectories under these video conditions, its output could not be reliably used for downstream extraction of copulation latency or mating duration.

      We next tested FlyTracker on the same type of cropped single-chamber mating videos. FlyTracker was developed to track multiple flies by estimating body position, orientation, size, wing and leg positions, and by maintaining fly identities across video frames; it also outputs per-frame features such as velocity, facing angle, and wing-angle-related measurements for downstream behavioral analysis (Eyjolfsdottir et al., 2014). In our videos, FlyTracker was able to generate a background model, indicating that the program could recognize the overall imaging field and chamber background (Fig. 5H). However, during calibration and segmentation, the default threshold setting produced inconsistent detection results (Fig. 5I). Although some frames were segmented relatively well under the default threshold (Fig. 5J), these successful examples were not representative of the overall tracking process, and the program frequently terminated with runtime errors during tracking or downstream feature computation. To improve detection stability, we lowered the segmentation threshold. Under this adjusted setting, FlyTracker produced more complete fly masks in representative frames (Fig. 5K), and the diagnostic output showed that two flies were detected in many sampled frames (Fig. 5L). Nevertheless, detection remained unstable in some sampled frames, including frames in which zero or one fly was detected despite the expected two flies per chamber (Fig. 5L). The full pipeline ultimately failed during feature computation, producing a runtime error before complete tracking and feature outputs could be generated (Fig. 5M). Therefore, even after threshold adjustment, FlyTracker could not provide a stable end-to-end workflow for extracting mating-duration metrics from these low-quality mating videos.

      This failure mode is relevant because FlyTracker depends on stable segmentation, identity maintenance, and per-frame feature extraction. In our assay videos, low contrast, chamber-edge artifacts, and prolonged male–female overlap during copulation interfered with these requirements. As a result, FlyTracker could occasionally identify the flies in individual frames, but it did not reliably complete the full analysis pipeline required for downstream behavioral quantification. This limitation is especially important for workflows such as JAABA, which use manually labeled examples to train behavior classifiers but still depend on upstream tracking-derived features. Thus, for our low-quality high-throughput mating recordings, FlyTracker-based analysis was substantially less practical than DrosoMating, which directly outputs mating-related timing metrics without requiring continuous high-fidelity two-fly pose tracking."

      (2) Evaluation of accuracy, robustness, and cross-species potential

      We addressed the three key points requested by the reviewer:

      (i) Accuracy: DrosoMating achieved 98–99% agreement with manual scoring. Under our experimental conditions, neither Ctrax nor FlyTracker completed an end-to-end workflow capable of reliably extracting copulation latency and mating duration. Ctrax produced fragmented trajectories with unstable target counts, and FlyTracker terminated with runtime errors during feature computation.

      (ii) Robustness: DrosoMating is highly robust to low-contrast lighting and common behavioral chamber setups, whereas conventional tools require high-quality videos and fail under fly occlusion.

      (iii) Cross-species testing: We did not perform cross-species validation in this revision. DrosoMating relies on mating state detection rather than species-specific morphology, suggesting potential transferability, although this remains to be experimentally validated. We have therefore revised the text to avoid claiming demonstrated cross-species performance and now describe this as a potential future application.

      The revised manuscript text is as follows (line 572-576):

      "Cross-species testing was not performed in this study. Although DrosoMating's state-detection approach is morphology-agnostic and therefore theoretically applicable across Drosophila species, we have revised the text to avoid claiming demonstrated cross-species performance. Formal validation across diverse species remains a promising future direction."

      We added Table1 to compare key features across DrosoMating, Ctrax, and FlyTracker, including low-quality video compatibility, multi-chamber support, robustness to fly overlap, direct output of mating timing metrics, accuracy.

      (3) Clarification of ideal use cases

      We revised the Discussion to emphasize that DrosoMating is not intended to replace general-purpose pose or tracking tools (e.g., FlyTracker, Ctrax) that provide fine-grained behavioral data.

      Instead, it offers a specialized, robust, and high-throughput workflow for scenarios requiring efficient extraction of mating timing metrics from low-quality videos, where conventional pipelines often fail.

      These revisions greatly improve the clarity and positioning of DrosoMating. We thank the reviewer for this valuable suggestion.

      The revised manuscript text is as follows (line 578-620):

      “Comparative Evaluation with Existing Courtship Analysis Tools

      A major advantage of DrosoMating is its compatibility with low-quality, high-throughput mating videos. Conventional tracking-based tools such as Ctrax and FlyTracker are powerful for trajectory- and pose-based behavioral analysis, but they generally require stable object segmentation, identity maintenance, and reliable feature extraction across frames. Ctrax was designed to estimate the positions and orientations of multiple walking flies while maintaining their identities, whereas FlyTracker aims to track detailed fly pose and generate per-frame behavioral features. Under our recording conditions, these requirements were difficult to satisfy because the videos were low contrast and the male and female frequently overlapped during copulation.

      In our tests, both tools showed limited compatibility with these videos. Even after cropping to single-chamber videos and adjusting detection parameters, Ctrax produced unstable target numbers, fragmented trajectories, and tracking errors. FlyTracker could generate a background model and occasionally segment flies successfully, but detection remained unstable and the full pipeline failed during feature computation. These issues are particularly relevant for copulation latency and mating duration analysis: during copulation, the male and female remain physically coupled for a long period, which makes identity-based tracking difficult. For these timing metrics, it is more important to robustly detect the onset and offset of the mating state than to reconstruct detailed individual trajectories. DrosoMating was purpose-built to address these specific challenges. It operates reliably on lower-quality video streams, requires no complex pre-processing or manual ROI definition, and is optimized for high-throughput multi-chamber analysis. While it does not offer the same level of pose or kinematic detail as other tools, it provides a unique solution for laboratories seeking a simple, fast, and robust pipeline to quantify core reproductive timing metrics—copulation latency (CL), courtship index (CI), and mating duration (MD)—without the overhead of more complex systems (Table 1).

      Although we directly tested only Ctrax and FlyTracker, this limitation may also affect workflows that depend on upstream tracking-derived features. For example, JAABA uses tracking-derived features to train behavior classifiers, and DANCE, a recent Drosophila aggression and courtship pipeline, uses JAABA-based classifiers and lists FlyTracker and JAABA as required software. Thus, our conclusion is not that these tools are generally unsuitable for Drosophila behavior analysis, but that DrosoMating provides a more practical workflow for low-quality, high-throughput videos focused specifically on mating timing.

      While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays.”

      We have also added this clarification to the Table 1 legend to avoid overgeneralizing the limitations of Ctrax and FlyTracker (line 786-790):

      "DrosoMating was designed to extract mating-related timing metrics from high-throughput videos without requiring continuous two-fly identity tracking. Ctrax and FlyTracker were tested on cropped single-chamber videos from the same recording setup. Their limitations described here refer specifically to these low-quality mating videos and should not be interpreted as general limitations of the tools."

      (2) The coarse nature of the metrics captured by DrosoMating may limit the usefulness of this tool for many researchers. Consider integrating DrosoMating with one or more behavioral classifiers (e.g., JAABA) and validating its performance to increase its utility across a wider range of uses. Demonstrating the feasibility of this would substantially increase the tool's appeal to those who need access to more discrete behavioral elements.

      We thank the reviewer for this constructive suggestion. We agree that DrosoMating is currently optimized for rapid, high-throughput quantification of core mating timing metrics (CL, CI, and MD) rather than discrete behavioral classification (e.g., wing extension, circling, or aggressive postures). This reflects a deliberate design trade-off: by prioritizing robust state detection over detailed pose tracking, DrosoMating achieves reliable performance on low-quality videos where conventional identity-based pipelines fail.

      We appreciate the reviewer’s vision for expanding the tool’s utility. While DrosoMating does not presently generate the per-frame kinematic features (position, orientation, wing angles, etc.) required as input for JAABA classifiers, its modular architecture and underlying video-processing framework provide a foundation for future integration. In the revised Discussion, we have clarified that extending DrosoMating to export trajectory-derived features compatible with downstream classifiers such as JAABA represents a promising future direction—one that would broaden its applicability to laboratories requiring granular behavioral elements, without necessitating a complete overhaul of the existing pipeline.

      We hope this clarification addresses the reviewer’s concern and accurately reflects both the current capabilities and future potential of the tool.

      The revised manuscript text is as follows (line 614-620):

      "While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays."

      (3) Please clarify how DrosoMating handles fly identification during mating? I would think mounted flies would be occluded, and if these frames are excluded from analysis, it would be expected to skew mating duration scores. It would be helpful if the authors could discuss these details.

      Occluded frames are excluded from CI calculation because individual courtship actions cannot be reliably assigned during prolonged overlap. However, these frames are not discarded from mating-duration analysis. Instead, prolonged merged contours are used as evidence for copulation-state detection, and the MD timer continues until physical separation:

      (1) When two flies are separate, the system detects two contours; when they overlap during mounting, it detects one merged contour. Our code processes both cases, so occluded frames are retained, not skipped.

      (2) To distinguish a single fly from two overlapping flies, we analyze the shape of the merged contour. Two overlapping flies produce a characteristically different aspect ratio (elongated shape) compared to one fly. This geometric cue is fed into the state classifier to label the frame as "mating."

      (3) Because these frames are classified as mating rather than excluded, the mating duration timer runs continuously through the occlusion period. Thus, mating duration is not artificially shortened.

      The high agreement between DrosoMating and manual scoring for mating duration (within 10 s, no significant difference) confirms that this approach does not introduce bias.

      We revised the manuscript as follow (line 191-193):

      "Occluded frames are excluded from CI calculation but retained for mating-state detection via merged-contour analysis, ensuring continuous MD measurement through copulation."

      (4) Please indicate the statistical tests used in Figures 3, 4, and S1. What methods were used to address multiple hypothesis testing?

      We thank the reviewer for this important comment. For comparisons between two groups, two-sided Student’s t-tests were used. For comparisons among three or more groups, one-way ANOVA followed by post-hoc tests were applied. For two-factor experimental designs, two-way ANOVA was used. We have clearly stated these statistical tests in the Statistical Analysis section and have added this information to the figure legends for Figures 3, 4, and S1 in the revised manuscript.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (5) Line 43: "High resolution video tracking enables...".

      We thank the reviewer for noting this imprecise expression. We have revised the statement to accurately describe that our method uses image-based video analysis to identify courtship and copulation events and quantify their temporal parameters. The description has been corrected in the revised manuscript line 44-46:

      "Our image-based video analysis enables precise identification of courtship and copulation events, as well as quantification of their timing and duration under controlled conditions. "

      (6) Line 51: remove quotation mark.

      We thank the reviewer for the careful correction. The unnecessary quotation mark at Line 51 has been removed in the revised manuscript.

      (7) Line 107: Chen et al, 2024 is listed twice in the references.

      We thank the reviewer for the careful correction. The duplicate reference of Chen et al., 2024 has been removed from the reference list in the revised manuscript.

      (8) Line 242: remove quotation mark.

      We thank the reviewer for the careful correction. The unnecessary quotation mark at Line 242 has been removed in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) Please re-analyze the data in Figures 3 and 4 using ANOVA followed by appropriate post-hoc tests (e.g., Tukey's HSD). Specifically, use a One-way ANOVA for strain comparisons in Figure 3 and a Two-way ANOVA for the learning assays in Figure 4. The interaction term (Genotype $\times$ Training) is critical for demonstrating specific learning deficits. Update the "Statistical Analysis" section and figure legends accordingly.

      We appreciate the reviewer’s recommendation. We have re-analyzed the data in Figures 3 and 4 using the suggested ANOVA approaches: one-way ANOVA with Tukey’s HSD for strain comparisons (Figure 3), and two-way ANOVA (including the Genotype × Training interaction term) with Sidak's post-hoc test for learning assays (Figure 4). The updated statistical methods are now described in the Statistical Analysis section and figure legends.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (2) To decouple motor defects from courtship deficits in $w^{1118}$ and $y^1$ mutants, please use your tracking data to calculate and present a "General Locomotion" metric (e.g., average velocity or total distance traveled in the absence of a female).

      We thank the reviewer for this suggestion. We have now added Figure S4 showing basal locomotor velocity of single-housed males across all four strains. As expected, w^1118 and y^1 mutants move more slowly than wild-type controls.

      These data reveal that both motor and sensory factors contribute to the observed courtship phenotypes. Reduced basal locomotion likely limits the males' ability to approach and follow females. However, this generalized hypoactivity is compounded by strain-specific sensory deficits: w<sup>1118</sup> males suffer visual impairment that compromises female detection (Krstic et al., 2013), while y<sup>1</sup> males display altered cuticular hydrocarbons that disrupt pheromonal communication (Drapeau et al., 2006). These sensory defects impair courtship initiation and female recognition independent of locomotor capacity. Thus, the reduced CI, CL, and MD in these mutants reflect the combined effects of slower movement and courtship-specific sensory impairments, rather than motor defects alone.

      We have revised the manuscript to incorporate the velocity data and clarify this interpretation (line 219-226):

      "Notably, reduced basal locomotor activity in w<sup>1118</sup> and y<sup>1</sup> mutants has been well documented in previous studies, independent of courtship behavior (Drapeau et al., 2006; Krstic et al., 2013). Consistent with these reports, our tracking data show that single-housed w<sup>1118</sup> and y<sup>1</sup> males exhibit lower average velocity than Canton-S and Oregon-R controls (Fig. S4A). These general locomotor differences are insufficient to fully explain the observed courtship and mating timing phenotypes, indicating that additional courtship‑related processes contribute to the observed behavioral differences."

      (3) Please expand the discussion or provide a small comparative dataset contrasting DrosoMating with established tools like JAABA. Explain the specific advantages of your pipeline (e.g., cost, simplicity, focus on CL/MD) to justify its adoption over these alternatives.

      We sincerely appreciate the reviewer’s insightful comments on benchmarking DrosoMating against existing tools for automated courtship behavior analysis. We fully agree that direct comparisons are essential to clarify the unique advantages and ideal use cases of DrosoMating. We have extensively revised the manuscript by adding comparative experiments, a new figure (Figure 5) and comparative table (Table 1), and expanded descriptions, as detailed below.

      (1) Direct comparison with conventional tracking tools

      We systematically tested Ctrax and FlyTracker on the same low-quality, high-throughput mating videos used for DrosoMating. The corresponding results are presented in Figure 5.

      Ctrax showed severe segmentation instability, including over-segmentation, under-segmentation, and fragmented trajectories, and failed to stably detect two flies.

      FlyTracker was able to generate background models and showed partially improved segmentation after threshold adjustment, but frequently crashed during feature computation and could not complete the end-to-end analysis pipeline.

      These results confirm that tracking-based tools are vulnerable to low contrast, chamber artifacts, and prolonged male–female overlap, whereas DrosoMating bypasses pose tracking and directly detects mating states.

      The revised manuscript text is as follows (line 265-342):

      " DrosoMating is more compatible with low-quality mating videos than conventional tracking-based pipelines

      To evaluate whether conventional fly-tracking pipelines could be used as alternative tools for extracting mating-duration metrics, we tested Ctrax and FlyTracker on representative low-quality single-chamber mating videos recorded under our standard high-throughput conditions (Fig. 5). Ctrax is designed to estimate the position and orientation of multiple walking flies while maintaining individual identities over time (Branson et al., 2009) (https://ctrax.sourceforge.net/), whereas FlyTracker aims to track fly pose, including position, orientation, body size, wing and leg positions, and to generate trajectory- and feature-based outputs for downstream behavior analysis (Eyjolfsdottir et al., 2014). Because these programs were not readily compatible with our full high-throughput behavioral recording setup, we first cropped the original videos and tested single-chamber videos containing one male and one female fly.

      Using Ctrax, we observed that target detection was highly variable even within single-chamber videos. In representative frames, Ctrax could occasionally identify the two flies correctly (Fig. 5A). However, imperfect segmentation was frequently observed. In some frames, parts of the fly body were detected as additional targets, resulting in over-segmentation and an apparent increase in the number of detected flies (Fig. 5B, C). Conversely, when the male and female were close to each other or physically overlapped, the two animals were sometimes detected as a single target (Fig. 5D). These examples indicate that Ctrax detection was sensitive to the low contrast and overlapping fly bodies present in our mating videos. We then adjusted the Ctrax detection threshold to improve segmentation quality. Although threshold optimization improved detection in some frames, abnormal detections remained evident across randomly sampled frames (Fig. 5E). For example, some frames still showed incorrect target numbers, and a severe segmentation failure was observed in the lower-right example, where the detected objects did not correspond to two clearly separable flies. Thus, even after parameter optimization, Ctrax did not consistently maintain the expected two-target detection state in single-chamber mating videos.

      The instability of Ctrax detection was also reflected in the trajectory output. In the first 500 frames, the generated trajectories were fragmented into multiple colored track segments rather than two continuous trajectories corresponding to the male and female (Fig. 5F). In addition, some trajectories extended outside the chamber boundary, indicating tracking errors and identity instability. Consistently, frame-by-frame quantification of detected target number showed that the detected object count did not remain stable at the expected value of two flies per chamber (Fig. 5G). Because Ctrax failed to maintain stable two-fly detection and continuous trajectories under these video conditions, its output could not be reliably used for downstream extraction of copulation latency or mating duration.

      We next tested FlyTracker on the same type of cropped single-chamber mating videos. FlyTracker was developed to track multiple flies by estimating body position, orientation, size, wing and leg positions, and by maintaining fly identities across video frames; it also outputs per-frame features such as velocity, facing angle, and wing-angle-related measurements for downstream behavioral analysis (Eyjolfsdottir et al., 2014). In our videos, FlyTracker was able to generate a background model, indicating that the program could recognize the overall imaging field and chamber background (Fig. 5H). However, during calibration and segmentation, the default threshold setting produced inconsistent detection results (Fig. 5I). Although some frames were segmented relatively well under the default threshold (Fig. 5J), these successful examples were not representative of the overall tracking process, and the program frequently terminated with runtime errors during tracking or downstream feature computation. To improve detection stability, we lowered the segmentation threshold. Under this adjusted setting, FlyTracker produced more complete fly masks in representative frames (Fig. 5K), and the diagnostic output showed that two flies were detected in many sampled frames (Fig. 5L). Nevertheless, detection remained unstable in some sampled frames, including frames in which zero or one fly was detected despite the expected two flies per chamber (Fig. 5L). The full pipeline ultimately failed during feature computation, producing a runtime error before complete tracking and feature outputs could be generated (Fig. 5M). Therefore, even after threshold adjustment, FlyTracker could not provide a stable end-to-end workflow for extracting mating-duration metrics from these low-quality mating videos.

      This failure mode is relevant because FlyTracker depends on stable segmentation, identity maintenance, and per-frame feature extraction. In our assay videos, low contrast, chamber-edge artifacts, and prolonged male–female overlap during copulation interfered with these requirements. As a result, FlyTracker could occasionally identify the flies in individual frames, but it did not reliably complete the full analysis pipeline required for downstream behavioral quantification. This limitation is especially important for workflows such as JAABA, which use manually labeled examples to train behavior classifiers but still depend on upstream tracking-derived features. Thus, for our low-quality high-throughput mating recordings, FlyTracker-based analysis was substantially less practical than DrosoMating, which directly outputs mating-related timing metrics without requiring continuous high-fidelity two-fly pose tracking."

      (2) Clarification of ideal use cases

      We revised the Discussion to emphasize that DrosoMating is not intended to replace general-purpose pose or tracking tools (e.g., FlyTracker, Ctrax) that provide fine-grained behavioral data.

      Instead, it offers a specialized, robust, and high-throughput workflow for scenarios requiring efficient extraction of mating timing metrics from low-quality videos, where conventional pipelines often fail.

      These revisions greatly improve the clarity and positioning of DrosoMating. We thank the reviewer for this valuable suggestion.

      The revised manuscript text is as follows (line 578-620):

      “Comparative Evaluation with Existing Courtship Analysis Tools

      A major advantage of DrosoMating is its compatibility with low-quality, high-throughput mating videos. Conventional tracking-based tools such as Ctrax and FlyTracker are powerful for trajectory- and pose-based behavioral analysis, but they generally require stable object segmentation, identity maintenance, and reliable feature extraction across frames. Ctrax was designed to estimate the positions and orientations of multiple walking flies while maintaining their identities, whereas FlyTracker aims to track detailed fly pose and generate per-frame behavioural features. Under our recording conditions, these requirements were difficult to satisfy because the videos were low contrast and the male and female frequently overlapped during copulation.

      In our tests, both tools showed limited compatibility with these videos. Even after cropping to single-chamber videos and adjusting detection parameters, Ctrax produced unstable target numbers, fragmented trajectories, and tracking errors. FlyTracker could generate a background model and occasionally segment flies successfully, but detection remained unstable and the full pipeline failed during feature computation. These issues are particularly relevant for copulation latency and mating duration analysis: during copulation, the male and female remain physically coupled for a long period, which makes identity-based tracking difficult. For these timing metrics, it is more important to robustly detect the onset and offset of the mating state than to reconstruct detailed individual trajectories. DrosoMating was purpose-built to address these specific challenges. It operates reliably on lower-quality video streams, requires no complex pre-processing or manual ROI definition, and is optimized for high-throughput multi-chamber analysis. While it does not offer the same level of pose or kinematic detail as other tools, it provides a unique solution for laboratories seeking a simple, fast, and robust pipeline to quantify core reproductive timing metrics—copulation latency (CL), courtship index (CI), and mating duration (MD)—without the overhead of more complex systems (Table 1).

      Although we directly tested only Ctrax and FlyTracker, this limitation may also affect workflows that depend on upstream tracking-derived features. For example, JAABA uses tracking-derived features to train behavior classifiers, and DANCE, a recent Drosophila aggression and courtship pipeline, uses JAABA-based classifiers and lists FlyTracker and JAABA as required software. Thus, our conclusion is not that these tools are generally unsuitable for Drosophila behavior analysis, but that DrosoMating provides a more practical workflow for low-quality, high-throughput videos focused specifically on mating timing.

      While DrosoMating currently prioritizes mating timing metrics over discrete behavioral classification, its modular architecture provides a foundation for future integration with behavior classifiers such as JAABA. Laboratories requiring granular behavioral elements—such as wing extension or circling—would benefit from an extended pipeline that exports per-frame kinematic features for downstream classifier training. Validating this integration represents a promising future direction to broaden the tool’s utility beyond core reproductive timing assays.”

      We have added this clarification to the Table 1 legend to avoid overgeneralizing the limitations of Ctrax and FlyTracker (line 786-790):

      "DrosoMating was designed to extract mating-related timing metrics from high-throughput videos without requiring continuous two-fly identity tracking. Ctrax and FlyTracker were tested on cropped single-chamber videos from the same recording setup. Their limitations described here refer specifically to these low-quality mating videos and should not be interpreted as general limitations of the tools."

      (4) Complement the aggregate bar plots with behavioral ethograms or raster plots for representative individual flies. Color-code these plots for specific states (resting, following, courting, copulating) to visually demonstrate the system's temporal resolution and detection accuracy.

      We thank the reviewer for this valuable suggestion. To address this comment, we have added a new supplementary figure (Figure S3) that presents behavioral ethograms for individual flies, directly complementing the aggregate bar plots in the main text. The figure displays the full temporal progression of courtship and copulation behaviors for all wells that exhibited successful mating. As requested, behaviors are color-coded (orange: courting, red: copulating) to clearly delineate different states. These ethograms visually demonstrate the system’s ability to resolve behavioral transitions with high temporal precision, confirming the accuracy of our automated detection of courtship initiation, duration, and copulation events. This addition provides critical individual-level validation that supports the aggregate statistical results presented in the main text.

      (5) Standardize the use of estimation statistics (DBMs). If used in Figure 3, they should also be applied to Figure 4, with appropriate statistical comparisons between groups.

      We thank the reviewer for this suggestion. We have standardized the statistical reporting format across all figure legends, with each legend explicitly stating the test used, the post hoc method (where applicable), and significance thresholds.

      Our approach is as follows:

      Fig. 2 and Fig.3D-I (two-group comparison, manual vs. automated scoring): DBM + Student's t-test.

      Fig. 3A-C and Fig. S1B-C, E-F (multi-group comparison, 4 strains or 2 rearing conditions): One-way ANOVA + Tukey's HSD.

      Fig. 4B-D and F (two-factor design, genotype × training): Two-way ANOVA + Sidak's post hoc.

      We have also added a sentence to the Methods clarifying that estimation statistics (DBM) are used for single two-group contrasts, while ANOVA-based approaches are used for multi-group or multi-factor designs. The statistical methods are now consistently documented in both the Methods section and the corresponding figure legends. We hope this clarification addresses the reviewer's concern.

      The revised "Statistical Analysis" section is as follows (line 512-524):

      "To ensure robust statistical analysis, each experimental group included at least 100 male flies (naïve, sexually experienced, or singly reared). Internal controls were incorporated into every experiment as recommended by Bretman et al. (2011) (Bretman et al., 2011). Normality of the mating duration data was confirmed using the Kolmogorov-Smirnov test (p>0.05). For group comparisons, two-sided Student’s t-tests were applied to calculate significance levels (****p<0.0001, ***p<0.001, **p<0.01, * p<0.05), Comparisons among three or more groups were performed using one-way ANOVA with Tukey’s HSD post-hoc tests. Two-factor experimental comparisons were performed using two-way ANOVA followed by Sidak’s post-hoc test. while estimation statistics (Claridge-Chang and Assam, 2016) were used to visualize effect sizes, mean differences, and precision, avoiding reliance solely on null hypothesis testing. All analyses, including data plotting, were performed using GraphPad Prism software."

      (6) Fix inconsistent labeling (e.g., "MB-247" vs. "MB247") and redundant axis labels.

      Thank you for pointing this out. We have now standardized all labels to "MB247-GAL4" throughout the text and figures.

      We have also removed redundant axis labels from multi-panel figures. And redundant DBM and metric labels have been streamlined in Figures 2 and 4. All statistical tests and metric definitions are now fully described in the corresponding figure legends rather than being repeated on each sub-panel.

      (7) Significantly increase the size of data panels to make individual data points and error bars legible.

      Thank you for this suggestion. We have standardized the figure formatting and simplified the panels by removing redundant axis labels and consolidating descriptive details into the figure legends. We believe the current panel sizes, combined with these clarifications, provide sufficient legibility for both data points and error bars in the final high-resolution PDF.

      Minor Corrections:

      (1) Abstract: Rephrase "lack of certain timing-related behavioral repertoires" (Line 108) to "lack of precise quantification for temporal parameters of post-copulatory behavior."

      Thank you for the suggested rephrasing. We have updated the sentence accordingly.

      (2) Define the physical parameter "s" (e.g., is it a pixel threshold?) to ensure reproducibility.

      Thank you for this important suggestion. We have now explicitly defined the parameter s in the Methods section. Briefly, s is the grayscale intensity threshold (0–255, 8-bit) used for binary segmentation of flies from the background. It is automatically calculated as the maximum grayscale value of three user-selected reference flies plus an offset of 28, and can be manually adjusted to accommodate varying illumination conditions.

      The revised manuscript text is as follows (line 499-503):

      “Note on the s value: The parameter s represents the grayscale intensity threshold (range: 0–255 for 8-bit images) used for binary segmentation of flies from the background. It is automatically calculated as the maximum grayscale value at the three reference fly positions plus an offset of 28, and can be manually adjusted to accommodate varying illumination conditions.”

      (3) Figure 1 Legend: Clarify the "clockwise selection" of pillars and their relation to well numbering.

      Thank you for this suggestion. We have revised the Figure 1 legend to clarify that:

      (1) The four pillars are selected in clockwise order starting from the top-left corner to define the chamber corners for perspective transformation (homography), which corrects for camera angle and standardizes the field of view.

      (2) Well numbering is independent of pillar selection order. After automated perspective correction, wells are numbered sequentially in a left-to-right, top-to-bottom order (1–36) based on the standardized chamber layout.

      The updated Figure 1 legend now reads:

      "Columns (pillars) should be selected in clockwise order starting from the top-left corner to define the chamber boundaries for perspective transformation (Fig. 1A, lower). Well numbering (1–36) follows a left-to-right, top-to-bottom sequence after perspective correction and is independent of pillar selection order."

      (4) Add the citation for Eastwood and Burnet (1977) regarding Courtship Latency.

      Thank you for pointing this out. We have added the citation Eastwood and Burnet (1977) to the Introduction where Courtship Latency is first defined

      (5) Select one term ("Mating Duration" or "Copulation Duration") and use it consistently throughout the text and figures.

      Thank you for this suggestion. We have now standardized the terminology throughout the manuscript and figures. "Mating Duration" (MD) is used consistently in all instances where "Copulation Duration" previously appeared. The abbreviation MD has been retained for consistency with existing figure labels

      References

      Branson, K., Robie, A.A., Bender, J., Perona, P., & Dickinson, M.H. (2009). High-throughput ethomics in large groups of Drosophila. Nature Methods, 6(6), 451–458.

      Claridge-Chang, A., & Assam, P.N. (2016). Estimation statistics should replace significance testing. Nature Methods, 13(2), 108–109.

      Drapeau, M.D., Cyran, S.A., Viering, M.M., Geyer, P.K., & Long, A.D. (2006). A cis-regulatory sequence within the yellow locus of Drosophila melanogaster required for normal male mating success. Genetics, 172(2), 1009–1030.

      Eastwood, L., & Burnet, B. (1977). Courtship latency in male Drosophila melanogaster. Behavior Genetics, 7(3), 359–372.

      Eyjolfsdottir, E., Branson, S., Burgos-Artizzu, X.P., Hoopfer, E.D., Schor, J., Anderson, D.J., & Perona, P. (2014). Detecting social actions of fruit flies. Lecture Notes in Computer Science, 8692, 772–787.

      Gil-Martí, B., Barredo, C.G., Pina-Flores, S., Poza-Rodriguez, A., Treves, G., Rodriguez-Navas, C., Camacho, L., Pérez-Serna, A., Jimenez, I., Brazales, L., Fernandez, J., & Martin, F.A. (2023). A simplified courtship conditioning protocol to test learning and memory in Drosophila. STAR Protocols, 4(1), 101572.

      Greenspan, R.J., & Ferveur, J.F. (2000). Courtship in Drosophila. Annual Review of Genetics, 34, 205–232.

      Hall, J.C. (1994). The mating of a fly. Science, 264(5163), 1702–1714.

      Kabra, M., Robie, A.A., Rivera-Alba, M., Branson, S., & Branson, K. (2013). JAABA: interactive machine learning for automatic annotation of animal behavior. Nature Methods, 10(1), 64–67.

      Kitamoto, T. (2001). Conditional modification of behavior in Drosophila by targeted expression of a temperature-sensitive shibire allele in defined neurons. Journal of Neurobiology, 47(2), 81–92.

      Krstic, D., Boll, W., & Noll, M. (2013). Influence of the White locus on the courtship behavior of Drosophila males. PLoS ONE, 8(9), e77904.

      Levin, L.R., Han, P.L., Hwang, P.M., Feinstein, P.G., Davis, R.L., & Reed, R.R. (1992). The Drosophila learning and memory gene rutabaga encodes a Ca2+/calmodulin-responsive adenylyl cyclase. Cell, 68(3), 479–489.

      Pavlou, H.J., & Goodwin, S.F. (2013). Courtship behavior in Drosophila melanogaster: towards a ‘courtship connectome’. Current Opinion in Neurobiology, 23(1), 76–83.

      Yapici, N., Kim, Y.J., Ribeiro, C., & Dickson, B.J. (2008). A receptor that mediates the post-mating switch in Drosophila reproductive behaviour. Nature, 451(7176), 33–37.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides a valuable advance in understanding how disordered proteins interact with cell membranes by identifying the sequence rules that enable aromatic residues to penetrate deeply into the membrane interior. The integration of complementary computational approaches, including molecular simulations, large-scale sequence analysis, and the development of an online prediction server, makes the work potentially impactful for the membrane protein and intrinsically disordered protein communities. The evidence supporting the main conclusions is generally convincing, although its transferability across diverse membrane compositions and its validity as a prediction tool for real protein-membrane systems remain to be further established.

      We thank the editors for recognizing our study as a valuable advance. This work lays a solid foundation for future developments to account for diverse membrane compositions and further refinements after additional experimental tests.

      Public review:

      Reviewer #1 (Public review):

      A primary limitation is the heavy reliance on computational modeling. Training for AroMIP is generated using PPM rather than direct experimental measurements, and so the model may primarily reproduce PPM behavior rather than true membrane insertion thermodynamics. Moreover, all simulations use a single lipid composition (POPC:POPS:<sub>2</sub> 70:25:5), but biological membranes vary substantially in cholesterol, cardiolipin, and acidic lipid content. Whether AroMIP's predictions transfer to diverse lipid environments remains untested. The 5% <sub>2</sub> concentration used in the simulations is higher than that of a normal mammalian cell and may therefore overemphasize electrostatic contributions. Applicability beyond short 9-residue motifs is unclear, as longer-range interactions or secondary structure in full-length IDRs could modulate insertion in ways the current model does not capture. This could be considered for future development.

      The reviewer’s point on our reliance on PPM for training, a single lipid composition, and potential effects beyond a 9-residue motif is well taken. Regarding PPM, we chose it as the optimal compromise for high-throughput data. However, we complemented the high-throughput PPM data with experimental data on an initial set of 10 peptides. Moreover, we validate AroMIP on an additional 12 IDRs (intrinsically disordered regions; Table S2). On membrane composition, we now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph).

      On potential effects beyond a 9-residue motif, we now add justification and note neglected factors for future developments (paragraph running from p. 19-20), as suggested by the reviewer.

      Reviewer #2 (Public review):

      (1) Aromatic residues have been shown to partition preferentially to the headgroup region of the lipid bilayer. Most of the papers on this problem were published in the mid 1990s to early 2000s. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev.

      Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602-608; Killian & von Heijne, TIBS 2000, 25, 429434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, none of these articles is cited.

      We have now citations to the Landolt-Marticorena paper and the von Heijne reviews [refs 25-27]. The Doyle paper is not particularly relevant. As for the Fleming paper, we cited a 2016 JACS paper (original ref 27; now ref 30) that specifically dealt with aromatic residues.

      (2) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth"."

      We emphasize that we have presented substantial experimental support for AroMIP. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM. Lastly, we now show good correlation between our insertion scores and binding free energies calculated from the scale determined experimentally by White and co-workers (new Figure S10; p. 15, second paragraph).

      (3) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is incorrect even in a qualitative sense.

      We have now revised these figures.

      Reviewer #3 (Public review):

      (1) Membrane composition and lipid shape characteristics: The authors chose to use a model membrane bilayer of a distinct lipid composition, POPC: POPS: PI4,5P2 (70:25:5 molar ratio), for their all-atom simulations of the various model peptides. While this may be pertinent for some of these peptides, it is not for many, such as sequence 2 derived from Drp1, which preferentially binds target conical lipids such as cardiolipin (CL) and phosphatidic acid (PA). The rationale behind using PI4,5P2, which can induce positive membrane curvature when sequestered, versus CL and PA, which both induce negative membrane curvature, is not explained.

      We now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph). In this Discussion paragraph, we also speculate that conical lipids, by promoting membrane defects, may facilitate membrane insertion.

      (2) Parallel vs. perpendicular peptide orientation of sequence 2 in peripheral Drp1-lipid interactions: On page 11, the authors state that their simulation results of sequence 2 derived from Drp1 "contrasts with a transmembrane orientation proposed by Mahajan et al." However, upon review, a transmembrane orientation for this region has never been proposed anywhere. Drp1 is a peripheral membrane protein that reversibly binds CL- and PA-containing membranes via its intrinsically disordered variable domain containing an aromatic-centered WRG motif. Indeed, the model presented in Figure 9 of Mahajan et al. displays a peripheral and parallel orientation of the transiently helical WRG-containing motif rather than a transmembrane (i.e., across the bilayer) orientation. While the authors can distinguish between a parallel vs. perpendicular orientation of this sequence relative to the plane of the membrane bilayer surface from their simulations, suggesting that previous studies indicated a transmembrane orientation for Drp1 is disingenuous and misleading. The term "transmembrane" should be removed or replaced, as it presents a wrong image.

      We have now deleted the sentence mentioning “transmembrane orientation”.

      (2) Mutational analysis of W vs. F in membrane insertion of W-centered insertion motifs and vice versa: The PPM-based workflow suggests that F-centered sequences have the highest membrane insertion properties as opposed to W-centered ones. A W552F mutation in the WRGML sequence of Drp1 was, however, found to impair function. How do the authors rationalize this? A crossmutational analysis of W vs. F in W-centered motifs and F-centered motifs is warranted. 

      AroMIP predicts a membrane insertion propensity of 0.782 for the WRGML sequence and a moderately higher propensity, 0.837, with a W552F mutation. This increase contradicts the experimental observation of a 3.6-fold increase in membrane binding affinity by Mahajan et al. We now speculate that the specific lipid, cardiolipin, as the reason for the discrepancy (p. 19, 3rd paragraph). This discrepancy provides a concrete example for the need to account for membrane composition in future developments.

      Recommendations for the authors:

      Reviewing Editor Comments:

      (1) The membrane composition used in this study is highly specific. The manuscript would benefit from a clearer justification of the lipid composition and discussing the transferability of the approaches to other relevant membrane systems.

      We now acknowledge the limitation of our work to a single composition (p. 19, 3rd paragraph). This composition was chosen because it is widely used in both computational and experimental studies (e.g., PMID: 21144818; 21344950; 29845130; 29995324; 34813727; 37406927). However, as we now point out, future developments should account for the effects of lipid composition.

      (2) This work heavily relies on computational modeling. Thus, it remains unclear to what extent the model captures the thermodynamics of membrane insertion, rather than reproducing the behavior of the PPM framework. Further comparisons with available experimental results will make this work more impactful.

      We note that the manuscript already has substantial experimental support. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM. Lastly, we now show good correlation between our insertion scores and binding free energies calculated from the scale determined experimentally by White and co-workers (new Figure S10; p. 15, second paragraph).

      (3) Overall, the manuscript is lengthy. Shortening will improve the readability, clarity, and accessibility to a broad audience.

      We have shortened some text, as explained below.

      Reviewer #1 (Recommendations for the authors):

      (1) The manuscript uses a single membrane composition… The high PIP<sub>2</sub> content (5%) in the simulations may overemphasize electrostatic contributions from basic residues. Please discuss how different membrane compositions (e.g., lower <sub>2</sub>, presence of cholesterol, cardiolipin in mitochondrial membranes) might alter the q parameters and whether AroMIP predictions would change qualitatively or quantitatively.

      We now discuss how different membrane compositions may alter q parameters (p. 19, 3rd paragraph). We believe that these alterations will change our predictions quantitatively but not qualitatively, given that our validation is against experimental results acquired on a variety of membrane compositions.

      (2) The limitations of the current framework should be discussed more explicitly. For example, the applicability of AroMIP beyond isolated 9-residue motifs remains unclear. In full-length IDPs, membrane insertion may be modulated by longer-range sequence interactions, transient secondary structure formation, multivalent interactions, or post-translational modifications.

      We now discuss the limitation of the 9-residue motif, and note these neglected factors for future developments (paragraph running from p. 19-20).

      (3) The AroMIP web server is user-friendly… The utility of the server could be enhanced by enabling visualization of full-length disordered proteins. For example, if users input a >300 residue IDP, the server could output residue-wise or sliding-window membrane insertion propensity profiles.

      We have revised the web server. We now display insertion propensity profiles as a plot and have added a link for users to download.

      (4) The manuscript is quite long and in several sections overly descriptive, e.g., in the first MD Results section and portions of the "Additional test cases" section.

      We have shortened the first subsection in Results and placed the expanded presentation in Supporting Information. However, we have kept the “Additional test cases” subsection, because Reviewer 2 appears to have overlooked the experimental validation of our method in these additional test cases.

      (5) The authors used the Berendsen barostat… known not to reproduce correct volume fluctuations and is generally considered less rigorous for equilibrium simulations, although this does not affect the main conclusions of the work. For future studies, the authors may consider using more modern barostats.

      Thank you for the suggestion! We will definitely be using the more modern barostats in future studies.

      Reviewer #2 (Recommendations for the authors):

      (1) The idea of the article seems very interesting. The problem of membrane association mediated by aromatic residues is definitely worth studying. Aromatic residues, especially Tryptophan (W), but also, albeit to a lesser extent, Phenylalanine (F), and Tyrosine (Y) are well known to partition preferentially to the headgroup region of the lipid bilayer. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev. Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602608; Killian & von Heijne, TIBS 2000, 25, 429-434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, however, none of these articles is cited. Have the authors read them?

      We now cite the relevant references as explained above.

      (2) The authors propose to decipher the sequence code for insertion of sequences containing aromatic residues in the membrane employing three types of calculation methods with decreasing order of detail and complexity, but increasing order of efficiency. First, all-atom MD simulations; second, the PPM method (protein positioning in membranes) from Lomize et al (2006), Protein Sci 15, 1318; and third, AroMIP, a mathematical model developed by the authors. Incidentally, I don't see anywhere in the text what AroMIP stands for. Is it Aromatic Membrane Insertion Prediction, or something like that? Please define. In any case, the proposed endeavor is commendable.

      We now spell out the acronym at its first occurrence in the Abstract and Introduction (Aromatic Membrane Insertion Predictor).

      (3) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth"." Thus, there must be a comparison with experiment, as elaborated in the next point.

      We emphasize that we have presented substantial experimental support for AroMIP. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM.

      (4) I understand that the authors are computational or theoretical physical chemists and would not expect them to perform the experiments. However, there are plenty of data in the literature that can be used to corroborate the calculations. First and foremost, though, the authors must calculate the binding constants for their set of peptides and then compare them with experiment, or use a different set of peptides for which the experimental results are available, and test how PPM and AroMIP perform on those peptides. There are two possible approaches to calculate the binding constants. First, calculate the Gibbs energy of binding from simulation or calculation, and, from that, calculate the binding constant via the Boltzmann factor. Second, and better, is to calculate the probability of binding in the simulations or calculations from the fraction of time that the peptide spends bound to the membrane or in water. According to the ergodic principle, the ratio of the two times is the binding constant. I understand that most amphipathic peptide sequences whose binding constants to membranes have been determined by experiment are long. But some are not. For example, Mastoparan X is a 14-residue antimicrobial peptide, and its dissociation constant from POPC vesicles is known to be about 300 μM.

      We would like to make it clear that the aim of our study is to predict membrane insertion propensities of aromatic-centred motifs, not the membrane binding affinity of peptides. These two properties are related but require different ways of validating predictions. For membrane insertion, validation requires that not only the motifs are bound to membranes but also the aromatic side chains are placed in the acyl chain region. Our validation of the 12 IDRs listed in Table S2 targeted these requirements.

      That said, we note that free energy of binding and probability of binding calculations, mentioned by the reviewer, have been reported previously, including the free-energy cost of Ala substitutions of aromatic residues located at various depths reported by Waheed et al. (ref 35) and residue-specific insertion depths reported by Wang et al. (ref 13). Both of these studies highlighted the propensities of aromatic side chains in inserting into the acyl chain region.

      (5) Furthermore, the entire Wimley-White interfacial hydrophobicity scale was determined using pentapeptides, measuring the equilibrium binding constant to POPC membranes experimentally and then using those data to build a residue-based Gibbs energy of binding (Wimley & White, Experimentally determined hydrophobicity scale for proteins at membrane interfaces. Nature Struct. Biol. 1996, 3, 842-848; White & Wimley Membrane protein folding and stability: Physical principles. Annu. Rev. Biophys. Biomol. Struct. 1999, 28, 319-365.) This allows for the calculation of the binding affinity for any sequence. The calculations have been compared to experiment and shown to have a high accuracy. A note of caution: Be careful, though, because Wimley and White used a mole fraction concentration scale, which makes the Gibbs energy of binding more favorable than the value calculated using the more common molar concentration scale by -2.4 kcal/mol. Calculating the binding constant for the peptides used by the authors from the WW interfacial scale and comparing the results with the authors' results would be a good place to start.

      This is an excellent suggestion! We now compare our insertion scores with the binding free energies calculated from the WW interfacial scale (new Figure S10; p. 15, 2nd paragraph). Interestingly, we found moderately higher correlations with binding free energies calculated from the octanol scale, which we suggest is more in line with our insertion scores since octanol mimics the hydrophobic region as suggested by White and Wimley in their 1999 Annu Rev paper.

      (6) When we speak of insertion in a bilayer, we normally mean insertion in the nonpolar core, not in the interfacial region, which is what the authors mean in this paper. This is extremely misleading and should be changed, namely in the title, but also throughout the paper.

      Actually, by insertion we precisely mean into the nonpolar core, NOT the interfacial region. Throughout the Introduction, when we used the word “insert”, we added “into the acyl chain region” (p. 3, line 6 from bottom; p. 5, lines 6-7 from top and line 7 from bottom; p. 6, line 4 from top). To avoid any confusion, we now also explicitly add “into the membrane hydrophobic core” when the word “insertion” first occurs in the Abstract and in the opening paragraph of Introduction (replacing the previous “deep insertion”).

      (7) What is q? It appears in equation (1) on page 14, and is referred to several times afterwards, but it is never defined.

      We now elaborate, in the text above and below equation (1), on the meaning of the q parameters: they represent the contributions of flanking residues to the insertion score of a central aromatic residue.

      (8) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is wrong even in a qualitative sense.

      We have revised these figures.

      (9) In Figure 1, the location of the bilayer midplane should be indicated, for example, with a line. Currently, there is a red line on the figures, but that is not the bilayer midplane. A reader may easily - and is likely to - misinterpret the figure.

      As explained in the Figure 1B, C caption, the red line is drawn at Z = -3.1 Å, which is the mean position of glycerol C2 carbon atoms (setting Z = 0 for the phosphate plane) and thus the start of the acyl chain region.

      Reviewer #3 (Recommendations for the authors):

      (1) Please refrain from using the term "transmembrane" when describing Drp1-membrane insertion, as Drp1 is a soluble, peripheral membrane-binding protein that reversibly associates with membranes.

      We have removed the sentence mentioning “transmembrane orientation”.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Naim et al. use genetically engineered mouse models and tissue culture cell lines to investigate the role of the SLAP adaptor protein in colonic epithelium and colon tumour formation. The SLAP adaptor protein is known to be a negative regulator of tyrosine kinase signaling in hematopoietic cells, but its role outside the immune system is less well defined. Here, the authors use genetically engineered SLAP-deficient mice, tissue-specific SLAP KO, and colonic organoids to demonstrate that SLAP is expressed in cells of the colonic epithelium, where it acts as a cell-autonomous regulator of proliferation and differentiation. In addition, they provide biochemical evidence that loss of SLAP expression in cultured colonic organoids results in increased Src family kinase activity and global tyrosine phosphorylation, consistent with its known role as a suppressor of tyrosine kinase activity in immune cells. Consistently, treatment with an SRC kinase inhibitor inhibited the growth of SLAP-deficient organoids. These data provide solid evidence of a cell-autonomous role of SLAP in the colonic epithelium.

      This work would be improved by further description and interpretation of the SLAP expression pattern shown in the constitutive and tissue-specific KO to further support the conclusions made. In Supplementary Figure 1, magnification of the colon epithelium areas with SLAP expression shown by b-gal and anti-SLAP staining, highlighting regions of interest, would better support the conclusions regarding SLAP expression in specific regions of the colon epithelium. In Supplementary Figure 1B, the authors should indicate that the SLAP staining referred to is epithelial and in resident immune cells, as is mentioned in the text. Also, magnification of the boxed area of LRG5 staining in Figure 1 would improve this figure.

      We thank the reviewer for their positive and constructive evaluation of our work.

      We have revised Fig 1 and S1 to better highlight SLAP expression patterns. Specifically, we have included higher-magnification images of the colonic epithelial regions, with clearly indicated regions of interest (new Figure S1). We have also clarified in the legend that SLAP staining is observed in both epithelial and resident immune cells, as described in the text. Additionally, we have provided a magnified view of the boxed area showing LGR5 staining in Figure 1 to improve clarity.

      Using a chemically induced model of colitis-associated cancer, the authors demonstrate that inactivation of SLAP shows a trend toward increased tumor formation (though this did not reach significance) as well as increased Src family kinase activity within tumors. Tumor spheres from SLAP-deficient animals showed enhanced growth that was suppressed by treatment with a Src family kinase inhibitor. Of note, the latter effect was specific to SLAP-deficient tumor spheres. These observations are convincing and support the authors' conclusion that SLAP has a tumor suppressor role in CRC through inhibition of SFK signaling.

      Mechanistically, elevated expression of the RTK, EphB2, was detected in immunoblots of SLAP KO colonic crypts, while overexpression of SLAP in CRC cell lines downregulated EphB2 protein levels. Using an EPHB2 inhibitor, the role of EPHB2 in the growth of SLAP-deficient colonic organoids was demonstrated. While these data generally support the authors' conclusion that SLAP limits colonic organoid growth by downregulating RTKS such as EphB2 and downstream Src family kinase activity, they do not show which cell types/regions in the colonic epithelium have increased EPHB2 protein and how this relates to SLAP and phospho-SRC expression, as shown in Figure 1 and Figure S1 immunocytochemistry. The expression of EphB2 and its role in colonic tumorsphere growth were not investigated.

      Overall, this work provides evidence of SLAP adaptor function in restricting tyrosine kinase signaling in the colonic epithelium, and suggests that loss of SLAP expression could promote tumorigenesis in this context.

      We thank the reviewer for their positive assessment of our tumour studies and for recognizing the evidence supporting a tumor suppressor role for SLAP through inhibition of SFK signaling.

      To address the reviewer’s mechanistic concerns, we performed additional experiments that are now included in the revised manuscript. We confirmed that loss of Slap is associated with increased EPHB2 expression in colonic crypts by IHC (new Figure 4B) and directly tested the role of EPHB2 in the Slap-deficient phenotype: EPH inhibition reduced both pTyr levels and SRC activation in Slap-deficient organoids (new Figure S2), demonstrating that SFK hyperactivation depends on upstream EPHB2 signaling. Consistent with this mechanism, we also observed increased EPHB2 tyrosine phosphorylation and active SRC (pSRC) association in isolated colonic epithelial cells following Slap deletion (new Figure 4A).

      To extend these findings to the tumour context, we examined the effect of EPH inhibition in human CRC tumoroids (new Figure 5). Pharmacological EPHB2 inhibition reduced tumoroid growth in CRC cells expressing low levels of SLAP, whereas this effect was largely lost upon SLAP overexpression. An SH2- or SH3-inactivating point SLAP mutant failed to suppress tumoroid growth and restored sensitivity to EPHB2 inhibition, further supporting EPHB2 as a critical target of SLAP-mediated tumour suppression. Together, these new data identify EPHB2 as a critical upstream activator of SRC that is negatively regulated by SLAP and strengthen our conclusion that deregulated EPHB2-SRC signaling drives the hyperproliferative phenotype associated with SLAP loss.

      Reviewer #2 (Public review):

      Summary:

      Protein tyrosine kinases are subject to diverse regulatory mechanisms controlling their activity in normal situations. The authors previously identified SLAP (Src-like adaptor protein), a negative regulator of receptor tyrosine kinase (RTK) signaling, as a key suppressor of the cytoplasmic tyrosine kinase SRC in the normal colon and demonstrated that SLAP is downregulated in a majority of colorectal cancers (CRCs).

      In this study, the authors further explored SLAP functions in mouse models using constitutive and inducible epithelial-specific Slap deletion (villin-CreERT2 model). They found that loss of SLAP augments colonic epithelial cell proliferation and that induction of tumorigenesis by the AOM/DSS protocol mimicking CRC leads to more aggressive tumors in the absence of SLAP. This effect is apparently cell-autonomous as growth of normal and tumoral colonic organoids is SLAP-dependent in in vitro settings. Finally, the authors define that, in colon, SLAP represses EphB2, an RTK lying upstream of SRC, and show that inhibitors of EphB2 can partially limit tumorigenic development in vitro.

      Strengths:

      The manuscript is clearly and concisely written, making it easy to follow. The data obtained in the mouse models are very convincing.

      Weaknesses:

      Direct evidence that EphB2 is activated/phosphorylated in the absence of SLAP is lacking, as conclusions are only based on results obtained with inhibitors. Some other issues have to be addressed before acceptance, in particular, the relevance of the findings in CRC patients.

      We thank the reviewer for their positive and constructive evaluation of our work.

      We agree that direct evidence linking SLAP loss to activation of the EPHB2-SRC pathway would strengthen the study. To address this point, we performed additional experiments that are now included in the revised manuscript. In addition to demonstrating increased EPHB2 expression upon Slap deletion, we found that loss of Slap enhances the EPHB2 tyrosine phosphorylation (an index of EPHB2 activity) and association between EPHB2 and pSRC in isolated colonic epithelial cells, supporting increased signaling through this pathway (new Figure 4A). Furthermore, pharmacological inhibition of EPHB2 reduced both SRC activation and the hyperproliferative phenotype observed in Slap-deficient organoids (new Figure S2). Together, these findings provide functional and biochemical evidence that deregulated EPHB2 signaling contributes to SRC activation in the absence of SLAP. We also examined the effect of EPHB2 inhibition in human CRC tumoroids (new Figure 5).

      Pharmacological EPHB2 inhibition reduced tumoroid growth in CRC cells expressing low levels of SLAP, whereas this effect was largely lost upon SLAP overexpression. Together, these new data identify EPHB2 as a critical upstream activator of SRC that is negatively regulated by SLAP and strengthen our conclusion that deregulated EPHB2-SRC signaling drives the hyperproliferative phenotype associated with SLAP loss.

      To address the relevance of our findings in CRC patients, we also extended our analyses to human datasets (new Figure 5D). We observed a significant inverse correlation between SLAP expression and a colorectal cancer stem cell-like activity score in TCGA tumours. In addition, co-expression of SLAP and SLAP2 with EPHB2 was associated with improved disease-free survival in microsatellite-stable (MSS) CRC patients, whereas no such association was observed in microsatellite instability (MSI) tumours. These findings support the clinical relevance of the SLAP-EPHB2 signaling axis and are consistent with a role for SLAP in restraining EPHB2-dependent CSC signaling in CRC.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Both reviewers have reported that the study of the EPHB2-SLAP-Src axis is not very developed, as most results are derived from using inhibitors. A key question is whether EPHB2 is activated by SLAP depletion and whether it is critical to SRC activation. Addressing these questions would greatly improve the paper.

      We thank the Reviewing Editor for this important suggestion. We would be happy for the editors to assess the revised version without involving the reviewers again. In the revised manuscript, we have substantially strengthened the mechanistic link between SLAP loss, EPHB2 activation, and SRC signaling. We show that Slap deletion increases both EPHB2 expression and tyrosine phosphorylation, enhances EPHB2-pSRC association in colonic epithelial cells. Importantly, pharmacological inhibition of EPHB2 suppresses SRC activation and rescues the hyperproliferative phenotype of Slap-deficient organoids. We further demonstrate that EPHB2 inhibition selectively impairs growth of CRC tumoroids with low SLAP expression, whereas this effect is largely abolished upon SLAP overexpression. An SH2- or SH3-inactivating point SLAP mutant failed to suppress tumoroid growth and restored sensitivity to EPHB2 inhibition, further supporting EPHB2 as a critical target of SLAP-mediated tumour suppression. Together, these new biochemical and functional data establish EPHB2 as a critical upstream activator of SRC that is negatively regulated by SLAP and significantly reinforce the central conclusions of the study.

      Reviewer #1 (Recommendations for the authors):

      (1) Evidence of SLAP expression in the colon is an important basis for these studies and could be moved to the main Figure 1 rather than being in the supplementary material.

      We thank the reviewer for this suggestion. While we agree that documenting SLAP expression in the colon is important, we have retained these data in Figure S1 to maintain a concise main figure set, consistent with the recommended format for Short Reports.

      (2) Define AOM/DSS and briefly describe the model at first mention. In addition, the model in 3A includes TAM treatment at 45 days, but this is not mentioned in the text. Why is this done?

      We have better defined the AOM/DSS protocol at first mention in the revised manuscript and specified the rationale for tamoxifen administration at day 45, which is required to maintain efficient SLAP deletion throughout the duration of the experiment (90 days).

      (3) Evidence that the SRC inhibitor decreased phospho-tyrosine levels in addition to inhibiting the growth of organoids should be included.

      We included data showing the inhibitory effect of the used SRC inhibitor on global phospho-tyrosine levels in organoids in the revised manuscript.

      (4) Further experiments investigating the involvement of EphB2 in colonic tumor formation are of interest and would increase the significance of this work.

      We included data showing that SLAP modulation affects the response of tumoroids derived from cell lines to EphB2 inhibition, providing complementary mechanistic insights.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should confront their findings with data obtained in normal and pathological tissues: are SLAP, SRC, and EphB2 co-expressed, at the single cell level, in normal colon and CRC? In which cell populations? Is loss of SLAP associated with poor prognosis in CRC patients?

      We thank the reviewer for this important suggestion. We agree that assessing the relevance of the SLAP-EPHB2-SRC axis in human CRC is important. However, transcriptomic datasets have inherent limitations in this context, as SLAP primarily regulates signaling at the post-transcriptional level and SRC activity cannot be reliably inferred from mRNA expression.

      To address the clinical relevance of our findings, we performed additional analyses of CRC patient datasets. We found a significant inverse correlation between SLAP expression and a colorectal cancer stem cell-like activity score in TCGA tumours. Furthermore, co-expression of SLAP and SLAP2 with EPHB2 was associated with improved disease-free survival in microsatellite-stable (MSS) CRC patients. These findings are consistent with a role for SLAP in restraining EPHB2-dependent signaling in CRC. Finally, while our data identify EPHB2 as a critical upstream regulator of SRC signaling controlled by SLAP, we do not exclude the possibility that additional receptor tyrosine kinases contribute to the effects of SLAP loss during colorectal tumorigenesis.

      (2) In Figure 4A, total EphB2 levels are increased in the absence of SLAP in colonic crypts. However, the level of EphB2 phosphorylation is not shown. This is an important point to address. Which ligand(s) may activate EphB2?

      We agree that assessing EPHB2 activation is important. To address this point, we have included new data in the revised manuscript showing that Slap deletion increases EPHB2 tyrosine phosphorylation in isolated colonic epithelial cells, providing direct evidence that EPHB2 signaling is enhanced in the absence of SLAP. In addition, we show that loss of Slap increases EPHB2-pSRC association and that pharmacological inhibition of EPHB2 reduces SRC activation and suppresses the hyperproliferative phenotype of Slap-deficient organoids. Together, these findings establish EPHB2 as a critical upstream regulator of SRC signaling following SLAP loss.

      Regarding EPHB2 activation, previous studies have shown that EPHB2 is primarily activated by ephrin-B ligands expressed within the intestinal crypt compartment (Batlle et al., 2002). EPHB2 signaling may also be reinforced through cooperation with other Eph receptors, particularly EPHB3, which is highly expressed in intestinal stem and progenitor cells (Holmberg et al., 2006; Genander et al., 2009). In addition, SRC has been reported to phosphorylate EPH receptors, raising the possibility of bidirectional signaling that could further amplify EPHB2-SRC pathway activity (Leroy et al., 2009; Hochgräfe et al., 2010).

      (3) In the absence of SLAP, inhibitors of EphB2 should also decrease SRC activity as EphB2 lies upstream of SRC (Figure S3C). Does this occur in organoids?

      We now show that EphB2 inhibition reduces SRC activity in SLAP-deficient organoids.

      (4) What is the status of EphB2 and SRC (total, phosphorylated) in SW620 and HT29 CRC cells in the absence of SLAP?

      SW620 and HT29 cells are SLAP-low CRC models. Given that SLAP expression is already minimal in these cells, further depletion is unlikely to provide meaningful additional insight.

      (5) Expression of SLAP is associated with a decrease in the stem cell compartment in CRC cell lines (Figure S2). Is there a stem cell signature associated with low SLAP levels in CRC?

      We analyzed TCGA colorectal cancer datasets using a published colorectal cancer stem cell (CSC) signature. We found a low but significant inverse correlation between SLAP expression and the CSC-like activity score, supporting our experimental observations that SLAP restrains stem cell properties in CRC cells and organoids.

      (6) Does overexpression of the mutant form of SLAP (SLAPmut) limit SLAP effects in SW620 and HT29 CRC cells in Figure S2?

      We have now performed the requested experiments and found that, unlike wild-type SLAP, SLAPmut failed to inhibit tumoroid growth in CRC cells. These results are consistent with our previous findings showing that SLAPmut lacks tumour suppressor activity in CRC cells (Naudin et al., Nat Commun, 2014) and further support the requirement of SLAP signaling functions for the regulation of CRC stem-like properties.

      (7) Total SRC is missing in Figure 2B

      Total SRC levels are now included in the revised figure.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary

      The manuscript by Kostanjevec et al. investigates the mechanism behind spiral pattern formation in the cornea. The authors demonstrate that the spiral motion pattern on the mammalian corneal surface emerges from the interaction between the limbus position, cell division, extrusion, and collective cell migration. Using LacZ mosaic murine corneas, they reveal a tightening spiral flow pattern and show that their cell-based, in silico model accurately reproduces these patterns without global guidance cues. Additionally, they present a continuum model that extends the XYZ hypothesis to describe cell flux on the cornea, offering a quantitative explanation for tissue-scale processes on curved surfaces.

      Strengths

      The manuscript is well-written, with a systematic approach that clearly explains experimental setups, model construction, assumptions, parameter selection, and predictions. The discussion also provides insightful perspectives on the broader implications of the results for both physics and biology.

      We thank the reviewer for their positive assessment of the manuscript. We are pleased that the reviewer found the work well written and systematic, and that the experimental design, model construction, assumptions, parameter selection, predictions and broader discussion were clearly presented. We have aimed to preserve these strengths in the revised manuscript while substantially expanding the discussion and analysis in response to the reviewer’s concerns.

      Weaknesses

      The central premise of the manuscript, that the spiral patterning of epithelial corneal cells occurs without guidance cues, is not fully supported. The authors overlook the potential role of axons in guiding epithelial cells, despite clear evidence of spiral axon patterns in their own Fig. 1b. Previous literature indicates that axon patterning precedes epithelial cell patterning, suggesting that epithelial migration might be influenced by pre-existing neural structures (e.g., Leiper et al. 2002, IOVS 2013). The authors need to address this point, possibly by exploring whether axonal patterns serve as a template for epithelial cell migration, or by providing experimental evidence to rule out axon-based guidance.

      The reviewer raises an important point that we now address in the revised manuscript. We respectfully disagree with the assertion that the premise of our work is faulty. At no point do we claim that global guidance cues are absent or ignore the fact that the corneal nerves swirl; rather, our results show that such global or contact-mediated cues are not required to explain the observed swirling patterns of epithelial cell radial migration in the adult cornea.

      Although our work shows that a swirling prepattern of nerves is not required to obtain radial patterns of epithelial cell migration, we now address why nerve swirling occurs and how it affects interpretation of our data. Previous literature indicates that nerve swirling is visible from approximately 3 weeks, before epithelial striping patterns become apparent in transgenic LacZ and GFP reporter mosaics at about 5 weeks. However, both experimental observations and our modelling indicate that spiral epithelial cell migration proceeds for many days before reporter stripe patterns become visible. Thus, epithelial migration may already be underway before epithelial striping is detectable. If axons are following epithelial migration, they would therefore be expected to become radially aligned before the reporter epithelial stripes are evident.

      We also considered the alternative possibility that epithelial cells follow an axonal prepattern. However, several experimental observations argue against the hypothesis that epithelial migration is primarily guided by axonal projections. In situations of genetic mutation or corneal injury, epithelial cell migration can progress independently of, or ahead of, corneal axon extension. In chimeric Pax6+/− LacZ<sup>+</sup> ↔ Pax6<sup>+/+</sup> LacZ<sup>-</sup> mice, where the normally disrupted radial migration of Pax6+/− epithelial cells is restored, the underlying nerves may continue to exhibit abnormal projection patterns. These findings suggest that axonal organization is at least partly dependent on epithelial behaviour rather than the reverse.

      Importantly, we do not exclude the possibility that additional cues, including axonal contact, neurotrophins or other environmental signals, may modulate or refine epithelial migration. We have therefore expanded the revised manuscript to discuss these possibilities and the relevant literature more fully.

      While the model is well-constructed, it currently falls short of its stated goal of elucidating the mechanisms of spiral formation. Key questions remain unanswered:

      Is the curvature of the cornea necessary for spiral formation, or would a simpler disk geometry suffice?

      What role do boundary conditions play?

      How well do the model's predictions quantitatively match experimental data?

      The current comparisons in Fig. 4c-f lack quantitative agreement, and this discrepancy should be discussed with possible explanations.

      We thank the reviewer for identifying these points, which we have now addressed in the revised manuscript.

      First, we have examined the role of geometry more systematically. The spiral pattern also appears in a disk geometry, and the same qualitative migration pattern is observed across a broader range of simulated geometries, including different curvatures, cap angles, a prolate ellipsoid, an oblate ellipsoid and a disk. The precise shape of the spiral depends on geometric features such as curvature and cap-angle opening, but spiral formation is robust across convex cornea-like geometries. These results are now discussed in the new section “Robustness of the spiral migration pattern and requirement for limbal stem cells” and shown in the revised Fig. 10.

      Second, we have clarified the role of boundary conditions, particularly the role of limbal epithelial stem cell proliferation. Without limbal stem cells, and with all other parameters unchanged, the cornea fails to produce the radial striping pattern. At the alignment strength where robust spiral formation normally appears, the simulated tissue flow is disordered and resembles our in vitro calibration simulations. At higher alignment strengths, spiral formation is still not recovered; instead, defects become anchored to the boundary. These findings show that ordered influx from the limbus is important for promoting the spiral state. They are now discussed in the same new section and shown in revised Fig. 9.

      Third, we have revised the quantitative comparison between model and experiment. We agree that the original comparison was limited. We have therefore reanalysed both experimental and simulation data, focusing on the time-averaged hydrodynamic velocity field rather than short-range fluctuations amplified by divisions in the numerical model. We also replaced the Fourier-space velocity correlation functions with spatial velocity correlation functions, which are more directly interpretable. The revised analysis shows that the experiments have mesoscale spatial and temporal correlations, of the order of 5-6 cell sizes in space and about one hour in time, and that these are well captured by the simulations for both plastic and explant substrates. We have also added representative experimental and simulation snapshots in revised Fig. 4g-j.

      For the full cornea, we acknowledge that direct quantitative comparison remains limited by the available experimental data. We can compare with inferred migration direction fields and resurfacing timescales, but we do not yet have direct live measurements of the full corneal velocity field.

      The authors emphasize polar alignment as a key feature of the spiral pattern based on simulation results. However, they do not provide experimental evidence for this polar alignment. The manuscript includes discussions of polar and nematic symmetries that, without supporting data, feel somewhat distracting. If direct experimental evidence for polar alignment is not available, the authors could instead quantify nematic alignment as the spiral forms. This would also allow them to explore potential crosstalk between nematic cell orientation and the polar alignment of self-propulsion, especially considering recent studies showing alternative mechanisms for vortex formation in similar systems.

      We thank the reviewer for pointing out that the discussion of polar and nematic alignment was confusing. We have substantially revised this part of the manuscript.

      We agree that we do not have direct experimental evidence for polar alignment. However, several observations support the interpretation that the system is dominated by substrate-based polar motility with weak polar alignment. We have now quantified nematic alignment of cell orientations and find no evidence of significant elongation or local nematic order in corneal epithelial cells, as shown in the new Fig. 3d.

      Our computational model begins from uncorrelated, substrate-based polar active cell migration. We tested whether polar alignment is necessary and found that the in vitro data are inconsistent with the complete absence of alignment: the flocking order parameter is too high, and the spatial and temporal correlations are larger than expected from persistent driving alone.

      We also discuss that cell-cell stress patterns in the epithelium may remain nematic, but any such effect must be sufficiently weak not to dominate the observed axon motion. We have further revised the discussion to include recent related work showing spiral formation in substrate-based cell migration models with polar dynamics, as well as recent work indicating that nematic-like phenomenology can arise from minimal ingredients such as uncorrelated polar activity and cell deformability. This revision is intended to make the interpretation clearer and to avoid overemphasising unsupported claims.

      Reviewer #2 (Public review):

      In K. Kostanjevec et al, the authors study a possible mechanism for the formation of spiral patterns in the cornea. First the authors analyze an inferred velocity field, which is deduced from images of fixed corneas, and then determine the position-dependent spiral angle of this velocity fields. Next, the authors analysed two possible markers of cell polarity: the direction of the centrosome-nuclei and the axis of mitosis. Then the authors introduce a stochastic agent-based model of self-propelled particles with over-damped dynamics and with aligning interactions to the orientation of the nearest neighbors and to the particle's velocity. The authors claim to be able to reproduce the equal-time autocorrelation function and the velocity Fourier spectrum. Then the authors introduce the geometry of the cornea by constraining the dynamics on a spherical cap and show that their model can reproduce a typical trajectory in experiments. Finally, the authors produce a phase diagram of the states at a fixed time point as a function of the spherical cap radius and the strength of the coupling aligning constant. Finally, the authors propose an interpretation of the cell fluxes based on the equation of mass conservation.

      We thank the reviewer for their careful assessment of the manuscript and for recognising the work as a solid theoretical study. We have revised the manuscript substantially in response to the reviewer’s major concerns, particularly regarding the terminology of topological defects and stagnation points, the comparison with experiments, and the role of corneal geometry.

      Regarding the terminology of topological defects, we have clarified the distinction between a velocity-field stagnation point and the topological classification of the velocity direction field. Stagnation points can be assigned a topological index when one considers the direction of the vector field away from the core, while ignoring the magnitude. We agree that the physical origin of interactions in a velocity field differs from that in a director field, and we have revised the text to avoid confusion. The revised manuscript now includes Box 1, which summarises the relevant topological concepts and caveats.

      Regarding the comparison with experiments, we have expanded and clarified the validation of the inferred velocity field. The LacZ reporter system was designed for lineage tracing and therefore reports coarse-grained cell motion and growth patterns. Direct live imaging of the full cornea remains technically difficult because of the macroscopic size, curved surface and long resurfacing time. However, live fluorescent reporter systems are consistent with our in vivo model and support the interpretation that inferred velocity fields recapitulate epithelial migration in vivo. We have also clarified the interpolation procedure using the XY model: approximately 30% of the corneal surface is covered by directly estimated velocity directions from stripe edges, rising to more than 50% near the central spiral. These measured directions provide sufficient boundary conditions for the annealing procedure to converge to a slowly varying field consistent with the observed stripe geometry.

      We have also revised the comparison between simulation and experiment. For in vitro data, we now compare spatial velocity correlation functions of the time-averaged hydrodynamic velocity field, rather than relying on Fourier-space correlations. The revised comparison shows good agreement between experiments and simulations for both plastic and explant substrates. For the full cornea, we acknowledge that the available quantitative data are limited to inferred migration direction fields and resurfacing timescales.

      Regarding the role of geometry, we have now simulated a wider range of substrate geometries, including spherical caps with different curvatures and cap angles, prolate and oblate ellipsoids, and a flat disk. The spiral migration pattern is robust across these convex geometries, although the detailed spiral shape depends on geometric features. We have also clarified that the boundary condition of inward limbal influx is crucial: without limbal stem cells, the radial striping pattern does not form, and defects may instead anchor to the boundary. These results are now shown in revised Figs. 9 and 10 and discussed in the new section “Robustness of the spiral migration pattern and requirement for limbal stem cells.”

      Overall, these additions clarify that the proposed mechanism relies on the combination of polar motility, weak alignment, limbal influx, cell division and extrusion, and cornea-like confinement, while also acknowledging the current experimental limitations.

      Recommendations for the authors:

      Reviewing Editor:

      The authors could strongly improve the manuscript by following the recommendations given below.

      Reviewer #1 (Recommendations for the authors):

      There are, however, substantial shortcomings that authors need to address to make their claims supported by enough evidence:

      Major concerns:

      (1) Neglect of Potential Axon Guidance

      As it stands the premise of the paper is unfortunately faulty. The authors overlook the potential role of axons in guiding epithelial cells, despite clear evidence of spiral axon patterns in their own Fig. 1b. Previous literature indicates that axon patterning precedes epithelial cell patterning, suggesting that epithelial migration might be influenced by preexisting neural structures (e.g., Leiper et al. 2002, IOVS 2013). The authors need to address this point, possibly by exploring whether axonal patterns serve as a template for epithelial cell migration, or by providing experimental evidence to rule out axon-based guidance.

      See for example:

      from: https://iovs.arvojournals.org/article.aspx?articleid=2126612:

      It is therefore not clear why the authors clearly show the axon vortex in Figure 1, then talk about prepatterning - and never mention the axon vortex again.

      The reviewer raises an important point that we now address in the revised manuscript. We respectfully disagree with the assertion that the premise of our work is faulty. At no point do we claim that global guidance cues are absent or ignore the fact that the corneal nerves swirl; however, our results show that such global or contact-mediated cues are not required to explain the observed swirling patterns of epithelial cell radial migration in the adult cornea. Although our work shows that a swirling ‘prepattern’ of nerves is not required to obtain radial patterns of epithelial cell migration, we address below why it occurs and how it affects our data.

      An intuitive interpretation of corneal anatomy would be that sensory axons follow the path of least resistance between migrating epithelial cells. We believe this is the most likely explanation in the normal wild-type cornea; however, the reviewer correctly notes that previous literature indicates that axon patterning precedes epithelial cell patterning. Nerve swirling is observed from approximately 3 weeks, earlier than the epithelial striping patterns that become apparent in transgenic LacZ and GFP reporter mosaics at about 5 weeks (e.g. Collinson et al., 2002 [PMID: 12203735]; Iannaccone et al., 2012 [PMID: 22347498]; McKenna and Lwigale 2011 [PMID: 20811061]). We now address this in the revised manuscript and below.

      The early appearance of radial axonal projections can be readily explained even if axons are following the epithelial cells. Experimental observations (Collinson et al., 2002 [PMID: 12203735]) and our modelling (Fig. 6a,b) both indicate that spiral epithelial cell migration proceeds for many days before stripe patterns become visible in reporter mosaics. Thus, epithelial migration is already underway before the striping pattern becomes detectable. If axons are following the epithelial migration, they would be expected to become radially aligned before the reporter epithelial stripes were evident. The apparent precedence of radial axonal projections over epithelial striping is therefore fully consistent with the biological scenario in which axons follow the migrating epithelial cells. Corneal epithelial basal cells have been shown to wrap around individual and grouped subbasal axons, acting as surrogate glia (Stepp et al., 2016, Investigative Ophthalmology & Visual Science 57, 1292), which represents a plausible mechanism to allow migrating epithelial cells to shepherd axons in their direction of movement.

      We considered that epithelial cells may follow an axonal prepattern, but several experimental observations argue against the hypothesis that epithelial migration is guided by axonal projections. In situations of genetic mutation or corneal injury, epithelial cell migration can progress independently of, or ahead of, the extension of corneal axons (Leiper et al., 2009 [PMID: 19029029]; Song et al., 2004 [PMID: 14744881]). Furthermore, in chimeric Pax6<sup>+/−</sup> LacZ<sup>+</sup> ↔ Pax6<sup>+/+</sup> LacZ<sup>-</sup> mice, where the normally disrupted radial migration of Pax6<sup>+/−</sup> epithelial cells is restored, the underlying nerves may continue to exhibit abnormal projection patterns (Leiper et al., 2009 [PMID: 19029029]). These results suggest that axonal organization is at least partially dependent on epithelial behaviour rather than the reverse.

      Importantly, none of this evidence excludes the possibility that additional cues (including axonal contact or other environmental signals) may modulate or refine epithelial migration. For example, Walczysko et al. 2016 [PMID: 27563231] showed that isolated epithelial cells cultured on de-epithelialised and de-nervated corneal stroma still migrate with a small but significant radial bias, indicating that epithelial cells can respond to physical features of their environment. Our model likewise includes a component of alignment with environmental structure. Since the central radial striping in many of our simulations is somewhat less ordered than in vivo, it is plausible that additional biological guidance cues or axonal neurotrophins help refine the pattern.

      We now discuss these issues and the relevant experimental evidence more fully in the revised manuscript.

      (2) Model Validation and Complexity

      While the model is well-constructed, it currently falls short of its stated goal of elucidating the mechanisms of spiral formation. Key questions remain unanswered:

      Is the curvature of the cornea necessary for spiral formation, or would a simpler disk geometry suffice?

      To address the Reviewers’ concerns about the role of geometry, we first note that the spiral pattern also appears in a disk geometry. In fact, the geometric parameters of the cornea vary across mammals, including humans, yet the spiral pattern is preserved (Dua, et al., 1993 [PMID: 8325424]; Zander and Weddell, 1951 [PMID: 14814019]). Similar variations arise in disease states, e.g. in keratoconus the cornea becomes elongated.

      On a spherical cap, and all shapes with the same topology, the boundary winding number fixes the interior index, so ongoing limbal influx maintains a total index of 1. To explore the role of geometry more systematically, we simulated a broader range of geometries, including different curvatures, cap angles, a prolate and an oblate ellipsoid, and finally a disk, and compared the resulting patterns with published data across mammals and with disease states.

      For all of these shapes, we find the same qualitative migration pattern – a central spiral, although its precise shape depends on geometric features such as curvature and cap-angle opening. These results are discussed in a new section “Robustness of the spiral migration pattern and requirement for limbal stem cells” and shown in Fig. 10 of the revised manuscript.

      What role do boundary conditions play?

      In the revised manuscript, we address the role of boundary conditions, in particular the presence of limbal epithelial stem cell proliferation. Without limbal stem cells, and with all other parameters kept unchanged, the cornea fails to make the radial striping pattern. We observe two distinct changes: First, at J=0.1, the amount of alignment at which robust spiral formation appears normally, the simulated tissue flow is instead disordered, in fact very similar to our in vitro calibration simulations (revised Figure 4). This shows that the boundary cue of an ordered influx from the limbus promotes flocking when it otherwise would not (yet) appear. Second, when we increase the alignment to J=0.15 or J=0.2, we still do not observe spiral formation. Instead of the expected central vortex shape, we have anchoring of defects to the boundary, facilitated by the effective compressibility of the tissue because of the density feedback in the division/extrusion rates.

      These findings are discussed in the new section “Robustness of the spiral migration pattern and requirement for limbal stem cells” and in Fig. 9 in the revised manuscript.

      How well do the model's predictions quantitatively match experimental data?

      The current comparisons in Fig. 4c-f lack quantitative agreement, and this discrepancy should be discussed with possible explanations.

      Regarding the in vitro cell data and matching simulations: We are aware that we have limited data to work with, and our match is intended as a rough estimate of physical parameters.

      For the revision, we have reanalysed both experimental and simulation data carefully. We realised that certain details of our numerical model amplify short-range fluctuations, namely the way cell divisions induce stress dipoles, and the fact that we did not include cell-cell friction forces. This is not merely a guess but emerged from related theoretical work by some of us (Keta and Henkes [PMID: 40556485]; Kammeraat et al, arXiv:2508.01046 (2025)).

      We therefore compared the time-averaged velocity field excluding divisions to the experiment, the same quantity that we already introduced as hydrodynamic velocity for the corneal surface. We also carried out further simulations that included cell-cell friction forces for comparison. They led to very similar results, albeit with a transition to flocking at somewhat lower alignment strength J.

      Instead of the Fourier-space velocity correlation functions that are hard to interpret, we switched to spatial velocity correlation functions. Our experiments have ‘swirly’ velocity fields with mesoscale spatial and temporal correlations, of the order of 5-6 cell sizes in space and one hour in time. As can be seen in the revised panels 4c,d for the spatial correlations of the hydrodynamic velocity, the match between experiment and simulations is in fact good for both plastic and explant substrates. There are also systematic changes in length and time scales between the two experiments that emerge without fine-tuning from simulation. We furthermore have included snapshots of both experimental in vitro conditions and matching simulations as new panels 4g-j, showing that the mesoscale correlations appear in the hydrodynamic velocity.

      Ultimately the parameter values and length and time scales that we infer for our systems (see Table 1) are quantitatively consistent with three other estimates of in vitro epithelia (Henkes et al. 2020 [PMID: 32179745]; Saraswathibhatla et al, Extreme Mechanics Letters 48, 101438 (2021), Kammeraat et al, arXiv:2508.01046 (2025)).

      For cell flows on the full cornea, regrettably we lack further quantitative data to compare to beyond the migration direction fields inferred in Figure 2, and the time scales of corneal resurfacing (Figure 6).

      (3) Importance of polar alignment.

      Polar alignment is put forward as one main feature for the observed patterns based on the simulation results. However, no experimental confirmation for such polar alignment is presented. The authors instead present various scattered discussions about polar versus nematic symmetry, which at times reads rather unnecessary and distracting from their main message.

      If they are not providing direct experimental evidence on the polar alignment, at least they could quantify nematic alignment of cell orientation as the spiral forms. One possibility is that nematic alignment of cell orientations has a crosstalk with the polar alignment associated with the self-propulsion of the cells. This is important because, as acknowledged by authors, several recent works in the context of in vitro epithelial under disk confinement have revealed alternative mechanisms for spiral vortex formation.

      We thank the Reviewer for pointing out the confusion, and we have therefore completely revised our discussion of alignment. While the evidence remains indirect, the following observations lead us to conclude that our system is dominated by substrate-based polar motility and weak polar alignment:

      We have investigated nematic alignment of cell orientations. As can been seen in new Figure 3d, corneal epithelial cells do not show evidence of significant elongation in any direction, i.e. we find no local nematic order.

      Our computational model starts from uncorrelated, substrate-based polar active cell migration. We have carefully investigated if polar alignment is a necessary ingredient and found that the in vitro data are inconsistent with the absence of alignment: the flocking order parameter is too high, and the spatial and temporal correlations are larger than those expected from persistent driving only.

      Still, cell-cell stress patterns in the epithelium may remain nematic. While one cannot exclude anything, the effect must be sufficiently weak to not affect axon motion. Their naturally long, thin shapes would strongly react to nematic stresses, but we do not see ±1/2 defects in their growth patterns, only polar ±1 defects.

      We were recently made aware the work of Lång et al. (2024) [PMID: 38630812], in which the authors also observe spiral formation in cell migration on a substrate, with +1 topological defects. They explain their observations using a polar model, and our results are consistent with their observations and model.

      In the active matter community, the ‘active nematic cell sheet’ paradigm is currently undergoing a revision. Notably, nematic-like phenomenology, in particular ±1/2 defects, can also arise from the minimal ingredients of uncorrelated polar activity and cell deformability (Chiang et al. 2024 [PMID: 39302997]).

      Minor comments:

      - It would help the reader to see some representative images of the cells, explants, and the simulation (maybe with vectors overlaid), as it could be hard to imagine what exactly is happening and the other relevant properties to compare.

      Please see new panels Fig. 4g-j, and new supplementary videos S3 and S4.

      - Add a colorbar that represents direction to Fig. 1a.

      We assume the Reviewer meant Fig. 2a. We have added a circular orientation colour chart to the image.

      - When do additional +1,-1 defects appear? The authors just say it is unlikely, but it is not clear when they are observed. Disease state?

      The reviewer is correct about pathology. We now cite (in conclusion) Collinson et al. 2004 [PMID: 15037575], which shows disruption and discontinuity in Pax6-mutant corneas with chronic corneal degeneration. We also have a publication in prep that shows extra +1 and -1 discontinuities occur during wound healing. We don’t want to include these data in this manuscript, but cite Sagga et al. (in prep).

      Reviewer #2 (Recommendations for the authors):

      Overall, the manuscript presents a solid theoretical work. However, I have a few major concerns on this manuscript. (1) The authors use the concept of topological defect to refer to stagnation points of the velocity field (2) The comparison to experiments remains qualitative and it is unclear whether the proposed mechanism is at work, (3) The role of geometry remains unclear.

      Major points:

      (1) On page 3 the authors claim that topological defects exist in velocity fields, however this statement is incorrect and can confuse readers. Unlike the director field of an ordered phase, their velocity field is a vector in R^2 with a norm that is not fixed, and therefore the velocity field has no topological defects. In fluid dynamics, these special points in the velocity field are called stagnation points, fluid sinks or sources. This distinction is important because the physical nature of the interaction forces between two "topological defects" in a velocity field is fundamentally different to the interaction forces between two topological defects in a director field. For these reasons, it can confuse readers to mix the two concepts (stagnation points vs topological defects). Note that if the authors address this concern, many parts of the main manuscript should be rewritten.

      Stagnation points are topological defects since the topological classification ignores the magnitude and looks only at the direction of the vector field away from the core. There are some caveats related to the role of boundary conditions, since in the case of a fluid, if boundary conditions are not fixed, one can eliminate the defect by setting the flow field to zero everywhere. Mathematically, a pair of point sources or sinks in an incompressible potential flow interacts through the same Green’s function that gives the elastic interaction between two-point disclinations in a 2D nematic director field, so their long-range pair potentials are formally identical (~ln(r) in 2D). The physical mechanism behind these forces is, as correctly pointed out by the reviewer, quite different. In addition, our coarse-grained velocity field is effectively compressible, which has consequences for the hydrodynamic equations we (can) write, see below.

      In the revised version, we clarify these points by adding Box 1, which summarizes the idea of topological defects.

      (2.1) How did the authors check that the velocity field that is inferred from the stripe edges matched the coarse-grained cell velocity field in a life sample? How did the authors validated the interpolation of the inferred velocity field using an XY model? What is the scale of the inferred velocity field? Can the authors clarify also this point?

      The murine cornea LacZ reporter system was specifically designed to allow for lineage tracing, i.e. following coarse-grained cell motion and growth patterns. The combination of macroscopic size (3.6 mm diameter), two-week resurfacing time and the curved surface however stymied our early attempts to directly measure the cell velocity field on the cornea using confocal time-lapse microscopy. However live sample fluorescent reporter systems are fully consistent with our in vivo model and shown conclusively that inferred velocity fields are recapitulated in vivo (Park et al., 2019 [PMID: 31843909]).

      For the XY model inference: We first note that approximately 30% of the corneal surface is covered by directly estimated velocity directions from the stripe edges (red arrows, Fig. 11f). This rises to more than 50% near the central spiral (red arrows, Fig. 11h) due to the way the stripes narrow near the centre due to cell extrusion. Therefore, the inferred areas are only slightly more than half of the cornea, and in regions where we expect the flow field to be largely uniform with no defects. The red arrows provide sufficient boundary conditions that a simple annealing simulation (or equivalently an energy minimisation) of the XY model rapidly converges to a slowly varying field consistent with those boundary conditions.

      Furthermore, in simulation we observe that stripe edges and the macroscopic velocity field correlate strongly with each other once the spiral has fully formed (Fig. 6b-c).

      The scale of the inferred velocity field is the distance between arrows in our digital version of the cornea, approximately 40 concentric rings over a 70° cone angle for a R = 1800 μm micron cornea, resulting in a spacing of 55 μm between velocity arrows. This is about at the scale of the in vitro velocity correlations. We are not able to obtain velocity magnitudes using this procedure.

      Furthermore, the spiral angle profile reported in Fig. 7b-c appear to be different to that found in experiments Fig. 2b. Can the authors clarify if their theoretical framework reproduces the spiral angle profiles?

      First, we note that empirically, simulations with the largest two radii (R = 1000 μm, R = 1500 μm), approaching the full experimental size, match the observed angle profile best. They both consist of a radially inward profile α(0) = 0° at the edges, only increasing, corresponding to a tightening spiral, below about θ = 20°. Note that very near the corneal centre, few cells contribute to data, and additionally the central defect position fluctuates somewhat. Therefore, the angle profile not reaching α(0) = 90° in the simulations is due to fluctuations and lack of statistics. As these simulations were run with our best fit experimentally matched parameters, this convergence is meaningful and cannot be scaled out.

      Second, our partial model (eq. 4, reproduced below) links the spiral profile angle α(θ) with the macroscopic velocity magnitude v(θ) and the net cell loss rate A(θ), and it depends explicitly on the corneal radius R.

      That is one equation for three radial fields. If we make the reasonable assumption that simulations at fixed 𝐽 that differ only in R have the same constitutive law A(ρ), and would follow the same continuum velocity equation that ultimately sets v(ρ), the radius R still appears explicitly in the equation. This is consistent with the different radial profiles for different R that we observe in Fig. 7c. Furthermore, in Fig. 7b we observe that above the flocking threshold, different J lead to very similar profiles. This indicates that the system enters a fully polar phase.

      The same is true in experiment: We expect that cell mechanics and planar cell polarisation coordination are local effects that will set J, A(ρ), and v(ρ). Thus, we do predict that the radius R will explicitly affect the spiral profile consistent with the equation above, but we would need more direct measurements of J, A(ρ), and v(ρ) to go any further.

      Properly answering this question would first require simulations of a non-dimensionalised model with different radii, boundary influx, alignment strengths and constitutive laws for the division / extrusion dynamics. Then one would want to construct a matching equation for the polarisation and / or velocity field, going beyond the Malthusian flock approximations of constant magnitude v(θ) and net cell loss rate A(θ) = 0. This is well beyond the scope of this publication, and there is certainly no experimental data to compare to.

      (3.1) It appears that the formation of the spiral velocity pattern results from a combination of a radial flow of cells due to cell division at the outer boundary and apoptosis at the geometrical center and azimuthal flow of cells due to flocking in confined geometries. Is this correct? Can the authors explain how does the 3d geometry of the spherical cap modify each of these flow fields?

      The Reviewer is broadly correct; however, the true picture is more subtle. Almost all of the divisions and extrusions occur in TA cells, neither at the limbus or at the corneal centre (please see the A(θ) profiles in Figure 8). Flow is then a spiral flock that is partially radial and azimuthal, and with a variable velocity magnitude. It ultimately all has to follow the flux equation 4, in steady state.

      For the modification of the flow field due to 3d geometry: Broadly, they simply change the amount of corneal surface available at different angles from the limbus when switching to isomorphic surfaces like, e.g. the disk. Thus, we still observe spirals, but of somewhat modified shapes. Please see the reply to Reviewer 1 above, and new Figure 10 for simulations of different geometries.

      A full theory of radially symmetric corneal shapes would again need additional velocity and constitutive equations and is beyond the scope of this publication.

      (3.2) In their model, it appears that the geometry is introduced by constraining the dynamics of agents, and it has not direct influence on the alignment of agents. Can the authors explain how the geometry of the spherical cap influences the emergent spiral states? Can the authors show whether their results are robust to changes in the substrate geometry? For example, by changing the substrate geometry from a spherical cap to another convex shape. Can the authors identify differences between the spiral patterns on a spherical cap vs that on a flat disk?

      Please see the response to Reviewer 1, above, and new Figure 10. Briefly, the results are robust to changes in the corneal geometry as long as they are other convex shapes. There are differences in details of the spiral shape.

      Minor comments:

      - On page 3, the authors claim that the Euler characteristic of a spherical cap or a disk is 1. Unfortunately, this statement is incorrect. The Euler characteristic of a spherical cap or a disk is determined by the winding of the vector field around the open boundary. In their case the Euler characteristic is 1 because the velocity field is oriented towards the top of the spherical cap, which give a winding of +1. The authors explain this correctly on page 13.

      The Euler characteristic is a topological invariant of the surface and does not depend on the tangent vector field. A disc or spherical cap has Euler characteristic 1. What depends on the boundary winding is the index formula for a vector field on a surface with boundary: the winding of the field along the boundary determines the corresponding boundary contribution, and hence the sum of interior indices. Thus, while the reviewer is correct about the role of boundary winding in computing the field’s index, this does not alter the Euler characteristic of the surface itself.

      We note that the orientation of the velocity field does indeed matter in our simulations: In new Figure 10, we show that in the absence of the boundary condition of inward flux at the limbus, we do not observe a spiral robustly, and we do see anchoring of defects to the boundary, with winding numbers that are now different.

      - On page 5, the authors claim the clockwise and counterclockwise -oriented spirals are equally likely, however no quantification is provided to support this claim. Can the authors clarify this point?

      The approximately equal likelihood is as described in Collinson et al. (2002) [PMID: 12203735].

      - On page 8 the authors state "Dipolar active force cannot cause a single cell to migrate, and the flow is an emergent collective phenomenon", here I was confused, because a bacteria swimming in a Newtonian fluid can self-propel by exerting a dipolar active force on the fluid. See for instance work by E. Lauga or I. Aronson. Can the authors clarify this point?

      The Reviewer is correct about how bacteria can swim using dipolar active forces. However, and unfortunately, that language was straight ported to very different conditions, that of cells migrating on a frictional substrate with no induced flow, with the equation of motion . Unless there is a net force arising from the stress profile, the cell cannot move, and with microscopic models where the active stress is a single value per cell, that statement is always true. Of course, more detailed cell models with active stresses exist, but their motion is still due to the net force arising from them. We have clarified the statement in the revised manuscript.

      - On page 16, the expression for the velocity seems to miss a parenthesis.

      Fixed. Thanks!

      - On page 16, the authors claim that the flux profiles in the simulations and in the XYZ model are in good agreement. However, there are clear differences for theta> 50 deg. Can the authors discuss the possible explanation for these differences? Note that one may expect a between agreement between two theoretical approaches.

      These disagreements are due to imperfections in the observed spiral patterns in simulations. They are not perfectly radially symmetric, and the defect is not always in the dead centre of the cornea. Thus, the theoretical predictions are not quite accurate. Numerically, what happens for θ > 50° is that the radial bins sometimes include part of the limbal zone with strong proliferation, and the corneal edge itself. Both are not described by the flux prediction. We decided not to crop out this region and rather explain where it comes from in the text.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Wang, Po-Kai et al., utilized the de novo polarization of MDCK cells cultured in Matrigel to assess the interdependence between polarity protein localization, centrosome positioning and apical membrane formation. They show that the inhibition of Plk4 with Centrinone does not prevent apical membrane formation, but does result in its delay, a phenotype the authors attribute to the loss of centrosomes due to the inhibition of centriole duplication. However, the targeted mutagenesis of specific centrosome proteins implicated in the positioning of centrosomes in other cell types (CEP164, ODF2, PCNT and CEP120), as well as the use of dominant negative constructs to inhibit centrosomal microtubule nucleation did not affect centrosome positioning in 3D cultured MDCK cells. A screen of proteins previously implicated in MDCK polarization revealed that the polarity protein Par-3 was upstream of centrosome positioning, similar to other cell types.

      Strengths:

      The investigation into the temporal requirement and interdependence of previously proposed regulators of cell polarization and lumen formation is valuable. The authors have provided a detailed analysis of many of these components at defined stages of polarity establishment, and well demonstrate that centrosomes are not necessary for apical polarity formation, but are involved in the efficient establishment of the apical membrane.

      Weaknesses:

      Key questions remain regarding the structure of the intracellular cytoskeleton following depletion of centrosomes, centrosome proteins, or abrogation of centrosome microtubule nucleation. The authors strengthen their model that centrosomes are positioned independently of microtubule nucleation using dominant negative Cdk5RAP2 and NEDD-1 constructs, however, the structure of the intracellular microtubule network remains unresolved and will be an important avenue for future investigation.

      We thank the reviewer for raising this important point. We agree that understanding the organization of the intracellular microtubule network following centrosome depletion, disruption of centrosomal proteins, or inhibition of centrosomal microtubule nucleation will be important for further mechanistic insight. However, a detailed analysis of cytoskeletal architecture under 3D culture conditions would require super-resolution or other advanced microscopy techniques and therefore falls beyond the scope of the current study. Nevertheless, several previous studies conducted under conventional 2D culture conditions provide relevant mechanistic context for interpreting our findings.

      (1) Centrosome depletion by centrinone treatment

      Previous studies have shown that upon centrosome loss induced by the Plk4 inhibitor centrinone, alternative microtubule-organizing centers (MTOCs), particularly the Golgi apparatus, can compensate by nucleating non-centrosomal microtubules (Chen et al., 2022; Gavilan et al., 2018; Martin, Veloso, Wu, Katrukha, & Akhmanova, 2018; Wu et al., 2016). Consistent with these findings, our microtubule regrowth assays in 2D MDCK cells (see Author response image 1) demonstrated that centriole depletion markedly altered microtubule organization. Control cells displayed the typical radial microtubule array emanating from a centralized centrosome, whereas centrinone-treated cells exhibited a dispersed microtubule network with enhanced microtubule growth from Golgi-associated sites.

      Author response image 1.

      Staining of control or centrinone (CN) treated MDCK cells for α-tubulin (magenta), Golgi GM130 green) and DNA (DAPI, blue) 1 min after nocodazole washout. Z-maximum projections of confocal images. The boxed cells in the overview images are magnified, and the microtubule regrowth regions are further enlarged. Scale bars: 50 μm (overview images), 10 μm (magnified cells), and 5 μm (enlarged microtubule regrowth regions).

      (2) Disruption of centriole or centrosomal proteins

      Previous studies have reported that depletion of the subdistal appendage protein ODF2 reduces centrosome–microtubule interactions and destabilizes centrosomal microtubules (Hung, Hehnly, & Doxsey, 2016; Ibi et al., 2011; Tateishi et al., 2013). In addition, in cells lacking the PCM protein pericentrin (PCNT), AKAP450 has been shown to partially compensate for centrosomal microtubule nucleation activity, although at a somewhat reduced level compared with wild-type cells (Figure 4—figure supplement 3A, B) (Gavilan et al., 2018).

      (3) Inhibition of centrosomal microtubule nucleation

      For dominant-negative Cdk5RAP2 and NEDD1 constructs, a previous study demonstrated that these constructs displace γ-tubulin from centrosomes and impair centrosomal microtubule nucleation (Vinopal et al., 2023). Consistent with this report, our MDCK cells expressing dominant-negative Cdk5RAP2 or NEDD1 also exhibited reduced γ-tubulin localization at centrioles (Figure 4—figure supplement 3C, D, and G).

      Together, these findings support the interpretation that centrosomal microtubules are not strictly required for polarized vesicle trafficking, centrosome migration, or epithelial polarization, but instead enhance the efficiency and robustness of these processes. We agree with the reviewer that future studies using advanced imaging approaches under 3D culture conditions will be important to resolve the spatial organization and dynamics of the intracellular cytoskeleton during epithelial polarization.

      Reviewer #3 (Public review):

      Here the Wang et al resubmit their manuscript describing the events in the establishment of polarity in MDCK cells cultured in vitro. As with the original version, the description is throughout and is important to the field to report as it establishes a hierarchy of events in polarization, placing Par3 upstream of centrosome positioning and apical membrane component trafficking. Unfortunately, in the revised version, the authors addressed almost none of my points. They did a cursory job of responding in the rebuttal letter but made little attempt to actually address what was being asked or to incorporate any of my suggestions into the manuscript. The particularly egregious examples are cited below:

      Comments on revisions:

      (1) My original main experimental concern was not addressed: I had originally asked what role microtubules play in the process of polarization (either centrosomal or non-centrosomal). An obvious model is that Gp135, Rab11, etc. are delivered to the AMIS on centrosomal microtubules. Centrosomes might also be pulled to the AMIS via cortically derived microtubules as is the case in the C. elegans intestine where the centrosome moves apically on apical microtubules via dynein directed transport to the cortically anchored minus ends. The authors do not explore the role of microtubules in the revision, citing that it was not possible to observe the microtubules directly or to perform nocodazole experiments during polarization.

      Instead, the authors use a relatively new genetic tool to disrupt centrosomal microtubules. They appear to succeed in displacing centrosomal g-tubulin using this tool, but without being able to observe microtubules, a remaining caveat of this experiment is that it is still unclear whether the authors have removed centrosomal microtubules. Compounding this issue is that this tool has never been used in MDCK cells. The authors conclude "we found that cells lacking centrosomal microtubules were still able to polarize and position the centrioles apically.", but they have not shown this, instead the data suggest this conclusion and the authors should acknowledge the caveat that they have no idea whether centrosomal microtubules are abolished.

      We appreciate the reviewer’s important comments regarding the role of microtubules during epithelial polarization. We agree that determining how centrosomal and/or non-centrosomal microtubules contribute to apical trafficking and centrosome positioning represents an important mechanistic question.

      We previously attempted to directly visualize microtubules during live imaging using SPY-tubulin labeling (see Author response image 2). However, under 3D Matrigel culture conditions, MDCK cells rapidly become rounded and densely packed, substantially reducing image contrast and making the majority of intracellular microtubule networks difficult to resolve, except for spindle microtubules and the cytokinetic bridge. In addition, nocodazole treatment caused mitotic arrest under our experimental conditions, thereby preventing de novo polarization from proceeding and precluding interpretation of polarity establishment.

      Author response image 2.

      Time-lapse maximum-intensity z-projections of MDCK cells expressing EGFP-PACT (yellow; centrosome marker) and H2B-mCherry (magenta; nuclei) embedded in Matrigel. Microtubules were labeled with SiR-tubulin (cyan) before live-cell imaging. Images show a representative dividing cell. Time stamps indicate hours and minutes relative to anaphase onset (0:00). Scale bar, 10 μm.

      To partially address the role of centrosomal microtubules, we used dominant-negative Cdk5RAP2 and NEDD1 constructs that have previously been shown to displace γ-tubulin from centrosomes and impair centrosomal microtubule nucleation (Vinopal et al., 2023). Consistent with this study, we observed substantial loss of γ-tubulin from centrioles in MDCK cells expressing these constructs (Figure 4— figure supplement 3C, D, and G). However, as the reviewer correctly points out, we were unable to directly visualize centrosomal microtubules under our 3D imaging conditions. Therefore, we cannot definitively conclude that centrosomal microtubules were completely abolished. We have revised the manuscript to clarify this limitation and to more cautiously state that our data suggest centrosomal microtubules may not be strictly required, for apical polarization and centrosome positioning under these conditions.

      Similarly, the authors also state: "Additionally, although PCNT knockout cells show reduced microtubule nucleation ability, they still recruit a small amount of γ-tubulin". Where are the data that show that microtubule nucleation is reduced in these PCNT knock out cells?

      We thank the reviewer for this comment and apologize for not sufficiently presenting these data in the previous revision. To directly assess microtubule nucleation activity in PCNT-KO cells, we performed microtubule regrowth assays in MDCK cells following nocodazole washout (Figure 4— figure supplement 3A, B). Compared with wild-type cells, PCNT-KO cells showed reduced centrosomal microtubule regrowth, indicating impaired microtubule nucleation capacity.

      Importantly, microtubule nucleation was not completely abolished in PCNT-KO cells, consistent with previous reports showing that AKAP450 can partially compensate for the loss of pericentrin and maintain residual centrosomal microtubule nucleation activity (Gavilan et al., 2018).

      (2) Many of my comments were addressed in the rebuttal, but not in the text.

      We sincerely thank the reviewer for the valuable suggestions. We have carefully considered all comments and incorporated many of the recommended revisions into the revised manuscript. However, we found that including every additional analysis, experiment, and discussion in the main manuscript would substantially reduce its coherence and readability. Therefore, while not all new analyses and experimental results are included in the revised manuscript, we have addressed every comment comprehensively in this response letter. Where necessary, we performed additional experiments and analyses to obtain the requested data, and the corresponding results and explanations are provided in our responses. We hope the reviewer will understand our effort to thoroughly address all comments while preserving the clarity and overall flow of the manuscript.

      The non-centrosomal GP135 in Figure 2 is not acknowledged or explained.

      We apologize for not sufficiently addressing the non-centrosomal Gp135 signal in Figure 2. We have now revised the manuscript to explicitly describe and discuss this point in the text (Page 5, Paragraph 4).

      That the polarity index does not actually measure polarity, but nuclear-centrosome distance is not acknowledged or explained in the paper.

      We have revised the manuscript to explicitly state that the “polarity index” represents the distance between the nucleus and the centrosome, which we use as a quantitative indicator of the degree of cell polarity (Page 5, Paragraph 1).

      I still don't believe that the quantification in Figure 3D matches the images I am being shown in Figure 3A. In the centrinone treatment condition, there is certainly an enrichment of GP135 at the AMIS that is not detected in the quantification. The method described in the rebuttal might miss this enrichment if it is offset from line drawn between the centroid of the two nuclei.

      We thank the reviewer for this comment. To better address this concern, we have now included a 3D view of the corresponding image data (see Author response image 3 and Author response image 4). This analysis clarifies that, in the centrinone-treated condition, the Gp135 signal is not localized at the geometric center of the cell doublet, but is instead offset from the axis used in our line-scan quantification. As a result, the enrichment visible in the projection image was not fully captured by the original quantification method. SiR-DNA Gp135

      Author response image 3.

      Author response image 4.

      3D reconstructions of p53-KO control and centrinone-treated MDCK cell doublets expressing EGFP-Gp135. Images are shown after 90° rotations about the x-axis (or y-axis) to visualize the spatial distribution of Gp135. Fluorescence intensity profiles of EGFP-Gp135 were measured along the line connecting the two nuclei. White arrows indicate the central fluorescence intensity value used to quantify Gp135 accumulation at the apical membrane initiation site (AMIS) (a.u., arbitrary units). Time stamps indicate hours and minutes.

      Cell height changes in the centrosome depleted cysts are still referenced in the text ("the cell heights of the centrosome-depleted cysts are less uniform"), but no specific data or image is called out. Currently, Figure 3G is referenced, but that is a graph of GP135 intensity

      We have revised the manuscript to indicate representative images, ensuring that the text and figures are consistent (Page 7, Paragraph 3).

      In my original review, I called on the authors to comment on the striking similarity of the mechanisms they documented in MDCK cells to what has been shown in in vivo systems. The authors did not do this, instead restating in the rebuttal some features of what they found. But, the mechanisms shown here are remarkably similar to the polarization of primordia that generate tubular organs in vivo. Perhaps most striking is the similarity to the C. elegans intestine where Par3 localizes to the cortex at the site of an apical MTOC that pulls the centrosome to the apical surface via dynein (Feldman and Priess, 2012). Instead of discussing this similarity, the authors state: "Par3 is likely to regulate centrosome positioning through some intermediate molecules or mechanisms, but its specific mechanism is still unclear and requires further investigation." Given the acetylated tubulin signal emanating from the Par3 positive patch in Figure 5E and F, I suspect similar mechanisms to the C. elegans intestine are at play here. Such a parallel should be noted in the Discussion.

      We thank the reviewer for this insightful suggestion. In the revised manuscript, we have expanded the Discussion section to compare our findings with epithelial polarization mechanisms described in in vivo systems, including the C. elegans intestine and other tubular epithelial tissues (Page 13, Paragraph 2).

      We agree that the hierarchical relationship we observe between Par3 localization, centrosome positioning, and apical membrane formation bears important conceptual similarities to mechanisms reported in the C. elegans intestine (Feldman & Priess, 2012). We also considered the possibility that Par3 may regulate centrosome positioning through dynein-dependent mechanisms.

      To examine this possibility, we performed immunofluorescence staining for the dynein cofactor dynactin subunit p150<sup>Glued</sup>. However, we did not observe enrichment of p150<sup>Glued</sup> at the center of cell doublets during the cytokinetic pre-abscission stage (see Author response image 5), suggesting that dynein is not strongly concentrated together with Par3 near the AMIS under our conditions.

      In addition, pharmacological inhibition of dynein resulted in cytokinesis failure and the formation of binucleated cells, preventing reliable assessment of centrosome migration and polarity establishment.

      We would also like to clarify that the acetylated tubulin signal observed in Figure 5E and F does not emanate from the Par3-positive patch. Rather, this signal corresponds to the cytokinetic bridge, adjacent to which Par3 accumulates during cytokinesis. Consistent with this interpretation, γ-tubulin was not detected at the Par3-positive region (Figure 1A).

      Author response image 5.

      Single MDCK cells after 12 h of culture in Matrigel. Immunostaining signals of the indicated markers are shown: centrosome marker PACT-mKO1, dynactin subunit p150Glued, Gp135, and DAPI. Single confocal sections through the middle of a cyst are shown. The order of polarization is arranged from single cell (1-cell), metaphase (Meta), telophase (Telo), cytokinetic pre-abscission (Pre-Abs), post-cytokinesis (Post-CK), to lumen open (LO). Scale bar: 5 μm.

      I had originally commented that "I find the results in Figure 6G puzzling. Why is ECM signaling required for Gp135 recruitment to the centrosome. Could the authors discuss what this means?" The authors responded that "The data in Figure 6G do not indicate that ECM signaling is required for the recruitment of Gp135 to the centrosome". In Figure 6G, the localization of GP135 to the centrosome appears significantly delayed compared to its localization to the centrosome in images where cells were cultured in Matrigel.

      Indeed, the authors argue that the centrosomal localization precedes and contributes to its localization to the AMIS. In the absence of ECM, GP135 localizes to the membrane before it localizes to the centrosome and its localization to the centrosome appears significantly reduced. Thus, my original and current interpretation is that ECM signaling is somehow required for the centrosomal targeting of GP135. One could make a competition argument, i.e. that the cortex in the absence of ECM is somehow a more desirable place to localize than the centrosome, but this experiment also argues that the centrosome does not need to be a source of this material in order for it to end up on the cortex.

      We agree that the absence of ECM substantially alters the trafficking behavior of Gp135.

      Our interpretation is that ECM primarily promotes the endocytosis and internal trafficking of Gp135, thereby enabling its redistribution to membrane domains lacking ECM contact and facilitating AMIS formation (Buckley & St Johnston, 2022; O'Brien et al., 2001; Yu et al., 2005).

      Under ECM-free conditions, a larger fraction of Gp135 remains associated with the plasma membrane, resulting in reduced internalized Gp135 available for centrosome-associated trafficking. During anaphase to telophase (Figure 6G, 0:05–0:20), Gp135 predominantly redistributes along the plasma membrane toward the cleavage furrow. Only after cytokinesis initiation (Figure 6G, 0:30) do we observe a small amount of internalized Gp135 associated with centrosomes near the center of the cell doublet.

      Importantly, we agree with the reviewer that these findings suggest centrosomal trafficking is not absolutely required for Gp135 to localize to the plasma membrane. Rather, our data support a model in which centrosome-associated trafficking contributes specifically to the efficient and spatially restricted delivery of Gp135 to the AMIS during epithelial polarization.

      We have revised the manuscript to clarify this interpretation and to avoid overstating the role of centrosome-associated Gp135 trafficking.

      (3) There needs to be precision in the language used in many places:

      I don't understand this line in the abstract: "When cultured in Matrigel, de novo polarization of a single epithelial cell is often coupled with mitosis." If a cell has divided, it is no longer a single cell.

      We have revised the sentence to: “When cultured in Matrigel, de novo polarization of a single epithelial cell is often coupled with cytokinesis (Page 1, Paragraph 1).” This indicates that polarization happens as the cell divides. We thank the reviewer for this helpful suggestion, which has improved the readability of the sentence.

      The authors state in the Introduction "Because of its strong ability to nucleate microtubules, the centrosome functions as the primary microtubule organizing center", but then state ""In polarized epithelial cells, the centrosome is localized at the apical region during interphase, which contributes to the construction of an asymmetric microtubule network conducive to polarized vesicle trafficking". In the latter statement, I assume the authors are describing the well-characterized apical microtubule network in epithelial cells that is non-centrosomal. Thus, the latter sentence is at odds with the former.

      We did not intend to refer to the apical non-centrosomal microtubule network present in mature, fully polarized epithelial cells. Rather, we were referring to the off-center centrosome functions as an off-center MTOC, creating an asymmetric microtubule network during the early stages of epithelial polarization, as mentioned in previous review papers (Meiring, Shneyer, & Akhmanova, 2020).

      The apical non-centrosomal microtubule network is a feature of mature, fully polarized epithelial cells. In fact, its formation is also driven by the release of microtubule minus-ends from the off-centre centrosome, which are then transferred to the apical membrane (Goldspink et al., 2017; Moss et al., 2007; Sanchez & Feldman, 2017).

      The authors continually refer to Par3 as a tight junction protein. "Par3, which controls tight junction assembly to partition the apical surface from the basolateral surface". To my knowledge, PARD3 is an apical protein with similar localization to C. elegans PAR-3 and Drosophila Bazooka. PARD3B is a junctional protein. I assume that the antibody that the authors are using is to PARD3 and not PARD3B? Can the authors please clarify this in the text?

      The antibody used for PARD3 staining was the Merck Millipore rabbit polyclonal antibody (Cat. No. 07-330), generated against a GST-tagged recombinant fragment corresponding to 288 amino acids from the internal region of mouse PAR-3.

      Canine PARD3 and PARD3B are encoded by distinct genes located on chromosomes 2 and 37, respectively. In our study, we used two independent shRNA constructs specifically targeting canine PARD3, both of which reduced the immunoblot signal detected by this antibody, supporting the conclusion that the antibody primarily recognizes PARD3 rather than PARD3B.

      However, in immunofluorescence staining, the signal was mainly localized at tight junctions, and we did not observe significant signal at the apical membrane. We will further clarify in the revised manuscript that this antibody targets PARD3 rather than PARD3B (Page 10, Paragraph 1).

      Reference

      Buckley, C. E., & St Johnston, D. (2022). Apical-basal polarity and the control of epithelial form and function. Nat Rev Mol Cell Biol, 23(8), 559–577. doi:10.1038/s41580-022-00465-y

      Chen, F., Wu, J., Iwanski, M. K., Jurriens, D., Sandron, A., Pasolli, M., . . . Akhmanova, A. (2022). Self-assembly of pericentriolar material in interphase cells lacking centrioles. Elife, 11. doi:10.7554/eLife.77892

      Feldman, J. L., & Priess, J. R. (2012). A role for the centrosome and PAR-3 in the hand-off of MTOC function during epithelial polarization. Curr Biol, 22(7), 575–582. doi:10.1016/j.cub.2012.02.044

      Gavilan, M. P., Gandolfo, P., Balestra, F. R., Arias, F., Bornens, M., & Rios, R. M. (2018). The dual role of the centrosome in organizing the microtubule network in interphase. EMBO Rep, 19(11). doi:10.15252/embr.201845942

      Goldspink, D. A., Rookyard, C., Tyrrell, B. J., Gadsby, J., Perkins, J., Lund, E. K., . . . Mogensen, M. M. (2017). Ninein is essential for apico-basal microtubule formation and CLIP-170 facilitates its redeployment to noncentrosomal microtubule organizing centres. Open Biol, 7(2). doi:10.1098/rsob.160274

      Hung, H. F., Hehnly, H., & Doxsey, S. (2016). The Mother Centriole Appendage Protein Cenexin Modulates Lumen Formation through Spindle Orientation. Curr Biol, 26(6), 793–801. doi:10.1016/j.cub.2016.01.025

      Ibi, M., Zou, P., Inoko, A., Shiromizu, T., Matsuyama, M., Hayashi, Y., . . . Inagaki, M. (2011). Trichoplein controls microtubule anchoring at the centrosome by binding to Odf2 and ninein. J Cell Sci, 124(Pt 6), 857–864. doi:10.1242/jcs.075705

      Martin, M., Veloso, A., Wu, J., Katrukha, E. A., & Akhmanova, A. (2018). Control of endothelial cell polarity and sprouting angiogenesis by non-centrosomal microtubules. Elife, 7. doi:10.7554/eLife.33864

      Meiring, J. C. M., Shneyer, B. I., & Akhmanova, A. (2020). Generation and regulation of microtubule network asymmetry to drive cell polarity. Curr Opin Cell Biol, 62, 86–95. doi:10.1016/j.ceb.2019.10.004

      Moss, D. K., Bellett, G., Carter, J. M., Liovic, M., Keynton, J., Prescott, A. R., . . . Mogensen, M. M. (2007). Ninein is released from the centrosome and moves bi-directionally along microtubules. J Cell Sci, 120(Pt 17), 3064–3074. doi:10.1242/jcs.010322

      O'Brien, L. E., Jou, T. S., Pollack, A. L., Zhang, Q., Hansen, S. H., Yurchenco, P., & Mostov, K. E. (2001). Rac1 orientates epithelial apical polarity through effects on basolateral laminin assembly. Nat Cell Biol, 3(9), 831–838. doi:10.1038/ncb0901-831

      Sanchez, A. D., & Feldman, J. L. (2017). Microtubule-organizing centers: from the centrosome to noncentrosomal sites. Curr Opin Cell Biol, 44, 93–101. doi:10.1016/j.ceb.2016.09.003

      Tateishi, K., Yamazaki, Y., Nishida, T., Watanabe, S., Kunimoto, K., Ishikawa, H., & Tsukita, S. (2013). Two appendages homologous between basal bodies and centrioles are formed using distinct Odf2 domains. J Cell Biol, 203(3), 417–425. doi:10.1083/jcb.201303071

      Vinopal, S., Dupraz, S., Alfadil, E., Pietralla, T., Bendre, S., Stiess, M., . . . Bradke, F. (2023). Centrosomal microtubule nucleation regulates radial migration of projection neurons independently of polarization in the developing brain. Neuron, 111(8), 1241–1263 e1216. doi:10.1016/j.neuron.2023.01.020

      Wu, J., de Heus, C., Liu, Q., Bouchet, B. P., Noordstra, I., Jiang, K., . . . Akhmanova, A. (2016). Molecular Pathway of Microtubule Organization at the Golgi Apparatus. Dev Cell, 39(1), 44–60. doi:10.1016/j.devcel.2016.08.009

      Yu, W., Datta, A., Leroy, P., O'Brien, L. E., Mak, G., Jou, T. S., . . . Zegers, M. M. (2005). Beta1-integrin orients epithelial polarity via Rac1 and laminin. Mol Biol Cell, 16(2), 433–445. doi:10.1091/mbc.e04-05-0435

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Li and colleagues examines how defensive responses to visual threats during foraging are modulated by both reward level and social hierarchy. Using a semi-naturalistic paradigm, the authors test how the availability of water or sucrose, with sucrose being more rewarding than water, shapes escape behavior in mice exposed to looming stimuli of different intensities, which are used to probe perceived threat level and defensive responses. In parallel, the study compares dominant and subordinate animals to assess how social rank biases the trade-off between reward seeking and threat avoidance. By combining behavioral analyses with computational modeling, the work addresses how reward level and social context jointly influence escape decisions in an ethological setting.

      Across the different experimental conditions, perceived threat level is the main determinant of behavior. The authors show that looming stimuli associated with higher threat (contrast) consistently elicit faster and more robust escape responses than lower threat stimuli. This effect is particularly evident during early exposures, when animals are highly vigilant and have not yet habituated to the looming stimulus (learned that it is not dangerous). Later they described that as animals gain experience and habituate, behavior becomes more flexible, and reward level begins to exert a graded modulation of the escape response. Importantly, the authors show that under high threat conditions increasing reward value leads to more frequent and faster escape rather than greater reward pursuit, specifically in dominant mice. This finding is particularly relevant, as it suggests that highly valued rewards can heighten vigilance and thereby enhance responsiveness to threat, highlighting that reward does not simply compete with defensive behavior but can also reshape it depending on the perceived level of danger, in contrast to low threat conditions, where threat can be more easily outweighed by reward. However, it is worth noting that the authors use an extremely low contrast for the low threat condition (20%), which may to some extent be insufficient to reliably trigger escape responses. Thus, an important conceptual contribution of the study is the introduction of vigilance as a useful framework to interpret these effects. Vigilance is treated as a behavioral state reflecting heightened attention to potential danger. In line with what is known from natural foraging, mice initially maintain high vigilance when confronted with an innate threat. This perspective helps clarify a finding that might otherwise appear counterintuitive. One might expect higher rewards to motivate animals to tolerate risk, explore more, and habituate faster in any scenario. Instead, the data suggest that highly rewarding outcomes can elevate vigilance, making animals more responsive to threat and leading to faster or more frequent escape under high threat conditions. In this sense, reward does not simply compete with threat but can also amplify sensitivity to it, depending on the internal state of the animal.

      We agree that the low-contrast condition (20%) represents a relatively weak threat signal by design. This level was intentionally chosen to fall within a regime that biases behavior toward freezing and escape after assessment rather than direct escape, allowing us to examine graded decision-making across threat intensities. The higher escape probability observed under low-contrast conditions relative to previous studies (Evans et al., 2018; Fratzl et al., 2021) is likely attributable to the linear arena used here, which generally promoted escape behavior over freezing.

      The social results are particularly interesting in this context as well. Dominant mice consistently prioritize avoidance over reward, showing stronger escape responses and slower habituation than subordinates. This behavior is well captured by the vigilance framework proposed by the authors: dominant animals appear to maintain higher vigilance, which biases decisions toward threat avoidance. The authors further suggest that stable social relationships sustain high vigilance and slow habituation, framing this as an evolutionarily conserved strategy that may enhance survival. This interpretation provides a valuable perspective on how social structure shapes defensive behavior beyond immediate physical interactions. At the same time, there are important limitations to this interpretation. All experiments were conducted in male mice, and it is possible that the relationship between social hierarchy, vigilance, and defensive behavior would differ substantially in females. In addition, the idea that stable social relationships sustain elevated vigilance should be interpreted carefully, as it does not fully align with broader views of social stability as protective against anxiety and stress and generally beneficial for mental health and resilience. These points do not undermine the findings but suggest that the social effects described here should be interpreted with caution and within the specific context of the task and sex studied.

      We thank the reviewer for this important comment and agree that findings obtained in male mice may not necessarily generalize to female mice. We have acknowledged the limitation in the Discussion and note that future studies will be required to determine the extent to which the present findings extend to female mice.

      We would also like to clarify our interpretation regarding vigilance and social buffering. Vigilance should not be conflated with stress or anxiety, and the slower habituation observed in group-housed mice reflects sustained sensitivity to repeated threat exposure rather than elevated anxiety. Thus, our data do not contradict the concept of social buffering; rather, they are consistent with it, as pair-housed mice exhibited reduced defensive responses compared with individually housed animals (Lenz et al., 2022). Furthermore, in an independent study (Li, Gao, and Li, 2026; eLife 15:RP109571), we directly compared responses to looming stimuli when mice were tested alone versus in the presence of a social partner and found clear evidence of social buffering. These findings suggest that social interactions can attenuate defensive responses while prolonging vigilance during repeated threat exposure. We have revised the Discussion accordingly.

      Another important limitation is that the neural mechanisms underlying these effects remain highly speculative. Although the manuscript includes an extensive discussion of candidate circuits, particularly involving the superior colliculus and downstream structures, these interpretations go far beyond the data presented in the study and are not directly supported by experimental evidence within the paper itself. The discussion gives substantial weight to potential circuit mechanisms based primarily on previous literature rather than on findings from the current study. Given the complexity and distributed nature of the circuits likely involved in integrating vigilance, reward, social context, and defensive behavior, the present work is better viewed as providing a strong behavioral framework rather than direct mechanistic insight into the underlying neural substrates. In this context, some references discussing how animals learn to suppress defensive responses to repeated looming threats and the neural mechanisms supporting this process could further strengthen the discussion (Salay et al 2021; Fratzl et al. 2021; Conway et al. 2025; Mederos et al. 2025).

      We agree that the proposed neural mechanisms remain speculative and that the circuits involved in integrating internal state, reward, and social context are likely far more complex. We have revised the manuscript to acknowledge this limitation and have included a discussion of learning-dependent suppression of defensive responses to repeated looming threats and the underlying circuit mechanisms.

      Methodologically, the behavioral paradigm is well suited for studying escape decisions in socially housed animals, and the machine learning based classification of defensive responses is a strength. The computational model provides a useful formalization of how threat level, reward level, and vigilance interact and may be valuable for other laboratories studying escape, approach avoidance, or conflict situations, particularly as a way to classify behavioral outcomes after pose estimation. More generally, the work will be of interest to the neuroethology community for its detailed characterization of escape behavior under naturalistic conditions. At the same time, some statements in the discussion slightly overstate the novelty of the methodological approach. For example, the claim that the study differs from earlier work by using machine learning rather than manual annotation overlooks that several previous studies have already implemented automated or semi-automated strategies to classify looming evoked defensive behaviors beyond manual scoring alone.

      We have revised the manuscript to moderate our claims regarding the novelty of the machine-learning–based classification approach.

      Given the ethological nature of the study and the high inter individual variability reported by the authors, clarity and precision in the methods are especially important for reproducibility. While the revised manuscript addresses many earlier concerns, some aspects remain slightly difficult to follow. For example, the main text states that animals were not water deprived to minimize differences in internal state across conditions, whereas parts of the methods describe experiments in which animals were water deprived. This distinction is not always clearly explained across the different experimental sections, despite internal state being central to the interpretation of the behavioral findings. A clearer separation and description of these conditions would further strengthen confidence in the work. In addition, it was somewhat surprising that the low contrast (20%) looming condition was still sufficient to trigger robust escape responses, and additional clarification or discussion regarding stimulus saliency at this contrast level could help readers better contextualize these findings.

      To improve clarity, we have revised the Methods section to clearly distinguish between experimental conditions that involved water deprivation and those that did not.

      Overall, this study provides a rich analysis of how reward level and social hierarchy modulate defensive behavior through changes in vigilance. It offers a useful conceptual advance for thinking about escape behavior in semi-naturalistic settings and lays a solid foundation for future work aimed at linking these behavioral states to underlying neural circuits.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      - Clarify which experiments involved water deprivation and which did not in methods section.

      We have revised the Methods section to clearly specify which experiments involved water deprivation and which were conducted without water deprivation.

      - Add discussion on why the 20% low contrast stimulus is still sufficient to trigger escape responses.

      We have expanded the Discussion regarding the 20% contrast stimulus to clarify why it was sufficient to elicit escape responses.

      - Tone down statements regarding the novelty of the machine learning based behavioral classification.

      We have revised the manuscript to moderate claims regarding the novelty of the machine-learning–based classification approach.

      - Include references related to learning-dependent suppression of looming responses and circuit mechanisms

      We have incorporated the suggested references and expanded the Discussion of learning-dependent suppression of looming responses and the underlying circuit mechanisms.

      - Moderate the discussion of candidate neural circuits, particularly the SC related interpretations, as these are not directly tested in the study.

      We have revised the discussion of candidate circuit mechanisms.

      - Clarify that the interpretation linking stable social relationships to elevated vigilance may be specific to this ethological context.

      We have revised the Discussion to clarify this interpretation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Garcia-Alcala, Kratz and Cluzel investigate to what extent our understanding of bacterial physiology in bulk experiments can be applied to single-cell observations. They find that intrinsic noise may be powerful enough to even inverse the trends found in the bulk. The authors hypothesize that the asymmetric distribution of ribosomes to daughter cells during cell division plays the dominant role in the intrinsic noise and is able to generate the observed phenomenon. They do not show it directly, but the data and its agreement with the model are sufficient to support this claim.

      Strengths:

      The experimental part is convincing: the positive correlation between the elongation rate and promoter activity of unnecessary protein is clear, as well as the negative correlation between the mean values while changing the promoter strength. This was demonstrated in both rich and poor media. The causality between the growth rate and the promoter activity was shown using the negative lag time of the cross-correlation function. A simple, reasonable model accounts well for the data. This paper demonstrates an interesting phenomenon and provides a plausible theory for it, advancing our understanding of bacterial physiology on the single-cell level.

      Weaknesses:

      (1) Mean-reversion timescales were assumed to be longer than the simulation time and much longer than the cell cycle time. It is not clear whether the results are robust in case mean reversion timescales become of the order of the cell-cycle or smaller. If not, is there an argument for such practically infinite reversion timescales?

      Due to an error in the simulation code, the incorrect mean-reversion timescales were reported in the manuscript and have now been corrected. Instead of 1000 h and 100 h for 𝜏<sub>R</sub> and 𝜏<sub>U</sub>, respectively, they are 6.25 h and 0.0625 h, which are both much shorter than the total simulation time (75 h). The error had no effect on simulation results and was purely a timescale conversion mistake. We appreciate the referee’s careful review of the manuscript which allowed us to catch this error.

      Given the correct values, 𝜏<sub>U</sub> is well within the cell-cycle time, as is expected since we assume the source of noise in unnecessary protein expression is from stochastic gene expression. In contrast, 𝜏<sub>U</sub> is still significantly longer than the cell-cycle time and needs to be in order to generate the experimentally observed behavior (i.e., the increase in growth rate with an increase in unnecessary protein expression at the single-cell level). This is consistent with our biological hypothesis that some cells inherit ribosomal surpluses from their mothers which enable bursts in protein production. 𝜏<sub>U</sub> less than the cell-cycle time would correspond to a case where ribosomal composition quickly decays to the population average, meaning that daughter cells would never have time to capitalize on the benefit of receiving a ribosome surplus. Furthermore, 𝜏<sub>U</sub> being longer than the cell-cycle is biologically justifiable as proteins such as ribosomes are passed from mother to daughter at division, thus allowing for memory to persist over longer timescales than a single generation.

      (2) It is not easy to understand the simulation part unless one reads Ref [14]. k(t) is assumed Equation (1) from Reference [14]? Is it crucial that the ribosome noise appears only at the division? The ribosome noise strength σ<sub>R</sub> =0.06 - is it lower or higher than the naively expected binomial division? Also, a more intuitive explanation of the Simpson paradox would help the reader.

      𝜅(𝑡) is indeed Eq. (1) from Ref. [14]. To make the computational results clearer, the methods section has been updated to include the full set of equations used to perform the simulations, and the code used to produce the computational figures is now on GitHub. The relative standard deviation expected by modeling ribosome distribution at division by a binomial distribution with equal probability of being inherited by either daughter cell is in the range of 1-3% (0.01-0.03), assuming N~10<sup>3</sup>-10<sup>4</sup> ribosomes. Thus, 0.06 is reasonable as it is in the same order of magnitude as what is predicted by binomial division. Furthermore, if clustering is present as already demonstrated in [39, 40, 43], we would expect the relative standard deviation to increase as clustering reduces the effective number of proteins which are distributed between daughters.

      (3) It would be useful for the reader to see the raw data and not only the filtered one to appreciate the measurement noise level.

      We have included the direct calculations of cell size and promoter activity in the time-lapse plots in Fig. 1. In addition, we included Fig. S2, which shows typical time traces of Class-2 activity and elongation rate, displaying both the direct measurements and the corresponding smoothed traces for the flagellar reporter strains.

      (4) Negative lag time of the cross-correlation function is visible, but consider adding a statistical test for it.

      We have included a statistical analysis of the lags of maximum correlation of elongation rate and activity for the flagellar reporter strains in Fig. S10. For each strain, we now show an overlay of the cross-correlation functions for all lineages and the distribution of the lag corresponding to the maximum correlation. In addition, we have added a section in Materials and Methods describing in detail how the cross-correlation and the strain-averaged lag were computed.

      (5) Can you make similar cross-correlation plots using the model? Can you infer by using it, whether the data agrees better with the assumption that ribosomal noise appears only at division or continuous fluctuations during the cell cycle?

      The model in its current form is unable to capture the observed cross-correlation (the correlation is sharply peaked at zero). This is because the model coarse-grains transcription and translation into one process of protein production and thus lacks any delay or memory mechanisms which would make a significant positive or negative correlation.

      Both continuous fluctuations and noise from division could in principle contribute to our observation of Simpson’s paradox. We are unable to use the model in its present form to dissect the contributions from both mechanisms. However, our model simulations clearly show that noise at division by itself is sufficient to explain the observed effect with parameter values that are biologically plausible. As Chao et al., [40] has previously shown experimentally that a significant source of ribosomal noise comes from unequal distribution at division, which is the hypothesis we retained for the model.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Garcia-Alcala et al. reports an interesting paradox: the cost of gene expression slows the population-average growth rate, whereas at the single-cell level, expression levels from these genes positively correlate with the growth rate. The effect is observed in the expression of flagellar genes and a gene under a synthetic promoter in E. coli. The findings are explained by the inheritance of growth factors, including ribosomes, during asymmetric division.

      Strengths:

      (1) The manuscript adds strength to an emerging body of literature showing that the population-level bacterial growth laws do not match correlations based on single-cell data. The evidence presented here is more striking than in previous works (such as Pavlou et al., Nat. Commun. 2025), as the trends in population-level data and single-cell data are reversed.

      (2) A relatively simple model correctly explains the trends in the data.

      Weaknesses:

      (1) It is not clear whether flagellar proteins are expressed proportionally to the reporter signal. Furthermore, it is questionable if E. coli bacteria in the mother machine channels are flagellated. If they are, they could potentially swim out of the channels, which is not the case when they do not carry the MotA E98K mutation. The authors should provide some evidence that E. coli expresses the actual filament proteins in the channels.

      We agree that it is important to demonstrate that our reporter reflects the production of functional flagellar structures under our experimental conditions. To this end, we first tested the swimming capabilities of our strains before introducing the MotA E98K mutation, using a standard soft‑agar motility assay. For strains in which only the Class‑1 promoter was modified (Pro2, Pro4, Pro5), after 12 h of inoculation, the diameters of the swarming rings followed the expected order based on promoter strength: Pro2 showed the smallest ring, followed by Pro4, WT, and Pro5. In contrast, control strains carrying Pro4 together with either MotA E98K (non‑rotating motors) or ΔfliC (no flagellin filament) did not form rings, consistent with their inability to swim. We also included an MG1655 strain carrying an IS5 insertion upstream of the Class‑1 promoter, which is known to enhance flagellar expression [56]; this strain showed a larger ring, as expected. A qualitative summary of ring diameters for all strains is provided in Author response table 1. We have included the figure and section “Experimental validation of functional flagella expression in reporter strains” on Supplementary Material.

      Author response table 1.

      We also tested whether cells assemble functional flagella on the mother‑machine. We compared MG1655 WT and MG1655 carrying the MotA E98K mutation under identical microfluidic growth conditions. Many WT cells left the channels during the experiment (∼35% of lineages over ~20 h after the onset of exponential growth inside the device), consistent with active swimming, whereas the non‑motile MotA E98K strain did not leave the channels. Because the only difference between these two strains is the MotA E98K mutation, which disables motor rotation but not flagellar assembly, this result indicates that (i) cells do express functional filaments and motors in the mother‑machine environment, and (ii) the strain used for our main experiments is immobilized by MotA E98K.

      Both experiments are included in the Supplementary Information section “Experimental validation of functional flagella expression in reporter strains” and Fig. S3.

      (2) It is unclear what fraction of the total proteome mVenus represents in different measurements. Some quantification is needed (for example, using the Coomassie staining). Using f_U as high as 14.4% in simulations is questionable.

      We agree that we do not currently know the exact fraction of the proteome occupied by mVenus in our experiments, as we only quantified fluorescence and did not perform Coomassie staining or proteomics. The primary goal of our simulations is not to reproduce exact experimental conditions, but to illustrate how unequal ribosome partitioning affects daughter cells across a range of protein synthesis burdens. Consequently, our use of values up to F<sub>U</sub> = 14.4% in the simulations was intended as an exploratory upper range rather than as a direct estimate of the experimental condition.

      To put our experimental burden in context, we compare our growth-rate reduction to a well– characterized high-burden case in the literature. In the study by T. Hwa’s group [6], overexpression of β–galactosidase such that it constituted approximately 27% of the total proteome led to a 67% reduction in growth rate for E. coli growing in a medium similar to ours (differing only in the carbon source: glucose in their case, glycerol in ours). In our system, overexpression of mVenus from a plasmid result in a substantially smaller growth–rate reduction of 9% relative to the non–expressing control, and this is observed on the poorer carbon source (glycerol), under which burden effects are typically less pronounced [6].

      Although we cannot convert our fluorescence measurements into an exact proteome fraction, the much smaller growth defect compared to the 27% β‑galactosidase case strongly suggests that mVenus does not approach such an extreme fraction of the proteome under our conditions. Under the simplifying assumption that the qualitative relationship between unnecessary‑protein fraction and growth‑rate reduction is similar in the two systems, our data are therefore consistent with a modest fraction of unnecessary protein and make it unlikely that mVenus reaches very high fractions such as 27%. This gives us confidence that exploring F<sub>U</sub> values up to 14.4% in the simulations represents a conservative upper range relative to our experimental burden, rather than an underestimate.

      (3) The data from the MC4100 strain does not directly match the trends of MG1655. The justification for filtering out the low-frequency components of MC4100 is not particularly convincing. It appears unlikely that ribosomes or other growth factors partition significantly differently in the MC4100 strain than in the MG1655 strain. Further discussion and a plot similar to Figure 1 (Left) for this strain are warranted.

      We thank the reviewer for this helpful suggestion. We agree that the filtering analysis from the initial manuscript was not clear enough to be used as robust supplementary information, and we therefore removed it. Instead, we carried out with the ‘unfiltered’ MC4100 strain, the same single-cell analyses as with MG1655, including the binning analysis of instantaneous elongation rate versus Class-2 promoter activity, and a cross-correlation analysis between such measurements.

      Both analyses did not reveal a significant correlation between growth and flagellar promoter activity in MC4100. We now present these results in the revised manuscript (Fig. S15), where we explicitly show the lack of association in a plot directly comparable to Fig. 1.

      While ribosomes partition mechanism is likely to be the same between these two strains, MC4100 is known to exhibit very long oscillations of growth rate over 10 generations, which are absent in MG1655.

      We now present MC4100 as an explicit counterexample to highlight that the positive correlation between short timescale growth fluctuations and flagellar expression observed in MG1655 is not universal across all E. coli strains, especially in strains like MC4100 whose growth rate fluctuations are dominated by long timescales, much longer than the division time. In MC4100 these slow modes are largely decoupled from flagellar gene regulation. We have also revised the text to clarify this point.

      (4) The model needs to be described in more detail. A closed set of equations that have been simulated must be presented, along with all values of the model parameters and their sources. The authors should consider depositing their code on GitHub or another publicly accessible repository.

      The methods section has now been updated to include the full set of equations used to perform the simulations along with all parameter values and their sources. Additionally, the code used to produce the computational figures is now on GitHub. For a full derivation and biological justification of each model component, we still refer readers to ref [14] where this model was first published.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The first paragraph of Results still belongs to the Introduction Section, I feel.

      Thank you for the recommendation, we agreed with the change.

      (2) Figure 1, left, is after some filter (Savitzky-Golay)? It might be useful to see the raw points.

      We have included the direct calculations of cell size and promoter activity in the time-lapse plots in Fig. 1. In addition, we included Fig. S2, which shows typical time traces of Class-2 activity and elongation rate, displaying both the raw (direct) measurements and the corresponding smoothed traces for the flagellar reporter strains.

      Reviewer #2 (Recommendations for the authors):

      (1) Is there a statistically significant positive slope in single-cell data of the elongation rate as a function of Class-2 activity? There should be an analysis of statistical significance for the slopes in Figure 1 (Right) and Figure 4A, including both population-average and single-cell data.

      Thank you for your recommendation. We have now added statistical analyses of the correlations between activity and elongation rate at both the single-cell and population levels. The results for the flagellar strains are presented in Fig. S5, S6, and S7. Also, the corresponding analysis for constitutive Venus expression is shown in Fig. S16.

      (2) How does FlhC-YFP data compare with the CFP signal in the first set of measurements? What would Figure 1, Left look like for this signal?

      At the single‑cell level, the relationship between Class‑1 promoter activity and elongation rate within each strain is still positive, but clearly weaker than for Class‑2, as shown in Fig. S8. When we bin the data, we can observe that the binned averages don’t display a clear positive trend, even though the Pearson correlation for each flagellar reporter strain is still positive. We think this is because Class‑1 controls a much smaller part of the proteome (it only encodes the two subunits of FlhDC) while Class‑2 promoters drive many structural and export proteins. So, changes in Class‑2 activity more directly reflect shifts in global translational capacity and are more tightly linked to growth, whereas Class‑1 activity adds only a small translational load.

      (3) The information from the ER-Activity cross-correlation functions is interesting but has not been interpreted or compared with the model. Which signal precedes the other? What can explain the observed lag time on the order of Tdiv? Why do some cross-correlations show negative values in Fig. S9 while others are positive (as they should)? Can the model explain the experimentally observed cross-correlation function?

      In our cross‑correlation analysis, a negative lag at the maximum means that fluctuations in elongation rate (ER) precede fluctuations in promoter activity (A) by that lag. We now describe this explicitly in Materials and Methods and quantify lag distributions for all reporter strains in Fig. S10.

      In Fig. 13 (previously Fig. 9), we mainly observe positive correlations between ER and PA with negative or near zero peak lags, i.e., ER tends to lead PA by a lag by about a division time. The reason governing the negative lags is not immediately clear, but it is consistent with a resource driven mechanism: cells that inherit more growth factors (e.g. ribosomes) at division, use them to prioritize housekeeping processes and thus increase growth rate first, and only subsequently use these extra resources to increase flagellar promoter activity.

      As for Class–1 activity, the correlation with growth is much weaker as it was already observed in Kim et al. Averaging the cross–correlations across all lineages yields a modest peak (mean correlation 0.10, SD 0.12; see Author response image 1) at a small positive lag of 0.5 h (while average division time is T<sub>div</sub> ~ 1.7h). This indicates a weak but real positive correlation between Class 1 activity and ER at short lags. However, the lags of the individual maxima are widely distributed (SD ≈ 9 h; mean −1.8 h, median −0.28 h, mode ≈ 0), with only ~53% of the cells showing a negative time lag. Thus, delays are roughly symmetrically spread around zero with only a slight negative bias.

      Author response image 1.

      Left: cross−correlation between elongation rate and Class−1 activity for each lineage (N = 96, colored lines), and their average as a function of lag (black line). Right: distribution of the lags at which each lineage’s cross−correlation attains its maximum. The dashed line indicates zero lag, and the full line marks the lag of the peak of the mean cross correlation.

      An in-depth analysis about the sign and magnitude of the lag would require measurements from a broader set of promoters. The lag is likely to depend on what genes the promoter controls (e.g. stress response, housekeeping, or large structural modules), on its strength and regulation, and on growth conditions.

      The model in its current form is unable to capture the observed cross-correlation (the correlation is sharply peaked at zero). This is because the model coarse-grains transcription and translation into one single process of protein production and thus lacks any delay or memory mechanism that could generate a phase shift.

      Nonetheless, this memory-free formulation shows that stochastic redistribution of growth factors at division is by itself sufficient to generate the Simpson’s paradox behavior. Capturing the experimentally observed lag time would require including explicit transcription/translation delays and additional regulatory dynamics between housekeeping and flagellar genes whose information we do not have.

      (4) Figure 3, Left - it is not clear what this plot shows. Red and blue are scattered over the whole plot. How have daughter 1 and daughter 2 been assigned? Perhaps choosing one of the daughters with a higher growth rate and then plotting the data could reveal some trends.

      We agree that the original left panel of Fig. 3 was difficult to follow. In the original version, the daughter labels were assigned as “Daughter–1” for the cell at the closed end of the channel and “Daughter–2” for the cell closest to the open side. After discussing this with Dr Camilla Ulla Rang (whom we now acknowledge in the manuscript), we relabeled the daughters as “new–pole daughter” and “old–pole daughter,” following the convention used in studies of aging and ribosome distribution in E. coli.

      We now show the class–2 promoter activity comparison between new pole vs old pole, and in a separate panel, we plot the mean ratios of elongation rate, and class–1/class–2 promoter activities between the new–pole and old–pole daughters. These ratios show that the new–pole daughter tends to have both greater growth rate and flagellar gene activity than the old–pole daughter. Importantly, this new plot is in line with Rang and colleagues who demonstrated that new–pole daughters have greater ribosome density and faster growth rates than their old–pole sisters. Together, these results further support our hypothesis that excesses of ribosomes inherited at division underlies the observed growth boosts.

      (5) Flagellar activity -> activity of flagellar gene synthesis (presumably no flagellar activity in these cells).

      We corrected the terms used to reference the flagellar gene activity.

      (6) Page 7: "The strain MC4100, known to exhibit slow, long period oscillations in growth ..." - some reference is- needed here.

      We placed the reference some lines after, as such paper also includes the information of MG1655 short-term oscillations. Now the reference is [44]: Tanouchi, Y., et al., A noisy linear map underlies oscillations in cell size and gene expression in bacteria. Nature, 2015

      (7) Page 7: Figure S11B - is Figure S11C perhaps meant?

      The correct panel indeed was Fig. S11C, and we have now corrected and updated the figure label and text accordingly.

      (8) Page 11: Savitsky-Golay filter.

      We have corrected it, thank you.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This study utilizes polarized second-harmonic generation (pSHG) microscopy to investigate myosin conformation in the relaxed state, distinguishing between the disordered, actin-accessible ON state and the ordered, energy-conserving OFF state. By pharmacologically modulating the ON/OFF equilibrium with a myosin activator (2deoxyATP) and inhibitor (Mavacamten), the authors demonstrate that pSHG can sensitively quantify the ON/OFF ratio in both skeletal and cardiac muscle. Validation with X-ray diffraction supports the accuracy of the method. Applying this approach to a hypertrophic cardiomyopathy model, the study shows that R403Q/MYH7-mutated minipigs exhibit an increased ON state fraction relative to controls. This difference is eliminated under saturating concentrations of myosin modulators, indicating that the ON/OFF balance can be pharmacologically shifted to its extremes. Additionally, ATPase assays reveal elevated resting ATPase activity in R403Q samples, which persists even when the ON state is saturated, suggesting that increased energy consumption in this mutation is driven by both a shift toward the ON state and inherently higher myosin ATPase activity.

      Strengths:

      This is a well-written and well-conducted study that clearly reveals the power of SHG microscopy. The study clearly establishes the great utility of SHG to study thick filament regulation.

      Weaknesses:

      (1) Several studies have shown that the ON state of the thick filament is sensitive to both temperature and filament lattice spacing, with a common recommendation to conduct skinned fiber experiments at temperatures above 27 °C and in the presence of dextran to better preserve physiological conditions. The authors should clarify the experimental temperature used in their skinned fiber studies, indicate whether dextran was included, and discuss whether adherence to these recommended conditions would have impacted their results.

      Additional experiments were performed on skinned psoas muscle strips to assess the influence of temperature and lattice spacing on the measured γ values. In particular, psoas strips were tested under relaxing conditions at three different temperatures (10 °C, 15 °C, and 20 °C). In a separate set of experiments, psoas strips were examined under relaxing conditions in the absence and in the presence of 5% dextran.

      These measurements confirmed previous reports indicating that the ON state of the thick filament is sensitive to both temperature and lattice compression. Specifically, increasing the temperature from 10 °C to 20 °C resulted in a shift toward lower γ values, consistent with a greater radial displacement of myosin heads. A similar trend, although less pronounced, was observed in the presence of dextran, which also produced a decrease in γ compared with control conditions.

      Figure 1 and text have been modified including these new findings (revised manuscript, pages 11-12).

      (2) On page 13, the authors report the proportion of disordered heads as approximately 30% in wild-type and 65% in R403Q fibers. They should clarify whether these values represent the percentage of total myosin heads, or rather the percentage of heads that are responsive to Mavacamten and dATP.

      The proportions reported in the original version referred to the fraction of myosin heads assumed to be responsive to Mavacamten and dATP. These estimates were obtained using a simplified model in which dATP was assumed to produce a complete (100%) shift of myosin heads toward the ON state, whereas Mavacamten was assumed to cause a complete depletion (0%).

      Since these assumptions are not fully supported by structural data, we decided to remove these values from the Results section and instead discuss these considerations in more general terms in the Discussion (page 19). Figure 3 and text have been modified accordingly (pages 14-15).

      (3) In Figure 5, regarding ATPase measurements, the content of contractile material per unit volume of muscle preparation will influence the results. Did the authors account for this variable, and if not, how might it have affected the conclusions?

      We thank the reviewers for raising this concern and agree that the content of contractile material per unit volume can influence the absolute values obtained in ATPase measurements. In our experiments, however, we primarily assessed the effects of compounds using paired measurements (e.g., the same strip measured before and after Mavacamten treatment).

      In addition, based on our previous structural data obtained from similar preparations of human myocardium (https://doi.org/10.1161/CIRCRESAHA.122.321956), the ratio of contractile tissue volume to total muscle volume is highly preserved. This supports comparisons between muscle strips from different experimental groups (e.g., WT and R403Q).

      (4) For readers primarily interested in assessing the ON/OFF state of thick filaments, could the authors list the specific advantages of polarized second harmonic generation (pSHG) microscopy compared to X-ray diffraction?

      In the present work, we deliberately chose to frame pSHG microscopy and X-ray diffraction as complementary approaches, rather than emphasizing a direct one-to-one comparison of their respective advantages and disadvantages.

      That said, we agree that for readers specifically interested in assessing the ON/OFF state of thick filaments, a clearer delineation of the respective strengths and limitations is useful. We have therefore slightly revised the Discussion (page 21) to better clarify the advantages and constraints of both approaches.

      (5) Given that many data points were derived from the same fiber or myocyte, how did the authors address the risk of type I errors due to non-independence of measurements? Was a nested or hierarchical statistical approach used?

      We thank the reviewer for raising this important point and agree that the original statistical analysis did not fully account for the non-independence of measurements derived from the same fibre.

      To address this issue, we have now reanalyzed the entire dataset with the support of Prof. Francesco Sera, who has been included as a co-author. A hierarchical (mixed-effects) statistical model was applied, explicitly accounting for the nested structure of the data, repeated measurements, and unbalanced group sizes.

      Importantly, this revised analysis substantially confirms the original results, with no major changes in the observed trends or in the statistical significance of the findings. All figures have been modified accordingly, and the statistical methods used are now described in the Methods section (page 8).

      Reviewer #2 (Public review):

      Summary:

      In striated muscle, myosin motors can dynamically switch between an energyconserving OFF state and an activated ON state. This switching is important for meeting the body's needs under different physiological conditions, and previous studies have shown that disease-causing mutations associated with cardiomyopathies can affect the population of these states, leading to aberrant contractility. Studying these structural states in muscle has previously only been possible via X-ray diffraction, which requires access to a beam line. Here, Arecchi et al. demonstrate that polarized second-harmonic generation microscopy (pSGH), a technique that is more accessible, can be used to probe the ON/OFF states of myosin in both permeabilized and intact muscle.

      Strengths:

      (1) There is an outstanding need in the field to better understand the regulation of the ON/OFF states of myosin. Currently, this is studied using X-ray diffraction, meaning that it is accessible to only a few labs. The authors demonstrate that pSGH can be used to probe the ON/OFF states of myosin both in intact and permeabilized muscle. This is a significant advance, since it makes it possible to study these states in a standard research laboratory.

      (2) The authors demonstrate that this approach can be employed in both skeletal and cardiac muscle. Importantly, it works with both porcine and mouse cardiac muscle, which are two of the most important animal models for preclinical studies.

      (3) The authors manipulate the ON/OFF equilibrium using both drugs and a genetic model of hypertrophic cardiomyopathy that has been shown to modulate the ON/OFF equilibrium. Their results generally agree with previous studies conducted using X-ray diffraction as well as biochemical measurements of myosin autoinhibition.

      Weaknesses:

      (1) While the application of pSGH to the ON/OFF equilibrium is an important advance, there are limited new biological insights since the perturbations used here have been extensively characterized in previous studies.

      We acknowledge that the biological insights provided in this study largely confirm previous findings reported in the literature. However, the primary aim of our work was to demonstrate the applicability and robustness of pSHG in probing the ON/OFF equilibrium of thick filaments. In this context, obtaining results that are consistent with established knowledge represents an important validation of the technique.

      We therefore believe that the combination of these findings with the advantages of pSHG makes our work noteworthy and supports the broader utility of this approach in future biological and physiological studies.

      (2) SGH has previously been applied to study the nucleotide-dependent orientation of myosin motors in the sarcomere (PMID: 20385845). The authors have previously interpreted the value of gamma as being a readout of lever arm position, but here, it is interpreted as a measure of ON/OFF equilibrium. When this technique is applied to intact muscle, it is not clear how to deconvolve the contributions of lever arm angle from the ON/OFF population (especially where there is a mix of states that give rise to the gamma value). This is an important limitation that is not discussed in the manuscript.

      We thank the reviewer for this insightful and important observation. The SHG signal arises from the coherent contribution of peptide bonds within the myosin molecule and therefore reflects an average angular distribution rather than a single structural state. In the study cited (PMID: 20385845), by reconstructing each contributing second-harmonic emitter within the actomyosin motor array at the atomic scale, we were able to establish a direct relationship between myosin conformation and the γ value. In particular, this analysis revealed an overall trend: γ increases as the average angle of myosin relative to the thick filament backbone becomes larger, and decreases as this angle is reduced. Based on these observations, we proposed that γ could serve as a proxy sensitive to changes in the ON/OFF equilibrium.

      However, we fully acknowledge that, especially in intact muscle, it is not possible to disentangle the respective contributions of lever arm orientation and population shifts between different structural states. The reviewer’s point is therefore well taken.

      To address this, we have revised the Discussion (pages 18-19) to more clearly acknowledge this limitation and to better articulate the interpretative framework underlying our analysis.

      Finally, we would like to emphasize that, in the present study, we aimed to minimize the contribution of actomyosin-bound states in intact trabeculae by performing experiments under conditions expected to strongly favor relaxation (low stimulation frequency, 0.1 Hz, and low temperature, 21 °C). Under these conditions, the number of strongly bound actomyosin cross-bridges during the diastolic phase is expected to be minimal, and variations in γ are therefore primarily associated with changes in the relaxed myosin population.

      (3) The R403Q mutation has previously been shown to cause an increase in ATP usage. Here, the authors measure an elevated basal ATPase rate under relaxing conditions, and they interpret this as showing increased myosin ATPase activity intrinsic to the motors; however, care should be used in interpreting these results. Work from the Spudich lab has shown that the R403Q mutation can appear as increasing motor function in some assays but depressing motor function in others (see PMID: 32284968, 26601291). Moreover, the actin-activated ATPase rate is an order of magnitude higher than the basal ATPase rate, and thus, small changes in the basal ATPase rate are unlikely to be important for physiology.

      We thank the reviewer for this thoughtful comment, and for pointing out the complexity of interpreting the functional consequences of the R403Q mutation. We agree that the functional effects of the R403Q mutation on myosin motor activity have been reported to vary depending on the experimental system and assay used, with studies showing both reduced and enhanced motor performance (PMID: 32284968, 26601291, 20560002).

      Consistent with this complexity, previous biochemical and physiological studies have reported altered energetic properties associated with the R403Q mutation in cardiac muscle. In a previous study, the relationship between cross-bridge kinetics and energetics was investigated in single cardiac myofibrils and multicellular cardiac muscle strips from human HCM samples with and without the R403Q mutation. In those experiments, cross-bridge relaxation was faster in R403Q samples and correlated with an increased energetic cost of tension generation. Basal ATPase activity measured in human samples was also elevated (4.4 ± 0.5 vs 6.6 ± 1.2 μmol L<sup>-1</sup> s<sup>-1</sup> in HCMsn and R403Q, respectively), supporting the idea that the mutation affects energetic balance in the relaxed state (PMID: 24928957).

      We also agree that actin-activated ATPase activity is substantially higher than basal ATPase activity. However, cardiac muscle spends a large fraction of the cardiac cycle in the relaxed (diastolic) state, during which myosin heads are predominantly detached from actin. Even relatively small changes in ATP turnover during this phase could therefore influence the overall myocardial energetics, and potentially contribute to the activation of signaling pathways involved in pathological remodelling.

      To address the reviewer’s concern, we have now expanded the Discussion (pages 19-20) to clarify the heterogeneous results reported in the literature for the R403Q mutation, and to more cautiously interpret the physiological implications of the observed resting ATPase activity.

      (4) The authors interpret some of their data based on the assumption that the high concentrations of drugs cause the myosin to either adopt 100% OFF or ON states. This assumption is not validated, limiting the ability to interpret the fraction of myosins in the ON/OFF states.

      We fully agree: these assumptions are not fully supported by structural data. We consistently decided to remove these analysis from the Results section and instead discuss these considerations in more general terms in the Discussion (page 18). Figure 3 and text have been modified accordingly (pages 14-15).

      (5) The ATPase measurements are innovative but hard to interpret. dATP and ATP do not have identical ATPase kinetics, meaning that it is hard to deconvolve whether the elevated ATPase rate with dATP is due to changes in the ON/OFF population and/or intrinsic ATPase activity. Similarly, mavacamten reduces the rate of phosphate release from myosin, and this effect is not strictly coupled to the formation of the OFF state (e.g., see PMID: 40118457). As such, it is difficult to deconvolve drug-based changes in the inherent ATPase kinetics of the myosin from changes in the OFF-state population.

      We thank the reviewer for this important comment. We recognize that dATP and ATP do not exhibit identical ATPase kinetics, and that Mavacamten slows steps in myosin nucleotide release that are well documented for ATP and not for dATP (for both S1; PMID: 28808052 and, more markedly, HMM; PMID: 30018063). At the same time, both compounds have been shown to perturb the regulatory state of myosin heads along the thick filament. These effects are, in turn, mediated by shifts in the distribution of ATPase intermediate states, which alter the likelihood of myosin adopting autoinhibited conformations and thereby biasing the system toward OFF or ON states (see, e.g., PMID: 39444161; PMID: 30018063). This configures a dual mechanism of action for small molecules, which we now describe more clearly in the Discussion (page 20). As suggested by the referee, we also highlight the intrinsic difficulty of disentangling compound-induced changes in the intrinsic ATPase kinetics of myosin from shifts in the population of the OFF state. Consequently, we have now better focused results and discussion section considering mainly differential effects of the drugs on SHG and ATPase measurements.

      Moreover, under the strongly relaxing conditions used in our experiments (pCa 10), the population of actomyosin-bound cross-bridges is expected to be extremely small (<0.1%), thereby minimizing any dATP-mediated activation of myosin arising from enhanced electrostatic interactions with actin (see, e.g., PMID: 31110001). This is supported by the absence of a significant effect of the nucleotide on the resting tension of myofibrils (new Fig. S1).

      We have now better clarified these points in both result and discussion section of the revised manuscript.

      Reviewer #3 (Public review):

      This is a very interesting paper extending the use of SHG to the study of relaxed muscle and its use to assess the order-disorder (and on /off) states of myosin heads in the thick filament. The work convincingly shows that SHG and the parameter gamma provide a reliable measure of the state of the myosin heads in a range of different relaxed muscle fibres, both intact and skinned, and in myofibrils. In mini pig cardiac fibres, the use of dATP and mavacamten increased or decreased the number of heads in the disordered state, respectively. On the assumption that these treatments push myosins fully into the disordered or ordered state, then this allows the fraction of ordered heads to be assessed under a wide variety of conditions. It is unfortunate that dATP treatment was not used (as mavacmten was) on rabbit psoas and mouse samples to further test this hypothesis.

      The results with the myosin mutant R403Q support the idea that this mutation reduces the fraction of myosin heads in the ordered state and that mavacamten can recover the WT situation.

      The results from SHG were compared with parallel studies using X-rays to validate the conclusions. Independent fibre ATPase data further support the conclusions.

      The work is solid and provides a novel approach to assessing the activity state of muscle thick filaments. The authors point out some of the potential uses of this approach in the future, including time-resolved SHG measurements. Indeed, jumps in mavacamten or dATP concentration with time-resolved SHG could measure the rates of entry and exit from the ordered, off state of the filament. A measurement is urgently needed in the field.

      Strengths:

      (1) The SHG signal is convincingly shown to assess the fraction of ordered/disordered myosin heads in the thick filament of a variety of muscle fibres.

      (2) The results are similar for rabbit psoas, mouse, and minipig cardiac fibres. Skinning the fibres and production of myofibrils do not change the SHG signal.

      (3) Use of myosin R403Q mutant in mini pig confirms a loss of ordered myosin heads, and the ordered heads can be recovered by mavacamten.

      (4) Parallel X-ray scattering and ATPase data support the conclusions.

      (5) Assuming that dATP and mavacamten generate 100% disordered vs ordered myosin heads respectively, then the percentage of ordered heads can be calculated for a variety of conditions.

      Weaknesses:

      (1) Issues like the effect of fibre disarray and lattice spacing on the SHG signal are not well defined.

      We thank the reviewer for raising this important point. Regarding fibre disarray, we took advantage of the spatial resolution of SHG imaging in thick samples to selectively analyse regions of the preparation in which myofibrillar organization was preserved. This approach proved particularly useful in samples such as those from HCM, where a pronounced global disarray is present but locally well-organized regions can still be identified and reliably analysed.

      Concerning lattice spacing, we have performed additional experiments on psoas muscle in which lattice spacing was modulated using dextran. This allowed us to directly assess the sensitivity of the technique to changes in inter-filament spacing.

      Figure 1 and the corresponding text have been revised accordingly to incorporate and clarify this point (page 11-12).

      (2) The, now well-defined heterogeneity of thick filament structure is not acknowledged.

      We agree that the heterogeneity of thick filament structure is an important aspect that should be acknowledged. In the revised manuscript, we have now explicitly addressed this point in the Discussion (page 21). In particular, we highlight that the capability of pSHG to probe the ON/OFF state with sub-sarcomere spatial resolution offers the future opportunity to investigate spatial heterogeneity in thick filament organization. This aspect is now clearly acknowledged and discussed in the context of the potential applications of the technique.

      (3) dATP was only used on minipig cardiac fibres. The effect of dATP on rabbit psoas and mouse cardiac fibres would be a useful comparison and would help validate the calculation of % ordered heads.

      We agree that, in the original version of the manuscript, there was a methodological imbalance in the use of dATP across preparations. To address this point, we have performed additional experiments in which the effect of dATP was also evaluated in rabbit psoas and mouse cardiac fibres.

      Importantly, the inclusion of these data has also proven useful in the discussion of the ON/OFF equilibrium across different muscle types and species. The corresponding results and discussion have been added to the revised manuscript (page 18-19).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      In addition to addressing the points in the Public Review, please also address the following points.

      (1) There are some issues with the calculated ionic strengths of the solution (or some details are missing). For example, on p. 4, it is stated that there is a 200 mM ionic strength solution that contains both 100 mM KCl and 2 mM MgCl2, meaning that the ionic strength is higher than 200 mM. There are other similar issues in other places.

      We thank the reviewer for pointing this out. We agree that there was an error in the reported ionic strength calculations and that some details were not sufficiently clear.

      We have now corrected the ionic strength values throughout the manuscript and revised the Methods section to provide a clearer and more consistent description of the solution composition (pages 6-7).

      (2) Please report standard deviations rather than standard errors.

      We thank the reviewer for this suggestion. Following the recommendations of Reviewer 1, we have reanalyzed the entire dataset using an updated statistical framework. As part of this process, we carefully evaluated the most appropriate measure of variability and have now consistently reported the corresponding error estimator throughout the manuscript.

      (3) For the statistical testing, please mention what tests were done for normalcy. It appears that some data is not normally distributed, in which case nonparametric tests should be used.

      We have now reanalysed the entire dataset with the support of Prof. Francesco Sera, who has been included as a co-author. A hierarchical (mixed-effects) statistical model was applied, explicitly accounting for the nested structure of the data, repeated measurements, and unbalanced group sizes. As part of this updated statistical framework, normality tests were performed for all datasets. When the assumption of normality was not met, appropriate nonparametric or model-based approaches were used.

      (4) Please discuss sex as a biological variable.

      Sex as a biological variable was not specifically investigated in the present study, and we agree that this represents a limitation. This point has now been explicitly acknowledged in the revised manuscript (page 5).

      (5) Please add an explicit section on limitations.

      We thank the reviewer for this suggestion. In the revised manuscript, the Discussion has been expanded to more clearly highlight the limitations of the technique and to better contextualize them in comparison with other approaches (page 21). We believe that integrating these aspects within the Discussion provides a more coherent and balanced presentation, and we have therefore chosen not to include a separate, dedicated limitations section.

      (6) Please discuss the limitations of using saturating concentrations of the drug. For example, the sensitivity of this method to detect changes in ON/OFF equilibrium at physiological concentrations is likely lower.

      The Discussion has been implemented to highlight that the sensitivity of this technique to the ON/OFF ratio could be further explored across species by employing a range of concentrations of Mavacamten and dATP, thereby better capturing physiologically relevant conditions (pages 18-19).

      (7) It is stated on p. 13 that R403Q has a higher sensitivity to mava versus dATP. I'm not sure this is supported by the data. The R403Q starts at a higher percentage of ON, and thus the effect size will be larger with mava, but this doesn't imply anything about sensitivity (which implies concentration dependence).

      Our statement was not intended to imply a difference in sensitivity in terms of concentration dependence, but was instead based on the statistical outcome of our measurements. Specifically, while a significant response was observed in the presence of mavacamten, no statistically significant response was detected upon dATP application in the R403Q condition.

      We acknowledge that this does not constitute evidence of differential sensitivity per se, and we have revised the text accordingly by removing the concept of sensitivity to avoid potential misinterpretation.

      (8) Please add a discussion of the potential contributions of RLC phosphorylation to ON/OFF regulation.

      We agree with the reviewer that the potential involvement of RLC phosphorylation in ON/OFF regulation is an important aspect to consider in future work, and we have now acknowledged this point in the revised Discussion.

      Reviewer #3 (Recommendations for the authors):

      Some things that need to be clarified:

      (1) The comparison of skinned and intact muscle. These appear to give an unaltered SHG gamma signal (Figure 2), but a change in lattice spacing is expected between skinned and intact fibres. There is no mention of the lattice spacing of the samples or whether this was controlled. The implications are that SHG is independent of lattice spacing and/or the fraction of ordered myosin heads is independent of lattice spacing - each of which would be a useful result.

      We thank the reviewer for this important observation and agree that the role of lattice spacing is a relevant factor in the interpretation of the SHG signal. To address this point, we have performed an additional series of experiments on rabbit psoas muscle in which lattice spacing was modulated using dextran. These measurements allowed us to directly assess the sensitivity of the SHG signal to changes in interfilament spacing. We observed relatively small effects, but in the expected direction.

      The lack of appreciable differences between skinned and intact preparations may therefore be explained by the limited sensitivity of the technique to detect the relatively small variations in lattice spacing associated with these conditions.

      (2) In comparing the WT and mutant mini pig cardiac data, the authors note that the mutant fibre has more disarray. A comment on the effect of disarray on the SHG signal would be helpful. How much of the difference between WT and mutant could be due to this disarray? What happens if more vs less disordered areas of the fibre are compared?

      Myofibrillar disarray is indeed a characteristic feature of HCM tissue and is more evident in the R403Q minipig samples compared to WT. To minimize potential polarization artefacts related to structural disorganization, the entire field of view was first examined to identify regions of interest (ROIs) where sarcomeres showed minimal local disarray. Data acquisition and analysis were restricted to these locally well-aligned regions, as described in the Methods section. Therefore, the comparison between WT and mutant samples was performed on locally well-oriented regions rather than on highly disorganized areas.

      We have now clarified this point in the revised manuscript to better explain how ROI selection, restricted to locally aligned regions, minimizes the potential contribution of myofibrillar disarray to the pSHG measurements (page 14).

      (3) I note in Figure S3 that there is a bigger dispersion in the minipig data than mouse or rabbit. The minipig may also show a non-normal distribution. Has this been considered?

      Following the reviewer’s suggestion to include dATP measurements in additional species, we have extended the interspecies analysis and revised Figure S3 to provide a direct comparison across rabbit psoas, mouse cardiac, and minipig cardiac samples under the investigated conditions.

      From this more comprehensive dataset, no substantial differences in data dispersion are apparent among the different muscle types. Moreover, as part of the updated statistical framework, normality tests were performed for all datasets

      (4) A note about the heterogeneity in the regulation of thick filaments along their length should be added. There is significant evidence for a difference in regulation between the MyBP-C regions and the rest from single molecule studies (Kad lab) and interference X-ray signals (London Kings group). This is probably beyond the current resolution of the SHG, but the complexity should be acknowledged. Heterogeneity in the thick filament is also apparent from cryo-EM images of relaxed muscle and thick filaments.

      This important point is now acknowledged the revised Discussion (page 21).

      (5) Interpretation of the effects of mavacamten on the ATPase data ae complicated by the observation that mava is an inhibitor of myosin independent of the effect on the order- disorder of thick filaments. I.e. mavacamten will inhibit myosin S1.

      This point has now been addressed (see response to Reviewer #2, point 5 weaknesses).

      Minor issues:

      (1) P7 line 2: where/were purchased from Sigma.

      (2) P7 last but one line: R403Q/R4303Q.

      (3) P8 last line: 2-deaoxyATP 2-deoxyATP.

      (4) Results, p9: in the section title and first line, replace psoas with rabbit psoas.

      (5) Line 5: ROI is not defined, only in the Figure legend.

      (6) Mavacamten: 50 uM used in psoas and 10 uM in cardiac fibres. A note on why would help those not familiar with this literature. Similarly, in Figure S3, presumably the mava concentration was saturating in each case.

      (7) P14, last paragraph, line 3: Figure 4 should be Figure 5.

      (8) P17, 5 lines from the end: as effective at 2-deoxyATP at/as.

      (9) P19, lt line: fibber/fiber.

      All these minor points have been fully addressed. We thank the reviewer for noting them.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This useful study uses creative scalp EEG decoding methods to attempt to demonstrate that two forms of learned associations in a Stroop task are dissociable, despite sharing similar temporal dynamics. However, the evidence supporting the conclusions is incomplete due to concerns with the experimental design and methodology. This paper would be of interest to researchers studying cognitive control and adaptive behavior, if the concerns raised in the reviews can be addressed satisfactorily.

      We thank the editors and the reviewers for their positive assessment and constructive feedback on our work. We also thank the editor for communicating with reviewer #1 regarding our thoughts on their comments. We hence revised the manuscript based the new feedback from reviewer #1. Please see below our responses to each comment raised in the reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study focuses on characterizing the EEG correlates of item-specific proportion congruency effects. In particular, two types of learned associations are studied. One association involves associations between stimulus features and control states (SC), and the other involves stimulus features and responses (SR). Decoding methods are used to identify time-resolved SC and SR correlates.

      The authors conclude that SC and SR associations can independently and simultaneously guide behavior. This conclusion is based on results showing that SC and SR correlates are (1) not entirely overlapping in cross-decoding, (2) simultaneously observed on average over trials, (3) independently correlate with RT, and (4) have a positive within-trial correlation.

      Strengths:

      Fearless, creative use of EEG decoding to test tricky hypotheses regarding latent associations.

      Nice idea to orthogonalize ISPC condition (MC/MI) from stimulus features.

      Thank you for acknowledging the strength in EEG decoding and design. We have addressed all your concerns raised below point by point.

      In my view, the ability to address this issue with additional analyses is relatively limited. I cannot think of a solid way to escape this issue in the present design. Adding a nuisance regressor to their RSA regression, which I suggested in my previous response, may reduce the bias, but the efficacy of this would be limited to controlling only a certain kind of phase-dependent confound (a 'main effect' component of study phase, i.e., one that is constant across other conditions; see next message for discussion).

      Rather than new analyses, I think a more straightforward revision might be to modify the conclusions advanced in the paper, so that they are more solidly supported by the design and evidence. In my opinion, this design is ill-posed to solidly identify SC and SR representations. As a result, I think that any framing that alleviates pressure on this design to yield 'solid' evidence for identification of SC and SR representations, and instead emphasizes stronger areas of this work, would be constructive.

      For example, one potential framing is to advance the idea of SC and SR representations, and discuss an idealized design that could identify them by orthogonalizing stimulus features from ISPC, which are genuinely novel and useful ideas. The current design could then be presented as an opportunistic or initial case study of testing this question, while acknowledging its limitation upfront. The goal here would be to frame the study in a way that allows for the results to be presented with an appropriate grain of salt, while also illustrating the authors thoughtfulness and creativity in devising analyses to test for latent associative representations. In this case the 'solid' label would reference the authors' reasoning and analyses rather than design and conclusions.

      I only intend this example as an illustration; there may be several ways of framing this paper so that it is on more 'solid' ground, and I don't want to dictate how exactly authors should write their paper.

      Nonetheless I think this issue is important, not only to avoid faulty inference, but also to avoid establishing counterproductive precedents in this field. For example, if students read a paper whose conclusions are labeled "solid" but that nevertheless has critical flaws in its design, then those students may be misled in their own work. But if the flaws were discussed transparently and critically, and the strength of the conclusions were de-emphasized relative to other aspects of the paper, students may not only be inspired by the ideas developed in the paper, but also come away with knowledge about the issues of experimental design.

      Discussion of the weaknesses in the conclusions and an additional potential analysis:

      The key goal of this study is to identify SC/SR representations, which requires decoupling stimulus features from item-specific proportion congruency (ISPC), but the study-phase confound contaminates this decoupling. I think this impacts both their cross-phase decoding and RSA analyses.

      In their results (lines 139-144):

      "This standard ISPC manipulation can test whether neural representations of controlled and non-controlled information are wrapped on the same trial by combing with the following EEG analysis (See Methods). However, it potentially mixes color identity with SC, word identity with SR, and the ISPC between SC and SR. To deconfound these factors when estimating SC and SR association representations on each trial, we modified this paradigm by flipping the ISPC contingencies across different phases of the task."

      SC/SR representations are higher-order conjunctive classes, formed by the interaction of lower order variables (Color, Word, and ISPC). This means that successfully identifying these representations relies on demonstrating that each class can be reliably individuated from every other class in a manner that cannot be explained by (1) representation of shared lower-order features, such as stimulus color or response, and (2) trivial nuisance factors such as study phase. However, in this design, the trivial factor of study phase is strongly confounded with the ISPC contingency flip.

      Regarding RSA: If study phase indeed leads to trivial separability, it seems that similarity among conditions within the same phase would be inflated because the study phase does not appear in the RSA model (Figure 12). In which case, the SC/SR coefficients would be inflated, as the SC and SR models are entirely within-phase.

      Nevertheless, an additional control analysis may be possible here. Looking at this regressor set, it seems possible to me to fit a model where the predominant phase is entered as an additional covariate. (If I am reading this correctly, phase would appear as a 2x2 block-diagonal matrix). I suggested this in my last review, but in their recent letter, authors refused. I do not understand why, as I think this regressor set should be identifiable, but perhaps I am wrong here.

      That being said, I do not think adding this nuisance covariate would fully solve the issue. The lower-order regressors (Color, Word, ISPC) are defined as the EEG responses shared across phases of the study. To interpret the higher-order SR/SC coefficients, the lower-order components must be fully partialled out. But the putative impact of study phase contaminates this interpretation, as study phase could trivially decrease the similarity of lower-order terms (e.g., decreasing Color similarity between phase 2 vs 3). In which case, the partialling would be expected to be incomplete.

      This incomplete partialling is the RSA analogue of the issue in interpreting cross-phase decoding discussed above. As there, so too here I do not see a solid way around it in the present design. This is why in my previous review I referred to the addition of a phase covariate in the RSA regression as a "band-aid": it controls for some problems (main effect of phase), but not all (interactions of phase and lower-order terms).

      To summarize, I think that RSA would offer an additional opportunity to control for this potential confound, albeit in a limited sense (study phase effects that are consistent across conditions). But in a more general and rigorous sense, to me, the SC and SR terms in the RSA regression also seem susceptible to the same weakness as the cross-phase decoding analysis.

      We thank the reviewer for taking the additional time and effort to provide the new comments. Following the reviewer’s suggestion, we revised the language regarding the conclusions of this project (page 2,5,27-28) and explicitly discussed the weaknesses of the design on decoding and outline a possible solution for future studies in the Discussion section:

      “A limitation of the current design is that in theory temporally structured noise (e.g., autocorrelation in EEG data) may bias the decoding accuracy due to the blocked design. Although the present data provided no evidence that the decoding results in this study were biased by temporally structured noise, future studies should aim to develop experimental designs that eliminate this potential confound at the source. One potential solution would be to introduce additional phases flipping ISPC manipulations. At the same time, enough trials must be included in each phase to ensure the strength of the ISPC effect within each phase. A careful balance between session number and length will be helpful to optimize the duration of such a design.”

      As eLife also publishes review report, below we also summarize the three control analyses we ran and our reasoning of how they (would) address the issue of temporally structured noise for interested readers to assess:

      We acknowledge the theoretical issue of temporally structured noise (TSN) in our design when classes were decoded across different phases. However, the key question for the current data is whether there is empirical evidence that the decoding results were actually driven by TSN. To clarify this issue, we summarize several lines of evidence suggesting that the decoding results were not attributable to TSN:

      (1) Split-half cross-validation. We split the EEG data from Phase 2 and the combined Phases 1 and 3 into chronological first and second halves. Phases 1 and 3 were combined because they shared the same MC and MI assignments. This resulted in four possible combinations, each consisting of eight classes drawn from different phases: combination 1 included the first half of Phase 2 and the first half of Phase 3; combination 2 included the first half of Phase 2 and the second half of Phase 3; combination 3 included the second half of Phase 2 and the first half of Phase 3; and combination 4 included the second half of Phase 2 and the second half of Phase 3. We trained the decoders on one combination and tested them on another and then averaged the decoding results across all possible training-test assignments. The similar decoding patterns observed across these analyses (Fig. 6a,b) further confirmed that the decoding results were not driven by TSN.

      This analysis is conceptually similar to the “cross-phase” decoding analysis suggested by the reviewer in the first round of review. We also performed an additional distance-based control analysis (see below) to further test whether the decoding results could be explained by TSN, without imposing the constraint used in the split-half cross-validation that trials from the two phases had to fall within a 400-trial window.

      (2) Distance analysis. We predicted that if a test trial is closer to a training trial of the same trial type, the higher similarity in TSN between the training and test data would more strongly inflate the decoding accuracy of the test trial, resulting in a negative correlation between distance between a test trial and its closest training trial of the same type and the test trial’s decoding accuracy. However, we did not observe such a negative pattern (Fig. 6c). Note that this distance was defined with respect to trials of the same type, rather than absolute chronological time.

      (3) Shuffled analysis. If the decoding results were primarily driven by TSN, either at a short-term or long-term timescale, then shuffling the condition labels within each mini block should preserve the TSN structure present in the real data. In that case, the decoding results from shuffled data should not differ from those observed from real data. However, we found the significant difference between real data and shuffled data as shown in Author response images.

      Author response image 1.

      Shuffling analyses with stimulus-locked data support separable SC and SR subspace. (a) Group average decoding accuracy of all 16 experimental conditions as a function of time after stimulus onset. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (b) Group average t values of representational strength for each factor over time. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (c) SC and SR association results from Fig. 1b.

      Author response image 2.

      Shuffling analyses with response-locked data support separable SC and SR subspace. (a) Group average decoding accuracy of all 16 experimental conditions as a function of time after stimulus onset. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (b) Group average t values of representational strength for each factor over time. Squares below the lines indicate the significant time points between real data and shuffled data (cluster-based permutation test, cluster-forming threshold p < 0.001, cluster-level p < 0.05). (c) SC and SR association results from Fig. 2b.

      We thank the reviewer for the suggestion on RSA with phase. There are some concerns for this analysis:

      First, we think that the suggested analysis may be difficult to interpret. Because the SC and SR conditions differ across phases. Regressing out phase in RSA could also remove SC and SR information.

      Second, based on the reviewer’s comment, we understand that the suggested analysis may still not provide a clear falsifiable criterion for determining whether the results could be driven by the theoretical TSN issue inherent in the design.

      Third, the three control analyses we have performed examine this issue from different perspectives and collectively provide no evidence that the results were driven by TSN.

      Other readers may, like me, be puzzled by the selection of this particular experimental design to test this question of SC and SR coding, given the temporal confound among SC/SR classes, and given that there would seem to be many possible designs that are less confounded. For example, why not use a design where ISPC was swapped/shuffled several more times within each subject, so that PHASE is more orthogonal to long-timescale noise? Isn't ISPC learning fast enough to support learning phases shorter than 700 trials? Such readers would likely appreciate a frank discussion of this dilemma, and a motivation for the choice of the present design, within the manuscript.

      Thank you for your suggestion regarding the design. It is possible that the (re-)learning of ISPC can be fast. That said, enough trials are required to obtain a robust ISPC effect for each phase after the ISPC flips. Given that the EEG scanning (not including capping) in current design was about 1.5 hours, it is impractical to have both more sessions for a more orthogonal design and long sessions for robust within-session ISPC effects. We chose to maximize the latter because flipped behavioral ISPC effect in each session is the basis for the following EEG analysis. We have included the reviewer’s suggestion as a potential design solution for future studies in the Discussion section mentioned above on page 27.

      Pre-stimulus coding:

      To explain the apparent pre-stimulus coding of several task variables, the newest version of the manuscript proposes that subjects were proactively coding these variables via predictive mechanisms. This is an interesting account of item-specific control. It is also surprising, given that item-specific control mechanisms are typically conceptualized as reactive or stimulus-driven phenomena. But I think support for a proactive control account was incomplete. The mechanistic logic was not presented, and no hypotheses under this account were developed or tested. So I would suggest pinning down some hypotheses here and actually putting this account to the test.

      Thank you for raising this important point. Although ISPC effects are considered reactive, in our design the long sessions may create a temporal context for the participants to differentiate the current control demand linked to each color. The maintenance of such contextual information needs to span across trials, leading to pre-stimulus coding that proactively guides the control demand for each color. This claim is not central to this manuscript, which investigates whether SC and SR representations simultaneously guide behavior. Additionally, we do not think the current design is well-equipped to test this hypothesis because the pre-stimulus onset is the only supporting evidence. In the revised manuscript, we discussed this as a future research direction and proposed a design that aims at better isolating proactive control signal on page 25.

      Random slopes were omitted due to convergence failure, but this can inflate false positive inferences (e.g., Barr et al. 2013), and doesn't really motivate a minimal model. I'd suggest trying a slightly reduced model (e.g., drop correlations via `slope || subject`) using buildMer automated selection, or switching to brms.

      We indeed tried both the full model of random effects (i.e., considering covariance between all slopes and intercept) and a reduced model without any covariance (i.e., listing each random slope separately without intercept in lme4). However, neither model converged at all time points.

      Reviewer #2 (Public review):

      Summary:

      In this EEG study, Huang et al. investigated the relative contribution of two accounts to the process of conflict control, namely the stimulus-control association (SC), which refers to the phenomenon that the ratio of congruent vs. incongruent trials affects the overall control demands, and the stimulus-response association (SR), stating that the frequency of stimulus-response pairings can also impact the level of control. The authors extended the Stroop task with novel manipulation of item congruencies across blocks in order to test whether both types of information are encoded and related to behaviour. Using decoding and RSA they showed that the SC and SR representations were concurrently present in voltage signals and they also positively co-varied. In addition, the variability in both of their strengths was predictive of reaction time. In general, the experiment has a solid design and the analyses are appropriate for the research questions.

      Strengths:

      (1) The authors used an interesting task design that extended the classic Stroop paradigm and is effective in teasing apart the relative contribution of the two different accounts regarding item-specific proportion congruency effect.

      (2) Linking the strength of RSA scores with behavioural measure is critical to demonstrating the functional significance of the task representations in question.

      We thank you for acknowledging our work on design and brain-behavior analysis. We have addressed your concerns raised below.

      Weaknesses:

      I still have some doubts on the effectiveness of the experimental manipulation on Phase 2: although the ISPC effect is still present, it is much weaker in comparison, suggesting the participants did not learn the contingency statistics in Phase 2 as well as they did in the other phases, due to either the lingering effect of the previous phase or an inherent bias towards one color pairs. Perhaps by separately plotting the earlier and later blocks of Phase 2 any difference can be revealed if it exists. This behavioral difference could result in unequal levels of SC/SR representation across phases, which may raise problems when data were combined for analyses that assume the neural effects are equivalent.

      Thank you for your concern about this important issue. We agree with the reviewer that the true SR/SC levels may not be equivalent between Phase 2 and Phase 1/3. Nevertheless, because the manipulation of ISPC is binary, the decoders were trained to test whether the neural signals represent the two levels of ISPC (i.e., a higher vs. a lower level) differ systematically. The decoding analysis does not require that the neural effects of SC and SR must be numerically equivalent between phases (i.e., it is not necessary that the two levels are equidistant from the center point of SC/SR. Indeed, the decoding analysis only requires that the two levels are different). For example, if the ISPC level ranges from -1 to 1 and the EEG signals can reliably decode ISPC levels of -0.5 and 0.7 (i.e., two unequal levels), it can still be treated as supporting evidence that ISPC levels are encoded in the EEG signals. The same logic applies to the representational subspace analysis. As to the RSA, as can be seen in Fig. 12, the regressors are also binary, encoding whether two experimental conditions share the same SC/SR level without assuming equivalent neural effects. Thus, we argue that the reported decoding and RSA can still test the encoding of SR and SC. We discussed this issue on page 24.

      Following the reviewer’s comment, we plotted the ISPC effects in first and second half of Phase 2 separately (the figure below). We also tested whether the ISPC effects differ qualitatively between the two halves using a 3-way ANOVAs separately on RT and Error rate. The results showed that the time (the first half vs. the second half) × Congruency × ISPC interaction was not significant for either RT data (F<sub>(1,39)</sub> = 3.40, p > 0.05) or error rate (F<sub>(1,39)</sub> = 2.49, p > 0.05), suggesting that ISPC effect did not systematically change over time in Phase 2 See Supplementary Figure 9.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In the revised manuscript, we have expanded the real-data analyses, clarified the relationship between CMP and prior modulated Poisson models, and added discussion of model limitations and future extensions. In summary, the major changes include:

      (1) We revised the Introduction, Results, Methods, and Goris-model appendix to clarify the relationship between CMP and prior modulated Poisson models. In particular, we now emphasize that the key distinction is CMP’s continuous-time stochastic gain process.

      (2) We moved the simulation-based recoverability analysis from Appendix 3 into the main Results section (Figure 3 in the revised manuscript), making the validation of the inference procedure more visible to readers.

      (3) We added new analyses of the inferred gain process. Specifically, we now show the cross-trial gain mean and cross-trial gain variance in Figure 4A to assess whether gain captures stimulus-locked structure, and we added an analysis of pre- versus post-stimulus cross-trial gain variability in Figure 4C to test for gain-variability quenching during stimulus presentation.

      (4) We clarified the definitions and implementation of the Baseline Poisson, Poisson-GP, and Goris-style comparison models, including the role of the smoothness prior on the stimulus drive.

      (5) We expanded the Discussion to describe future extensions to population recordings, including a GPFA-inspired extension with low-dimensional shared gain activity across neurons.

      (6) We added a Discussion paragraph clarifying that the current CMP model captures Poisson and super-Poisson variability, but not sub-Poisson variability, and outlined possible extensions using spike-history terms, renewal-process likelihoods, or alternative count distributions.

      eLife Assessment

      This work of fundamental significance introduces a novel statistical model of spiking activity that incorporates continuous−time gain modulation. The authors provide exceptional evidence that the model outperforms earlier approaches and alternative candidates in capturing spiking responses across multiple visual areas in the macaque. Beyond its methodological contribution, the study offers new insights into how stimulus−driven variability and internally generated gain fluctuations evolve over time and between brain areas. The framework is likely to find broad application beyond the datasets examined here.

      We sincerely thank the Senior Editor, Reviewing Editor, and both reviewers for their careful evaluation and constructive feedback. We are encouraged by the positive assessment of the work and by the recognition of its methodological and conceptual contributions. We especially appreciate the acknowledgement that the continuous-time formulation provides a useful framework for modeling gain modulation in spiking activity, improves upon earlier approaches in capturing responses across multiple visual areas, and offers new insights into how stimulus-driven variability and internally generated gain fluctuations evolve over time and across brain regions.

      In the revised manuscript, we have addressed the reviewers’ comments by clarifying the relationship between CMP and prior modulated Poisson models, strengthening the presentation of the simulation-based recoverability analysis, adding new validation analyses of the inferred gain process, and expanding the Discussion of model scope, limitations, and future directions. In particular, we now more clearly distinguish the continuous-time gain process in CMP from Goris-style models with constant or piecewise-constant gain, move the simulation recoverability analysis into the main Results, examine trial-averaged inferred gain and gain-variability quenching, clarify the definitions of the baseline and comparison models, and discuss extensions to population recordings and sub-Poisson variability.

      We believe these revisions improve the clarity, rigour, and scope of the manuscript. Below, we address each reviewer comment in turn and describe the corresponding changes made in the revised manuscript.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      In this manuscript, Rupasinghe and co−authors introduce a new statistical model for spiking neurons. Building on earlier work, they propose to model spikes as arising from a Poisson process whereby the firing rate is the product of stimulus drive and astimulus−independent gain signal. The critical innovation of this work is that the gain signal is modeled in continuous time. Earlier explorations of this statistical construction treated the gain−signal as constant within a trial. This innovation is elegant and important. It makes the model richer, more plausible, and more broadly applicable. The authors show that the model parameters are recoverable from realistic amounts of data and then apply the framework to previously studied datasets. They show that the new model outperforms earlier models and alternative candidates in capturing spiking data across four visual areas of the macaque monkey. Analysis of the model parameters replicates some earlier findings and uncovers several new insights. The model and fitting methods can be broadly applied to partition different types of signals and noise from spiking data and are likely to be widely adopted in the systems neuroscience community.

      Strengths:

      (1) Through clever use of advanced statistical techniques, the authors manage to infer critical information from single−trial single−cell data.

      (2) The question of which aspect of a spike train is signal and which is noise is omnipresent in neuroscience. By improving our ability to characterize the distinct factors that shape spiking activity, this work makes a fundamental contribution to the literature.

      We sincerely thank the reviewer for the thoughtful and detailed evaluation of our manuscript. We are pleased that the continuous-time formulation and its methodological contributions were viewed as elegant, important, and broadly applicable. We also appreciate the reviewer’s recognition that the framework provides a useful way to separate stimulus-driven and modulatory components of neural variability from single-trial, single-cell data. The reviewer’s comments helped us improve the precision of our framing, clarify the relationship between CMP and prior modulated Poisson models, and strengthen the validation of the inferred gain process. Below, we respond to each point in turn and describe the revisions made in the manuscript.

      Weaknesses:

      Overall, I find the work impressive and important. I have a couple of questions and suggestions.

      (1) The work is entirely focused on single−cell data. While this is a great starting point, expanding the approach to spiking activity in neural populations is an importantfuture goal.

      We thank the reviewer for this important suggestion. We agree that extending the CMP framework to population recordings is a natural and important direction for future work. In the present study, we focus on single-neuron responses to establish the continuous-time model, validate the inference, and characterize how stimulus-driven activity and stochastic gain fluctuations can be separated at the level of individual cells. However, the same modeling principles could be extended to simultaneously recorded neural populations by introducing shared latent structure across neurons. For example, one natural direction would be to combine CMP with ideas from Gaussian Process Factor Analysis [Keeley et al., 2020], using low-dimensional shared gain activity to capture population-wide fluctuations, while retaining neuron-specific stimulus-driven components. Such an extension would allow the model to capture correlated variability and shared modulatory dynamics across neural ensembles. In the revised manuscript, we have expanded the Discussion to describe this possible future extension to population recordings.

      To address this comment, we expanded the Discussion (Page 14: lines 473-478) to describe a possible GPFA-inspired extension of CMP to population recordings.

      (2) Line 49−53: These statements seem incorrect to me. The modulated Poisson model , as introduced in Goris et al (2014), is a process model that can perfectly be used to generate spike trains (within a trial, spiking emerges from a Poisson process, which canbe homogeneous or inhomogeneous). Moreover, the model contains a parameter thatrepresents the duration of the counting window (delta t). The dependency of over− dispersion on the size of the time bins for real neurons is shown in Figure 1b (inset plot) of that paper (and shown to resemble the model prediction). This time− dependency was further explored by the same authors in Goris et al (2018 − Journal ofVision) and also in Henaff et al (2020 − Nature Communications). I suggest that the authors rephrase this argument (here and at some later points in the paper). They could just say that the Goris model makes the simplistic and implausible assumption that, within a given trial, gain does not fluctuate. This is clearly an important limitation and the key difference with the continuous model introduced here.

      We sincerely thank the reviewer for identifying this lack of clarity in our original description. We agree that our original description was not sufficiently precise. The modulated Poisson model introduced by Goris et al. (2014) is indeed a generative process model and can be used to generate spike trains, with spiking arising from a Poisson process that may be homogeneous or inhomogeneous within a trial. We apologize for implying otherwise.

      Our intended point was that, in the original formulation, the modulatory gain is represented as a scalar random variable associated with a counting window or trial, and therefore does not explicitly model gain as a continuously time-varying process within a trial. Thus, the key limitation addressed by CMP is not the use of a Poisson process, but the assumption that gain is constant or piecewise constant over the relevant interval.

      In the revised manuscript, we have rephrased the Introduction to clarify this distinction. We now describe the Goris model more accurately as a modulated Poisson framework in which gain is constant over the counting window, and we emphasize that CMP extends this framework by replacing this assumption with a continuous-time stochastic gain process. We have also added discussion of related time-dependent analyses and extensions [Goris et al., 2018, H´enaff et al., 2020], as thoughtfully suggested by the reviewer.

      In addition, we revised the Results and Methods to clarify how the Goris-style baselines were implemented in our comparisons. Specifically, all Goris-style results reported in the main model comparisons use versions with a smoothness prior on the stimulus drive, where the stimulus-dependent firing rates are set to the smooth firing-rate estimates obtained from the Poisson-GP model. This ensures that the comparisons focus on different assumptions about the temporal structure of the gain process, rather than differences in stimulus-drive estimation. We also clarified the comparison to Goris-style variants without this smoothness prior, in which the stimulus-drive parameters are estimated directly under the corresponding Goris-style likelihood (Figure 5 - figure supplement 2). These results show that the smoothness prior on the stimulus drive substantially improves model performance. Finally, we revised the Figure 1 caption and the Goris-model appendix to make these distinctions explicit.

      To address this comment, we revised the Introduction (Pages 2-3: Lines 49-77), Results (Page 9: Lines 263-266, 273-277, Page 11: Lines 319-326), Methods (Page 21), Figure 1 caption, and Goris-model appendix to clarify that CMP extends the Goris framework by modeling gain as a continuously time-varying process within trials.

      (3) Line 54−55: I think the first part of the claim is a bit misleading. There is nothing in the Goris model that would inherently limit it to homogeneous Poisson processes, as seems to be implied by this description. The model is built on the assumption thatspike generation within a trial arises from a Poisson process. This may very well be an inhomogeneous Poisson process (i.e., a stimulus−dependent time−varying firing rate). Homogeneous and inhomogeneous Poisson processes both give rise to Poisson distributed spike counts (and thus a mixture of Poisson distributions across trials in the Goris model). I suggest the authors clarify this description a bit. Note that the two model variants illustrated in Figure 1b and c were also explored in Henaff et al (2020 − Nature Communications).

      We thank the reviewer for this helpful clarification. We agree that the Goris model is not limited to homogeneous Poisson spiking and can incorporate a stimulus-dependent, time-varying firing rate within trials. We did not intend to imply otherwise, and we have revised the relevant text to avoid this misunderstanding.

      Our intended point was that, in formulating continuous-time extensions of the modulated Poisson framework, we explicitly model the time-varying stimulus drive using a smoothness prior, as in the CMP framework, and then consider different assumptions about the temporal structure of the gain process, including constant gain and independently resampled gain across time bins. This highlights the distinction between piecewise-constant gain assumptions and the fully continuous gain process introduced in CMP.

      In the revised manuscript, we have clarified this distinction in the Introduction, Results, and Methods. We now state that the Goris-style variants use stimulus-dependent, time-varying Poisson firing rates, and that the main difference between these variants and CMP lies in the temporal structure assumed for the gain process. We have also acknowledged related variants explored in Goris et al. [2018] and H´enaff et al. [2020], and clarified that our continuous-time formulations of the Goris model differs by imposing a smoothness prior on the stimulus drive. This allows us to estimate a regularized time-varying stimulus component while comparing different assumptions about gain dynamics, ensuring that the comparison focuses on the temporal structure of the gain process rather than differences in stimulus-drive estimation. We also highlight in Figure 5 - figure supplement 2 that even for the Goris-style models, versions that use a smoothness prior on the stimulus drive outperform versions that do not, which are closer to the original modulated Poisson formulation.

      To address this comment, we revised the Introduction (Pages 2-3: Lines 49-77), Results (Page 9: Lines 263-266, 273-277, Page 11: Lines 319-326), Methods (Page 21) to clarify that the Goris-style variants allow stimulus-dependent time-varying firing rates and differ from CMP primarily in their assumptions about gain dynamics. We also added citations to related time-dependent extensions of the modulated Poisson framework.

      (4) The extension to the continuous case is very elegant!

      We thank the reviewer for the positive comment and are pleased that the continuous-time formulation was viewed as elegant.

      (5) I find the result shown in Appendix 3 critically important. The recoverability of the model for realistic amounts of data is foundational for the rest of the paper. I wouldconsider including this analysis in the main results section. Not all readers may check Appendix 3, but they should know about this result.

      We thank the reviewer for emphasizing the importance of this result. We agree that demonstrating parameter recoverability is foundational to the paper and should be visible to readers in the main Results section. In the revised manuscript, we have moved the simulation-based validation from Appendix 3 into the main Results. This section now describes the synthetic CMP dataset, the inference procedure used to estimate the latent stimulus-drive and gain processes, and the comparison between true and inferred GP hyperparameters. These results show that the proposed inference framework can accurately recover the ground-truth stimulus drives, gain processes, and hyperparameters from realistic amounts of simulated data.

      To address this comment, we moved the simulation-based recoverability analysis from Appendix 3 into the main Results section (Page 6: Lines 200-211 and Figure 3).

      (6) Figure 3: I am wondering whether the inferred gain is capturing some response fluctuations that originate from the cell’s phase−selectivity. Could the authors compute the trial−averaged inferred gain (ideally, aligned to stimulus−phase at the start of the trial if this experimental parameter varied across repeats)? If they have successfully partitioned the response variance, the trial−averaged gain should have no systematic temporal structure. If it has a sinusoidal modulation, it may partially capture stimulus−drive. This could be an interesting test to run on all model fits to further validate that the partitioning into a signal and noise component succeeded as intended.

      We thank the reviewer for this insightful suggestion. We agree that verifying that the inferred gain does not capture stimulus-driven structure is an important validation of the model. In the revised manuscript, we have added the trial-averaged inferred gain to Figure 4A for the example neuron. This analysis shows that the trial-averaged inferred gain is relatively flat and neither resembles the inferred stimulus drive nor exhibits clear stimulus-locked temporal structure. This suggests that trial-specific gain fluctuations largely average out across repeats, consistent with the interpretation that the gain process captures random trial-to-trial variability rather than stimulus-driven activity.

      We also note that a direct comparison of this inferred gain trace across methods is not possible for the Goris-style baselines, because these models do not infer a continuous trial-specific gain process. Instead, they marginalize over scalar or time-bin-independent gain variables when computing likelihoods and Fano factor curves. Thus, the trial-averaged gain diagnostic is specific to the CMP model, where the posterior over the continuous-time gain process is explicitly inferred.

      To address this comment, we added the trial-averaged inferred gain to Figure 4A and clarified that it does not show a clear stimulus-locked temporal structure (Page 7: Lines 229-236).

      (7) One common observation that is currently not explored is the quenching of neuronal response variability following stimulus onset (Churchland et al 2010 − NatureNeuroscience), which was suggested to reflect a quenching of gain variability in Goris et al (2024 − Nature Reviews Neuroscience). Building on the previous suggestion, the authors could compute the temporal evolution of cross−trial gain variability from the inferred gain traces. Do they recognize a reduction in gain variability following stimulus onset? If so, it would be worthwhile to show this.

      We sincerely thank the reviewer for this valuable suggestion. We agree that examining whether gain variability decreases following stimulus onset provides an important test of the inferred gain process. In the revised manuscript, we have added an analysis of the temporal evolution of cross-trial gain variability before and after stimulus onset.

      First, in Figure 4A, we now show the cross-trial variance of the inferred gain for the example neuron. This trace shows larger gain variability during the stimulus-off period and a reduction following stimulus onset, suggesting that the inferred gain captures a stimulus-related quenching of trial-to-trial variability. To quantify this effect across the population, we also added a pre- versus post-stimulus comparison in Figure 4C. Following the approach of Churchland et al. [2010], we compared gain variability in two matched 400-ms windows: a pre-stimulus window ending at stimulus onset and a stimulus-period window beginning 100 ms after stimulus onset. For each neuron and stimulus condition, we computed the cross-trial variance of the inferred gain at each time bin, averaged this quantity within each window, and then compared the pre- and post-stimulus values across neuron-stimulus pairs.

      This analysis revealed a significant reduction in inferred gain variability following stimulus onset (one-sided paired Wilcoxon signed-rank test, p≤ 10<sup>−15</sup>), consistent with gain variability quenching [Churchland et al., 2010, Goris et al., 2024]. We now report this result in the main text and illustrate it in Figure 4A and Figure 4C. This provides additional evidence that the inferred CMP gain captures meaningful trial-to-trial variability and its temporal modulation around stimulus presentation.

      To address this comment, we added the cross-trial gain variance trace to Figure 4A and a population-level pre- versus post-stimulus gain-variability quenching analysis (Page 7 and 8: Lines 239-248) in Figure 4C.

      (8) Line 543−565: I want to make sure I understand the Baseline Poisson model and Poisson−GP correctly. For the baseline model, I had imagined that the authors would simply use the stimulus−conditioned PSTH as an estimate of the time−dependent firing rate, coupled with an inhomogeneous Poisson process assumption. But they additionally assume a Gamma prior on the firing rate to compensate for the sparsenessof the data (sometimes only 5 repeats per condition). The Poisson−GP includesexactly the same model components, but now the time−dependent firing rate is modeled by a Gaussian process. Doing this massively improves the goodness−of−fit (Fig 4A). Do I understand this correctly?

      We thank the reviewer for this careful reading. Yes, this understanding is broadly correct, and we have revised the manuscript to clarify the relationships among the Baseline Poisson, Poisson-GP, and Goris-style models. The Baseline Poisson model estimates a stimulus- and time-dependent firing rate independently for each stimulus condition and time bin, using a Gamma-Poisson formulation to regularize the estimate when the number of repeats is limited. The Poisson-GP model uses the same conditionally Poisson observation model, but replaces these independent time-bin-wise rate estimates with a smooth stimulus-specific Gaussian process model for the log firing rate.

      We have also clarified how the Goris-style models were implemented. All Goris-style results reported in the main model comparisons use versions with a GP prior on the stimulus drive. In these versions, the stimulus-dependent firing rates are set to the smooth firing-rate estimates obtained from the PoissonGP model, and the gain parameters are then fit under either the independent-gain or constant-gain assumptions. We used these GP-smoothed versions as stronger baselines. In Figure 5, Figure Supplement 2, we additionally compare these models to Goris-style variants without the GP prior on the stimulus drive, in which the stimulus-drive parameters are estimated directly under the corresponding Goris-style likelihood. This comparison shows that adding a GP smoothness prior to the stimulus drive substantially improves held-out model fit. Together, these analyses clarify that the GP-smoothed stimulus drive improves the Goris-style baselines, while the continuous-time gain process in CMP provides an additional improvement by capturing temporally structured trial-to-trial variability.

      To address this comment, we clarified the definitions of the Baseline Poisson, Poisson-GP, and Goris-style models (Pages 8-9: Lines 255-259, 263-266, 273-277), and revised the text (Page 11: Lines 319-326) describing Figure 4 - figure Supplement 2 to make explicit how this existing comparison isolates the effect of the GP prior on the stimulus drive.

      Reviewer #2 (Public Review):

      Summary:

      Neurons have varied responses to external stimuli that cannot be explained by naive Poisson models. Previous work has quantified and partitioned higher−than−Poisson variability in the brain into different components. The authors improve on these methods to infer how both the stimulus drive and internal gain dynamics impact neuronal variability continuously in time. The clean and well−reasoned model is rigorously developed and then applied to neural data across the visual hierarchy. This lends new insights into how variability is partitioned, agreeing with and extending previous work on how that variability changes from early visual areas (LGN, V1) through to higher, motion−sensitive areas (area MT). Another key contribution is that this partitioning can be fully addressed as a continuous−time process, which allows for the dissection of how the timescale of fluctuations in these two components changesacross the brain’s processing arc.

      Strengths:

      (1) The model is cleanly derived and thoroughly documented, including usable code shared in a GitHub repo. This makes the method immediately portable to other neural systems.

      (2) This is a clear and well−presented piece of work. The figures and writing are clear and understandable, and all pieces of the derivations are included in the main text and supplementary information.

      (3) Comparisons to other models, particularly the one from Goris et al., 2014 shows how this Continuous Modulated Poisson (CMP) model outperforms previous work.

      (4) New insights about how variability partitioning changes across the visual stream from LGN to MT are revealed, including how the gain fluctuates on longer timescales in higher visual areas. Another key result about the anticorrelation between the variance in stimulus drive and gain fluctuations comports with theories about how neurons maintain efficient, reliable encoding.

      (5) In addition to the results reported here, this work will serve as an excellent tutorial for students and postdocs first delving into the sources of variability in the brain.

      We sincerely thank the reviewer for the thoughtful and positive assessment of our work. We are pleased that the model development, empirical analyses, and presentation were viewed as clear, rigorous, and useful for the broader neuroscience community. We also appreciate the reviewer’s recognition that the continuous-time formulation meaningfully extends prior variability-partitioning approaches by allowing stimulus drive and internal gain dynamics to be characterized across temporal scales. The reviewer’s comments helped us further clarify the positioning of the work, expand the Discussion of model scope and limitations, and better articulate future extensions. Below, we address the specific suggestions raised by the reviewer and describe the revisions made in the manuscript.

      Weaknesses:

      The work is somewhat incremental, building on previous studies of the partitioning of variability in the brain, but it provides important new extensions, as noted above.

      Regarding the comment on incremental contribution, we agree that our framework builds directly on previous variability-partitioning approaches, especially the modulated Poisson framework of Goris et al. However, the main goal of this work is to move this class of models from a count-based formulation to a continuous-time spike-train framework. This extension is important because it allows us to model gain as a temporally structured latent process, characterize how variability depends on the timescale over which spikes are counted, and infer the temporal covariance structure of stimulus-independent fluctuations. In addition, the CMP framework provides analytic expressions for the Fano factor as a function of bin size, introduces the EPL covariance function for slowly decaying gain dynamics, and enables direct comparisons of gain amplitude and timescale across visual areas. In the revised manuscript, we have clarified this positioning and emphasized how CMP extends prior variability-partitioning models while preserving their interpretability.

      To address this comment, we revised the Introduction (Pages 3-4: Lines 108-111 and Lines 118-121) and Discussion (Page 13: Lines 425-429) to clarify better how CMP builds on prior variability-partitioning models while extending them to continuous-time spike-train data.

      The only major gap I would suggest addressing in the Discussion is the observation of sub−Poisson variability in the brain. It seems clear that this model can extend to sub− Poisson variability and its partitioning and perhaps even show how that varies in real time, with an animal’s attentional state. That is, of course, beyond the scope of the current work, but could be mentioned in the Discussion.

      We thank the reviewer for this suggestion. We agree that sub-Poisson variability is an important phenomenon observed in neural data. Because the CMP model uses a conditionally Poisson observation model with stochastic gain modulation, it naturally captures Poisson and super-Poisson variability but does not generate sub-Poisson spike count statistics in its current form. In the revised manuscript, we have clarified this limitation in the Discussion and outlined possible extensions that could address sub-Poisson variability, including spike-history terms, renewal-process likelihoods, and alternative count distributions [Truccolo et al., 2005, Paninski et al., 2007, Aghamohammadi et al., 2024]. We also note that such extensions could allow future models to examine how sub-Poisson and super-Poisson components vary with behavioral state, attention, or arousal.

      To address this comment, we added a Discussion paragraph describing the current model’s limitation for sub-Poisson variability and possible extensions to capture it (Page 14: Lines 460-471).

      References

      Stephen Keeley, Mikio Aoi, Yiyi Yu, Spencer Smith, and Jonathan W Pillow. Identifying signal and noise structure in neural population activity with gaussian process factor models. In H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 13795–13805. Curran Associates, Inc., 2020.

      Robbe L. T. Goris, Corey M. Ziemba, J. Anthony Movshon, and Eero P. Simoncelli. Slow gain fluctuations limit benefits of temporal integration in visual cortex. Journal of Vision, 18(8):8–8, 08 2018. ISSN 1534-7362. doi: 10.1167/18.8.8. URL https://doi.org/10.1167/18.8.8.

      Olivier J H´enaff, Zoe M Boundy-Singer, Kristof Meding, Corey M Ziemba, and Robbe L T Goris. Representation of visual uncertainty through neural gain variability. Nat. Commun., 11(1):2513, May 2020.

      Mark M Churchland, Byron M Yu, John P Cunningham, Leo P Sugrue, Marlene R Cohen, Greg S Corrado, William T Newsome, Andrew M Clark, Paymon Hosseini, Benjamin B Scott, David C Bradley, Matthew A Smith, Adam Kohn, J Anthony Movshon, Katherine M Armstrong, Tirin Moore, Steve W Chang, Lawrence H Snyder, Stephen G Lisberger, Nicholas J Priebe, Ian M Finn, David Ferster, Stephen I Ryu, Gopal Santhanam, Maneesh Sahani, and Krishna V Shenoy. Stimulus onset quenches neural variability: a widespread cortical phenomenon. Nat. Neurosci., 13(3):369–378, March 2010.

      Robbe L T Goris, Ruben Coen-Cagli, Kenneth D Miller, Nicholas J Priebe, and M´at´e Lengyel. Response sub-additivity and variability quenching in visual cortex. Nat. Rev. Neurosci., 25(4):237–252, April 2024.

      Wilson Truccolo, Uri T. Eden, Matthew R. Fellows, John P. Donoghue, and Emery N. Brown. A point process framework for relating neural spiking activity to spiking history, neural ensemble, and extrinsic covariate effects. Journal of Neurophysiology, 93(2):1074–1089, 2005. doi: 10.1152/jn.00697.2004. URL https://doi.org/10.1152/jn.00697.2004. PMID: 15356183.

      Liam Paninski, Jonathan Pillow, and Jeremy Lewi. Statistical models for neural encoding, decoding, and optimal stimulus design. In Paul Cisek, Trevor Drew, and John F. Kalaska, editors, Computational Neuroscience: Theoretical Insights into Brain Function, volume 165 of Progress in Brain Research, pages 493–507. Elsevier, 2007. doi: https://doi.org/10.1016/S0079-6123(06)65031-0. URL https://www.sciencedirect.com/science/article/pii/S0079612306650310.

      Cina Aghamohammadi, Chandramouli Chandrasekaran, and Tatiana A. Engel. A doubly stochastic renewal framework for partitioning spiking variability. bioRxiv, 2024.

    1. Author response:

      Reviewer #1:

      We thank the reviewer for their comments. They raised an issue with the correlational nature of the analysis, suggesting alternative explanations, including stable individual differences in response styles, latent third variables, or measurement properties of the confidence scales. We agree with the reviewer’s comments that a latent third variable may confound our findings and higher-order metacognitive beliefs (which we did not assess) are a possible candidate. We will broaden our discussion to include other plausible confounds that may jointly relate to depression and confidence dynamics. Regarding stable individual differences in response styles, our analysis decomposed confidence into within and between-person components and standardised the within-person component by each participant’s own variability. This allowed the interaction and moderated mediation analyses to separate within-person associations from stable between-person differences (e.g., range of confidence scale used; Epskamp et al., 2018). With regards to measurement properties, we will conduct additional analyses investigating whether trait depression is associated with altered response patterns (e.g., non-linear or more variable mapping of latent evidence into confidence reports).

      The reviewer also noted that the temporal resolution we selected may not be optimal to measure co-fluctuation of depression and confidence. We agree this remains a challenge for the depression research (Jamalabadi et al., 2025; Tamm et al., 2024), and metacognitive confidence research (da Fonseca et al., 2023; Wright et al., 2024). We chose the interval as it allowed us to balance retention and data quality across an 8-week study (Eisele et al., 2022) and address a gap in the literature – frequent but shorter (~1-week) EMA-based assessments of existing metacognitive confidence studies (da Fonseca et al., 2023; Wright et al., 2024) precludes insights into the dynamics of mood and metacognitive confidence across longer timescales. We have explored one-day and half-day lagged associations but did not find significant effects across most symptoms for both local and global confidence. This was omitted from the original manuscript as the analyses could not symmetrically control for local/global confidence (as these were only assessed bi-daily) but will be included in the revised manuscript’s supplement.

      The reviewer noted that the temporal ordering of the moderated mediation analysis cannot be conclusively established. We agree that there is a lack of research investigating alternative orderings. We selected the local-to-global ordering because prior work has similarly modelled and demonstrated global confidence estimates as a cumulative integration of local confidence estimates across the block (Cavalan et al., 2023; Katyal et al., 2025; Lee et al., 2021; Rouault et al., 2019, 2022). Nevertheless, we modelled an alternative process (i.e. whether global confidence interacted with trait-level depression in predicting subsequent local confidence) but did not find a significant frequentist interaction effect. We will elaborate on our rationale for the current ordering and include analyses of this alternative process in our revised manuscript and supplement.

      Finally, we agree with the reviewer’s comments that the sample constraints generalisability. The revised manuscript will more clearly discuss the generalisability of our findings to clinical populations and the implications of self-selection and limited within-person variability in depressive symptoms for interpretation and future work. We will also tone down the causal language in the manuscript.

      Reviewer #2:

      We thank the reviewer for their comments. We agree with the reviewer that our strongest effects live at the trait level, whereas our cross-lagged findings provided insufficient evidence for depression driving underconfidence or vice versa (at least with a two-day lag). However, our moderated mediation finding centres on a different angle - trait level depression moderating how within-person confidence is integrated into more global beliefs. Nevertheless, we agree that the findings provide a plausible account for why underconfidence (and possibly metacognitive beliefs) remain persistent in depression, but not how either underconfidence or depression arose in the first place. We will make this clearer in our revised manuscript.

      Consistent with reviewer #1’s comments, we agree that the non-clinical nature of sample does constrain generalisability and inference. We will conduct a sensitivity analysis with a larger sample (more lenient inclusion criteria) and discuss its implications in more detail in our revised manuscript. We also agree with the reviewer’s comments that our findings do not preclude the possibility of temporal precedence and will further clarify in our revised manuscript with reference to shorter time lags. The reviewer noted that our use of terminologies (e.g., “depression”, “depressive symptoms”, “depressive mood”) was not clearly defined. Our revised manuscript will make this point clearer whilst also making the usage of terminologies consistent.

      References

      Cavalan, Q., Vergnaud, J.-C., & de Gardelle, V. (2023). From local to global estimations of confidence in perceptual decisions. Journal of Experimental Psychology: General, 152(9), 2544–2558. https://doi.org/10.1037/xge0001411

      da Fonseca, M., Maffei, G., Moreno-Bote, R., & Hyafil, A. (2023). Mood and implicit confidence independently fluctuate at different time scales. Cognitive, Affective, & Behavioral Neuroscience, 23(1), 142–161. https://doi.org/10.3758/s13415-022-01038-4

      Eisele, G., Vachon, H., Lafit, G., Kuppens, P., Houben, M., Myin-Germeys, I., & Viechtbauer, W. (2022). The Effects of Sampling Frequency and Questionnaire Length on Perceived Burden, Compliance, and Careless Responding in Experience Sampling Data in a Student Population. Assessment, 29(2), 136–151. https://doi.org/10.1177/1073191120957102

      Epskamp, S., Waldorp, L. J., Mõttus, R., & Borsboom, D. (2018). The Gaussian Graphical Model in Cross-Sectional and Time-Series Data. Multivariate Behavioral Research, 53(4), 453–480. https://doi.org/10.1080/00273171.2018.1454823

      Jamalabadi, H., Koosha, T. A., Stocker, E., Jansen, A., Ebner-Priemer, U. W., Proppert, R. K. K., Rieble, C. L., Tutunji, R., & Fried, E. I. (2025). Optimizing the frequency of ecological momentary assessments using signal processing. Psychological Medicine, 55, e358. https://doi.org/10.1017/S003329172510264X

      Katyal, S., Huys, Q. J., Dolan, R. J., & Fleming, S. M. (2025). Distorted learning from local metacognition supports transdiagnostic underconfidence. Nature Communications, 16(1), 1854. https://doi.org/10.1038/s41467-025-57040-0

      Lee, A. L. F., de Gardelle, V., & Mamassian, P. (2021). Global visual confidence. Psychonomic Bulletin & Review, 28(4), 1233–1242. https://doi.org/10.3758/s13423-020-01869-7

      Rouault, M., Dayan, P., & Fleming, S. M. (2019). Forming global estimates of self-performance from local confidence. Nature Communications, 10(1), 1141. https://doi.org/10.1038/s41467-019-09075-3

      Rouault, M., Will, G.-J., Fleming, S. M., & Dolan, R. J. (2022). Low self-esteem and the formation of global self-performance estimates in emerging adulthood. Translational Psychiatry, 12(1), 1–10. https://doi.org/10.1038/s41398-022-02031-8

      Tamm, J., Takano, K., Just, L., Ehring, T., Rosenkranz, T., & Kopf-Beck, J. (2024). Ecological Momentary Assessment versus Weekly Questionnaire Assessment of Change in Depression. Depression and Anxiety, 2024, 9191823. https://doi.org/10.1155/2024/9191823

      Wright, A. C., Palmer-Cooper, E., Cella, M., McGuire, N., Montagnese, M., Dlugunovych, V., Liu, C.-W. J., Wykes, T., & Cather, C. (2024). Experiencing hallucinations in daily life: The role of metacognition. Schizophrenia Research, Hallucinations: Neurobiology and Patient Experience, 265, 74–82. https://doi.org/10.1016/j.schres.2022.12.023

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting and well-written manuscript in which the authors set out to answer a simple, old question with a modern toolkit: where in crab evolution did sideways walking arise, how often has it been lost or regained, and is it plausibly linked to the ecological and taxonomic success of true crabs. To do this, they record locomotion from 50 live species, convert each species' movements into a quantitative index that compares forward versus sideways bouts, and then map the resulting states onto a recent crab phylogeny to infer the most likely evolutionary history of locomotor direction.

      We thank the reviewer for this positive summary of the study and for recognizing the value of our comparative behavioral dataset and phylogenetic approach.

      Strengths:

      The strongest part of the study is the dataset itself. Comparable behavioral measurements across dozens of crab species are rare. The authors have done the field and husbandry work needed to make this possible. The overall pattern they recover, that most true crabs are strongly biased toward sideways movement (while a smaller set of lineages move predominantly forward), is interesting and likely to be useful to others. The phylogenetic mapping is also a reasonable way to address the "how many times" question (although this is peripheral to my expertise). The manuscript makes a convincing case that sideways locomotion is not simply a trivial byproduct of a crab-like body plan.

      We appreciate the reviewer’s recognition of the dataset and the overall value of the study. We have revised the manuscript to make the conclusions more robust and better aligned with the strength of the evidence.

      (1) Where I am less convinced is in how strongly the authors describe the discreteness of the behavioral categories and the absence of intermediates. The manuscript states that the Forward-Sideways Index shows a clear separation between two locomotor types with little evidence for intermediates, and it cites a statistical test rejecting a single peak in the distribution. However, the histogram in Figure 3 appears structured within each labeled category, with subclusters inside both the forward and sideways groups rather than a single tight peak per group. This matters because the index is built by first placing each movement bout into "forward" versus "sideways" bins using a fixed angle boundary and then collapsing the result into a single ratio. That approach is simple and transparent enough, but it can also hide mixed strategies. For example, a species that produces substantial amounts of both forward and sideways walking can still end up with a strongly positive or negative index, and therefore be classified as a pure "type," even though the underlying behavior is mixed. In that context, rejecting a single peak in the across-species distribution does not, by itself, justify the stronger claim that intermediates are rare or absent.

      Related to this, a key methodological choice is the use of 60 degrees as the cutoff between forward and sideways bouts. This boundary may be reasonable as a convention, but the paper does not explain why it is the right place to draw the line, and there is a plausible biological concern that a fixed angular cutoff does not mean the same thing across taxa.

      Crabs vary in body shape and in how the legs are arranged around the body. In my own comparative work, for example, some species show an elliptical stance pattern elongated along the preferred direction of travel, while others show a more circular leg arrangement, and the latter can express more mixed forward and sideways behavior. When limb arrangement and body geometry differ across species, the same measured angle can correspond to different underlying mechanics and different functional "degree of sidewaysness." The practical implication is that the reported binary separation may partly reflect the imposed classification rule, rather than a sharp biological divide.

      We thank the reviewer for this important point. We agree that the across-species distribution of FSI values alone does not justify a strong statement that intermediate or mixed locomotor tendencies are absent. We also agree that reducing continuous bout-angle distributions to a single index could potentially obscure mixed directional strategies. We have therefore revised the manuscript to avoid implying a strict absence of intermediates and have added an additional analysis of the underlying continuous angle distributions (Abstract, lines 27-29; Results, lines 191-208; Table S2).

      Specifically, we fitted one- and two-component mixture models to the continuous bout-angle distributions of each taxon and examined the supported number of components, peak locations, and mixture weights (Results, lines 196-204; Table S2). This analysis showed that 14 taxa were best described by a one-component model, whereas 36 taxa were best described by a two-component model. Importantly, among the 36 taxa best described by a two-component model, 33 had a dominant component explaining at least 70% of the distribution, whereas only three taxa showed relatively balanced two-component distributions. Thus, although some taxa do show mixed directional tendencies, most taxa are dominated by a primary directional component rather than showing an even mixture of forward and sideways locomotion.

      We also clarified the rationale for the 60° threshold used in the FSI calculation (Methods, lines 125-130). This threshold was not intended to represent a taxon-specific biological boundary between forward and sideways locomotion. Rather, it was used to divide the 360° space into three equal directional sectors: forward, sideways, and backward. This equal partitioning provides a consistent reference under a null expectation of uniformly distributed movement directions.

      To further assess whether our classification depended on the original FSI-based classification, we performed an additional data-driven check based on continuous bout-angle distributions (Results, lines 204-208; Fig. S2). We extracted the dominant peak location from each taxon’s continuous bout-angle distribution (Table S2) and estimated a boundary from the distribution of these dominant peak locations using a Gaussian mixture model. This yielded a data-informed cutoff of approximately 49.4°. The resulting peak-based classification was identical to the original FSI-based classification, with 15 forward-moving and 35 sideways-moving taxa. Thus, no taxon changed category under this independent classification approach.

      (2) Another limitation that affects interpretation is the decision to use one individual per species. I understand the logistics, and for some questions, a single representative individual can be a reasonable first pass. But it is not strong support for negative claims about intermediates, especially in a group where individuals can change substantially with growth and allometry. Crabs can grow dramatically, often with pronounced allometric shifts in limb proportions that can alter the center of mass location. Size alone can alter the kinematics and choice of locomotor behaviors in crustaceans. In species where appendage proportions change with size, or where certain legs become disproportionately large (or calcified), it is plausible that locomotor direction and the distribution of movement angles shift across ontogeny. That makes it hard to treat a single individual as a complete description of a species-level strategy, particularly for species that fall closer to the boundary between categories.

      We thank the reviewer for raising this important limitation. We agree that using one representative individual per species cannot capture the full range of within-species variation, including ontogenetic, size-dependent, or allometric changes in locomotor behavior. We also agree that this limitation is particularly relevant to strong claims about the absence of intermediates.

      As described in our response to Comment #1, we have therefore toned down statements implying a strict absence of intermediates and added analyses of the underlying continuous angle distributions (Abstract, lines 27-29; Results, lines 191-208; Table S2). These additional analyses showed that some taxa do exhibit mixed directional tendencies, although most taxa were dominated by a primary directional component.

      We have also revised the manuscript to clarify the scope of our conclusions. Specifically, we now state that our single-individual sampling design does not capture possible ontogenetic, size-dependent, or allometric variation within species (Methods, lines 105-107). We also clarify that our conclusions are intended to identify broad interspecific patterns in the predominant direction of locomotion across major brachyuran lineages, rather than to describe the full range of locomotor variation within each species (Methods, lines 107-108). Thus, we no longer treat a single individual as providing a complete description of species-level behavioral variation, but instead use it as a standardized representative observation for broad comparative and phylogenetic analyses.

      In sum, this is a valuable and useful behavioral comparative study with a dataset that many in the field will appreciate. The main conclusions about the likely evolutionary placement of sideways walking are plausible, but several of the stronger claims about discrete locomotor types, the absence of intermediates, and the relationship to diversification would be more convincing if the analysis were less dependent on a fixed angular cutoff and on single individuals per species, or if the manuscript framed those points more cautiously so the conclusions track the strength of the evidence.

      We thank the reviewer for this constructive summary and for recognizing the value of our behavioral comparative dataset. We have addressed these concerns in detail in our responses above and revised the manuscript to make the main claims better aligned with the strength of the evidence.

      Reviewer #2 (Public review):

      Summary:

      The current work investigates the evolution of sideward locomotion in Brachyura in light of a single evolutionary origin. To this end, the authors first analysed the mode of locomotion in 50 crab species and observed mutually exclusive presence of sideways vs. forward movement. The phylogenetic analysis confirmed that there is indeed a single evolutionary origin for sideways movement, which was sometimes followed by several reversions to forward locomotion. This way, authors demonstrate how locomotor movement modes shape evolutionary diversification in animals by showing that species richness is much higher in side-ways-moving crabs than in the nearest groups. This is an interesting work that integrates behavioural analysis and phylogenetic relations, capitalising largely on crabs. I have a few suggestions and questions.

      We thank the reviewer for the positive assessment of the study and for recognizing the value of integrating behavioral analysis with phylogenetic relationships. We address the specific suggestions and questions below.

      (1) Firstly, I think the paper spends too much time on a straightforward analysis of the mode of locomotion.

      We agree that the final classification of taxa into predominantly forward- and sideways-moving groups is conceptually simple. However, because our study compares locomotor behavior across a broad range of crab taxa and then uses these behavioral data for phylogenetic reconstruction, we considered it important to describe the behavioral quantification in a transparent and reproducible way. The purpose of this section is therefore not to make a simple endpoint unnecessarily complex, but to show how discrete locomotor states were derived from raw trajectory data using a standardized procedure. For this reason, we retained the current analytical description.

      (2) I was also wondering whether the phylogenetic analysis could be simply achieved by maximising an objective function in which the modes of movement are inversely coded for two putative groups, with all values calculated at all possible nodes.

      The proposed objective-function approach may be useful for identifying a node that best separates two predefined locomotor groups. However, in the present study, we aimed not only to locate a possible boundary between forward- and sideways-moving lineages, but also to reconstruct the evolutionary history of locomotor transitions under an explicit phylogenetic model.

      For this reason, we used standard ancestral state reconstruction and stochastic character mapping rather than maximizing an ad hoc objective function across possible nodes. This approach allowed us to compare alternative transition-rate models (ER and ARD), estimate uncertainty in ancestral states at internal nodes, and quantify the posterior distribution of gains and reversals. We therefore retained the current phylogenetic framework, as it provides a model-based and more informative reconstruction of locomotor evolution across true crabs.

      (3) Unfortunately, I find that the authors did not sufficiently discuss differences in the ecological niches of species with forward vs. sideways locomotion modes (including challenges of locomotion and substrate).

      Likewise, what are the anatomic correlates of forward vs. sideways locomotion? For instance, how are the advantages assumed for sideways movement associated with a flattened body? Is it possible that the mode of motion is secondary to flattened/narrow body structure, which basically limits the distance between legs and thus makes the forward movement difficult - under this logic, the mode of movement would be a secondary phenomenon to body shape traits. How can one differentiate between this alternative and the one that puts the mode of movement in the centre of the story? On a related note, how do different modes of movement relate to the ability to fit into tight spaces - how does it relate to differences in leg joints?

      Is it possible that the sideways movement maximises the scanned visual field per unit time/displacement, which may be beneficial for mostly forward-moving predators?

      We thank the reviewer for this helpful comment. We agree that the previous version did not sufficiently address the possible relationship between locomotor mode and body shape, especially the alternative explanation that sideways locomotion may be secondary to carapace flattening. In response, we added a new morphological analysis using two carapace shape indices: relative carapace length (CL/CW) and relative carapace depth (CD/CS) (Methods, lines 178–185). In the revised Results, we report that relative carapace length differed significantly between forward- and sideways-moving taxa (phylogenetically informed ANOVA: F = 26.90, p < 0.001), whereas relative carapace depth did not differ significantly between the two groups (F = 1.18, p = 0.403) (Results, lines 209–214; Fig. S3). We also added this interpretation to the Discussion, noting that locomotor mode is associated with some aspects of carapace shape but is not explained by simple carapace flattening alone (Discussion, lines 318–325).

      We also revised the Discussion to address the reviewer’s suggestions about possible functional advantages of sideways locomotion beyond rapid bidirectional escape. Specifically, we now mention that other possible advantages may include movement through confined spaces and visual-field sampling during locomotion (Discussion, lines 310-318).

      Finally, we retained the existing discussion of ecological specializations in forward-moving lineages, including coordinated collective movement in soldier crabs, decoration and concealment in majoid crabs, and life inside confined host spaces in pea crabs. This discussion supports the broader point that the adaptive value of sideways locomotion may depend on ecological context.

      (4) It is really difficult to decipher the information contained in the nodes (circles) in the printed black-and-white version of the manuscript.

      We have changed the color scheme and strengthened the outlines of the node pie charts so that the ancestral-state probabilities can be more easily distinguished (Fig. 5). We also applied the same revised color scheme to Figure 3 and Figure S4 for consistency across the manuscript.

      (5) Briefly, although I find the study interesting, the presented complexity may not be necessary given the endpoints; it can be achieved much more simply. Furthermore, the degree to which the conceptual analysis of different modes of locomotion was exercised was limited. The general approach may serve as a good model for the evolutionary analysis of other traits. The demonstration of traceability of the relations in question is a major contribution of the work.

      We thank the reviewer for this constructive summary and for recognizing the broader value of our approach. We have addressed the methodological and conceptual points raised here in our responses to the specific comments above.

      Strengths:

      The research question and the novel combination of different data types.

      We thank the reviewer for highlighting the research question and the novel combination of different data types as strengths of the study. We have revised the manuscript to further strengthen this integrative framework.

      Weaknesses:

      The complexity of the methods used, along with a limited discussion of the potential dynamics that may underlie the evolution of the sideways movement mode.

      We have addressed these concerns in our responses to the specific comments above, particularly by clarifying the rationale for the behavioral quantification and expanding the discussion of morphology, ecological context, and functional hypotheses.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Unbiased analysis of angle distribution. The authors already extract continuous bout angles prior to binning. I recommend using these distributions directly to assess modality at the species level (e.g., unimodal vs bimodal, peak locations, mixture weights) before collapsing behavior into the Forward-Sideways Index. Even a simple circular density estimate or mixture model would clarify whether species classified as "forward" or "sideways" are behaviorally pure or mixed, and would provide a quantitative basis for claims about intermediacy.

      We added mixture-model analyses of the continuous bout-angle distributions, including modality, peak locations, and mixture weights (Results, lines 196-204; Table S2).

      (2) Justify or stress-test the 60° cutoff. The manuscript should either provide a clear biological or data-driven justification for using 60° as the boundary between forward and sideways bouts, or demonstrate that the main conclusions are robust to reasonable alternative cutoffs (e.g., 45°, 75°). A brief sensitivity analysis in the supplement would be sufficient and would greatly strengthen confidence in the classification. Alternatively (my preference) would be to let the data inform the cutoff.

      We clarified the rationale for the 60° sector definition used to calculate FSI (Methods, lines 125-130; Fig. 2) and added an independent data-driven boundary analysis based on dominant peak locations, which yielded the same forward/sideways classification (Results, lines 204-208; Fig. S2).

      (3) Sampling justification. I recommend explicitly acknowledging that sampling a single individual per species limits the ability to detect ontogenetic, size-dependent, or allometric variation in locomotor strategy. If feasible, adding even limited replication across size classes or individuals for a small subset of taxa (particularly those near the classification boundary) would substantially strengthen the conclusions; otherwise, the manuscript should more clearly delimit which claims do and do not rely on the assumption of within-species invariance.

      We clarified that our conclusions concern broad interspecific patterns of predominant locomotor direction, rather than the full range of within-species variation (Methods, lines 102-108).

      (4) I would suggest toning down or reframing statements about "no intermediates". If additional analyses are not added, I recommend revising statements that imply a strict absence of intermediates to language that reflects what is directly shown (e.g., bimodality in an index derived from binned data). This would better align the claims with the current evidence.

      We revised the manuscript to avoid implying a strict absence of intermediates and now acknowledge that some taxa show mixed directional tendencies (Abstract, lines 27-29; Results, lines 191-208).

      (5) Framing and claims about diversification. The discussion of sideways locomotion as a key innovation would benefit from clearer separation between observed correlations and causal inference. If trait-dependent diversification analyses are not added, I suggest consistently framing this section as a hypothesis supported by comparative patterns rather than a demonstrated mechanism.

      We revised the Discussion to more clearly frame sideways locomotion as a possible key innovation associated with diversification, rather than as a demonstrated causal mechanism (Abstract, lines 31-34; Discussion, lines 294-309).

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors perform an analysis of the relationship between the size of an LMM and the predictive performance of an ECoG encoding model made using the representations from that LMM. They find a logarithmic relationship between model size and prediction performance, consistent with previous findings in fMRI. They additionally observe that as the model size increases, the location of the "peak" encoding performance typically moves further back into the model in terms of percent layer depth, an interesting result worthy of further analysis into these representations.

      Strengths:

      The evidence is quite convincing, consistent across model families, and complementary to other work in this field. This sort of analysis for ECoG is needed and supports the decade-long enduring trend of the "virtuous cycle" between neuroscience and AI research, where more powerful AI models have consistently yielded more effective predictions of responses in the brain. The lag analysis showing that optimal lags do not change with model size is a nice result using the higher temporal resolution of ECoG compared to other methods like fMRI.

      We thank the reviewer for their thoughtful assessment! We agree that the “virtuous cycle” between neuroscience and AI research has been, and will continue to be, a driving force in advancing our understanding of brain function through more powerful predictive models. We are especially pleased that the reviewer appreciated the lag analysis, as we view this as a valuable complement to the existing fMRI work.

      Weaknesses:

      I would have liked to have seen the data scaling trends explored a bit too, as this is somewhat analogous to the main scaling results. While better performance with more data might be unsurprising, showing good data scaling would be a strong and useful justification for additional data collection in the field, especially given the extremely limited amount of existing language ECoG data. I realize that the data here is somewhat limited (only 30 minutes per subject), but authors could still in principle train models on subsets of this data.

      We thank the reviewer for their valuable suggestion. For the revised manuscript, we performed a new analysis where we trained encoding models using subsets of the data (randomly sampling contiguous chunks of 50%, 25%, and 10% of all words in each of the training folds) and tested these models on all words in the test fold. As expected, we found that encoding performance increases as the training dataset size increases, suggesting that model performance scales with data quantity even within the constraints of our relatively small dataset. This result reinforces the importance of collecting dense ECoG data. We have added the following text to our Results section: “We also built encoding models using subsets of the data and found that encoding performance increases as the volume of training data increases (Fig. S6)” and included the results as a supplementary figure 6 in the revised manuscript.

      Separately, it would be nice to have better justification of some of these trends, in particular the peak layerwise encoding performance trend and the overall upside-down U-trend of encoding performance across layers more generally. There is clearly something very fundamental going on here, about the nature of abstraction patterns in LLMs and in the brain, and this result points to that. I don't see the lack of justification here as a critical issue, but the paper would certainly be better with some theoretical explanation for why this might be the case.

      We thank the reviewer for this insightful comment. The general inverted U-shaped trend of encoding performance across layers has been a frequently observed phenomenon in studies comparing LLM representations to brain activity (Goldstein, Ham, et al., 2025; Schrimpf et al., 2021). A potential explanation is the existence of a “two-phase abstraction process” within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the initial layers, models begin by processing relatively low-level input features. As layers get deeper, representations become increasingly abstract and richly contextualized in semantic features relevant for understanding language. These intermediate layers often show the highest correlation with brain activity in language areas, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. These observations suggest that it is primarily the abstractive, contextual features developed in the intermediate layers of LLMs that drive their alignment with brain activity. As models become more potent at prediction, their most predictive layers (for the LLM’s natural language task) and their most generalizable layers (for brain activity) can diverge.

      A key finding in our study is that the initial processing phase does not scale and take up more layers as models scale up in size and layers. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Consequently, the prediction phase may begin relatively earlier in these larger models, and the later layers could develop highly specialized representations that are increasingly divergent from the more general linguistic processing captured in brain activity. For example, these layers may specialize in capturing very specific patterns of language (thus lowering their perplexity) that do not actually occur often or at all in our naturalistic dataset. It is also possible that the later layers of larger models are overall underutilized and do not contribute as much to linguistic processing and next-word prediction (Csordás et al., 2025).

      We have added the following text to our Discussion section:

      “The inverted U-shaped trend of encoding performance commonly found in previous research is likely due to a "two-phase abstraction process" within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the early and intermediate layers of the model, a composition phase occurs, where low-level input features become increasingly abstract and contextualized. The intermediate layers of the model show the highest correlation with brain activity, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers of the model, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. Our results indicate that the initial composition phase does not take up more layers as models scale up in size. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Thus, as LLMs increase in size, the later layers of the model may contain representations that are increasingly divergent from the more general linguistic processing captured in brain activity. It is also possible that the later layers of larger models are overall underutilized and may not significantly contribute to benchmark performances during inference (Csordás et al., 2025; Fan et al., 2024; Gromov et al., 2024).”

      Lastly, I would have wanted to see a similar analysis here done for audio encoding models using Whisper or WavLM as this is the modality where you might see real differences between ECoG and other slower scanning approaches. Again, I do not see this omission as a fundamental issue, but it does seem like the sort of analysis for which the higher temporal resolution of ECoG might grant some deeper insight.

      We appreciate this suggestion. In a separate project, we focused on multimodal audio-to-speech-to-language large language models (LLMs), building encoding models using Whisper embeddings (from both the encoder and decoder stacks) to predict electrocorticographic (ECoG) signals during naturalistic conversations (Goldstein, Wang, et al., 2025). The higher temporal resolution of ECoG enables us to trace the temporal flow of information from the superior temporal gyrus (STG) and somatomotor areas (SM) to the inferior frontal gyrus (IFG) during speech comprehension. Conversely, during speech production, encoding in IFG peaked significantly earlier than in the STG and SM. We agree that scaling encoding models using multimodal approaches and our ECoG conversation datasets could yield valuable insights, and we look forward to exploring this in future work. However, we feel that the added complexity of multimodal encoding models falls beyond the scope of this paper.

      We have modified the following text to our Discussion section:

      “Since we exclusively employ textual LLMs, which lack inherent temporal information due to their discrete token-based nature, future studies utilizing multimodal LLMs integrating continuous audio or video streams, like Whisper or WavLM may better unravel the relationship between model size and temporal dynamic representations in LLMs (Goldstein, Wang, et al., 2025; Millet et al., 2023; Vaidya et al., 2022).”

      Reviewer #2 (Public review):

      Summary:

      This paper investigates whether large language models (LLMs) of increasing size more accurately align with brain activity during naturalistic language comprehension. The authors extracted word embeddings from LLMs for each word in a 30-minute story and regressed them against electrocorticography (ECoG) activity time-locked to each word as participants listened to the story. The findings reveal that larger LLMs more effectively predict ECoG activity, reflecting the scaling laws observed in other natural language processing tasks.

      Strengths:

      (1) The study compared model activity with ECoG recordings, which offer much better temporal resolution than other neuroimaging methods, allowing for the examination of model encoding performance across various lags relative to word onset.

      (2) The range of LLMs tested is comprehensive, spanning from 82 million to 70 billion parameters. This serves as a valuable reference for researchers selecting LLMs for brain encoding and decoding studies.

      (3) The regression methods used are well-established in prior research, and the results demonstrate a convincing scaling law for the brain encoding ability of LLMs. The consistency of these results after PCA dimensionality reduction further supports the claim.

      We thank the reviewer for their thoughtful and positive feedback.

      Weaknesses:

      (1) Some claims of the paper are less convincing. The authors suggested that "scaling could be a property that the human brain, similar to LLMs, can utilize to enhance performance", however, many other animals have brains with more neurons than the human brain, making it unlikely that simple scaling alone leads to better language performance.

      We thank the reviewer for this insightful comment. We agree that simply having more neurons does not automatically confer more complex or human-like cognitive or linguistic capabilities. This suggestion deserves a more nuanced treatment than we had included in the original manuscript.

      Research in comparative neuroscience has argued that human cognitive abilities emerge from scaling up the primate brain (Herculano-Houzel, 2012). However, the critical aspect is not merely the number of neurons, but how these neurons contribute to computational power within a specific evolutionary and cultural context. The uniqueness of human cognition has been argued to result from a global adaptation for increased information processing capacity (Cantlon & Piantadosi, 2024). Moreover, the language network in humans is likely grounded in the evolution of particular structural networks in the primate brain (Friederici & Becker, 2025). This suggests that the way brain regions are connected and the expansion of certain pathways are critical, not just the overall scale. Furthermore, the specialized structure must be tuned by its learning environment and training data. For example, both humans and LLMs learn from language data generated by other humans, which reflects world knowledge that has accumulated over many generations.

      We have modified the following text in the Introduction:

      “Research in comparative neuroscience has suggested that uniquely human cognitive abilities emerge from scaling up the primate brain (Herculano-Houzel, 2012).”

      We also added a caveat to the Discussion on this point:

      “As in the human brain, while scaling alone may yield emergent cognitive abilities (Cantlon & Piantadosi, 2024; Herculano-Houzel, 2012), specialized architectural features likely also play a critical role (Friederici & Becker, 2025).”

      Additionally, the authors claim that their results show 'larger models better predict the structure of natural language.' However, it remains unclear to what extent the embeddings of LLMs capture the "structure" of language better than the lexical semantics of language.

      We appreciate the reviewer's point about how well LLM embeddings capture the "structure" of language versus just lexical semantics. It's true that distinguishing these aspects is complex. From our perspective, a model's ability to predict/produce natural language entails that the model captures various levels of linguistic structure, including morphology, syntax, semantics, and contextual dependencies. We use "structure" inclusively in this sense. A model cannot achieve high predictive accuracy without representing, to some extent, all of these structural elements (Linzen & Baroni, 2021; Manning et al., 2020; Pavlick, 2022). There is a very active field of research into understanding exactly how these models represent these different structures of language (Ameisen et al., 2025; Chemla et al., 2024; Elhage et al., 2021, 2022; Hewitt & Manning, 2019). Our results confirm the core trend that larger models tend to better reproduce the various structures of language (i.e., yield lower perplexity; Fig. 2A).

      In previous work, we have shown that LLM embeddings better predict neural activity during natural language processing than lexical embeddings (e.g., GloVe) that do not contain other elements of linguistic structure (Goldstein et al., 2022; Kumar et al., 2024; Zada et al., 2024). In response to the following comment, we also compare LLMs to simpler models capturing specific speech and language features (see next comment). To clarify our intended use of the word “structure”, we’ve added a brief explanation in the Methods section:

      “In this study, we use the term “structure” to refer to a variety of linguistic patterns (e.g., morphology, syntax, semantics, context) that LLMs encode in order to better predict natural language.”

      (2) The study lacks control LLMs with randomly initialized weights and control regressors, such as word frequency and phonetic features of speech, making it unclear what the baseline is for the model-brain correlation.

      We’ve added several supplementary analyses to the revised manuscript to address these concerns. To establish a baseline, we extracted embeddings from each layer of the SMALL model with randomly initialized weights and constructed encoding models. The encoding performance is significantly higher for pretrained SMALL than for untrained SMALL for every layer (Fig. S4). For the untrained model, the performance is the highest for the 0th layer and decreases in subsequent layers. This is because at the 0th layer, every instance of the same word receives an identical, albeit random, embedding (See Supplementary Figure 4).

      We also compared the encoding performance of LLMs with more classical speech/language features and static GloVe embeddings (Goldstein, Wang, et al., 2025; Kumar et al., 2024). First, we extracted features capturing lower-level speech features. Using the stimulus transcript as input, we created one-hot vectors for phonetic and articulatory features. Phoneme classes (39 total classes) were obtained from the Carnegie Mellon Pronouncing Dictionary (The CMU Pronouncing Dictionary, n.d.). We further classified the phonemes based on their place of articulation (9 classes), manner of articulation (9 classes), and voiced or voiceless status (3 classes), according to the general American English consonants of the International Phonetic Alphabet. Given that each word consists of multiple phonemes, we averaged the one-hot vectors for all phonetic and articulatory features for each word.

      Second, we extracted linguistic features using spaCy (Honnibal et al., 2020), including part of speech (17 classes), tag (50 classes), function or content word (3 classes), dependency (45 classes), whether the word is an alpha character (binary), and whether the word is a stop word (binary). We also extracted prefix (30 classes) and suffix (44 classes) information using the Cambridge Dictionary. We constructed one-hot vectors for each multi-class feature and one-dimensional vectors for each binary feature.

      Third, for each word, we obtained word frequency from the Google Web Trillion Word Corpus (Brants & Franz, 2006) and from our own dataset.

      Fourth, we generated static word embeddings of dimension 50 using GloVe (Pennington et al., 2014).

      We then built encoding models in the same way as the contextual embeddings for each of the three categories of speech features, all speech features concatenated, and the GloVe embeddings. To control for the different dimensions of the embeddings, we also standardized all embeddings to the same size (50 dimensions) using principal component analysis (PCA) and trained linear encoding models using ordinary least-squares (OLS) regression. For both ridge and OLS encoding, our contextual embeddings from LLMs showed significantly better performance than the classic speech features and GloVe embeddings.

      We have added the following text to our manuscript and updated our Figures S4, S5, Table S1, and the methods section:

      “To establish a general baseline for encoding performance, we built encoding models using embeddings from the SMALL model with randomly initialized weights. The trained SMALL model exhibits significantly higher encoding performance across all layers compared to the untrained SMALL model (Fig. S4). We also assessed the encoding performance of contextual embeddings from LLMs against classic speech features and static GloVe embeddings (Table S1). The SMALL and XL embeddings achieved markedly higher encoding correlations than the speech features and GloVe embeddings (Fig. S5).”

      (3) The finding that peak encoding performance tends to occur in relatively earlier layers in larger models is somewhat surprising and requires further explanation. Since more layers mean more parameters, if the later layers diverge from language processing in the brain, it raises the question of what aspects of the larger models make them more brain-like.

      We thank the reviewer for this insightful comment; this point was also highlighted by Reviewer 1. We agree that this result is somewhat surprising, and we aim to provide a more detailed explanation in the revised manuscript. The general inverted U-shaped trend of encoding performance across layers has been a frequently observed phenomenon in studies comparing LLM representations to brain activity (Goldstein, Ham, et al., 2025; Schrimpf et al., 2021). A potential explanation is the existence of a “two-phase abstraction process” within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the initial layers, models begin by processing relatively low-level input features. As layers get deeper, representations become increasingly abstract and richly contextualized in semantic features relevant for understanding language. These intermediate layers often show the highest correlation with brain activity in language areas, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. These observations suggest that it is primarily the abstractive, contextual features developed in the intermediate layers of LLMs that drive their alignment with brain activity. As models become more potent at prediction, their most predictive layers (for the LLM’s natural language task) and their most generalizable layers (for brain activity) can diverge.

      A key finding in our study is that the initial processing phase does not scale and take up more layers as models scale up in size and layers. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models.

      Consequently, the prediction phase may begin relatively earlier in these larger models, and the later layers could develop highly specialized representations that are increasingly divergent from the more general linguistic processing captured in brain activity. For example, these layers may specialize in capturing very specific patterns of language (thus lowering their perplexity) that do not actually occur often or at all in our naturalistic dataset. It is also possible that the later layers of larger models are overall underutilized and do not contribute as much to linguistic processing and next-word prediction (Csordás et al., 2025).

      We have added the following text to our Discussion section:

      “The inverted U-shaped trend of encoding performance commonly found in previous research is likely due to a "two-phase abstraction process" within LLMs (Cheng & Antonello, 2024; Csordás et al., 2025). In the early and intermediate layers of the model, a composition phase occurs, where low-level input features become increasingly abstract and contextualized. The intermediate layers of the model show the highest correlation with brain activity, presumably because they capture complex semantic and contextual information in a way that generalizes well across a variety of tasks (including prediction of human neural activity) (Antonello & Huth, 2024). Subsequently, a prediction phase happens in the later layers of the model, where the representations become more specialized for the LLM's specific training objective (e.g., next-word prediction). This specialization can effectively constrict the more generalized feature representations, making these layers less optimal for predicting brain activity. Our results indicate that the initial composition phase does not take up more layers as models scale up in size. Larger models develop the necessary rich, abstract representations in the same number of layers as smaller models. Thus, as LLMs increase in size, the later layers of the model may contain representations that are increasingly divergent from the more general linguistic processing captured in brain activity. It is also possible that the later layers of larger models are overall underutilized and may not significantly contribute to benchmark performances during inference (Csordás et al., 2025; Fan et al., 2024; Gromov et al., 2024).”

      Reviewer #3 (Public review):

      This manuscript studies the connection between neural activity collected through electrocorticography and hidden vector representations from autoregressive language models, with the specific aim of studying the influence of language model size on this connection. Neural activity was measured from subjects who listened to a segment from a podcast, and the representations from language models were calculated using the written transcription as the input text. The ability of vector representations to predict neural activity was evaluated using 10-fold cross-validation with ridge regression models.

      The main results are that (as well summarized in section headings):

      (1) Larger models predict neural activity better.

      (2) The ability of language model representations to predict neural activity differs across electrodes and brain regions.

      (3) The layer that best predicts neural activity differs according to model size, with the "SMALL" model showing a correspondence between layer number and the language processing hierarchy.

      (4) There seems to be a similar relationship between the time lag and the ability of language model representations to predict neural activity across models.

      Strengths:

      (1) The experimental and modeling protocols generally seem solid, which yielded results that answer the authors' primary research question.

      (2) Electrocorticography data is especially hard to collect, so these results make a nice addition to recent functional magnetic resonance imaging studies.

      We thank the reviewer for their thoughtful and positive feedback.

      Weaknesses:

      (1) The interpretation of some results seems unjustified, although this may just be a presentational issue.

      (a) Figure 2B: The authors interpret the results as "a plateau in the maximal encoding performance," when some readers might interpret this rather as a decline after 13 billion parameters. Can this be further supported by a significance test like that shown in Figure 4B?

      We agree that this could be a subjective interpretation, so we conducted an additional analysis. We performed paired two-sided t-tests between best layer encoding performances averaged across electrodes (df = 159 electrodes), comparing all models with larger models. We found that after 13 billion parameters, only the encoding performance for OPT-66B, the largest model in the OPT family, is significantly worse than the encoding performance of some other smaller models, supporting the claim that the maximal encoding performance declines after 13 billion parameters. However, we did not find conclusive statistical evidence of a decline in encoding performance for other model families.

      We have added the statistical results as Supplementary Figure 1.

      We have also modified the following text in the manuscript:

      “We also observed a plateau in the maximal encoding performance, occurring around 7 billion parameters (Fig. 2B), with a decline in performance for the OPT-66B model (Fig. S1).”

      (b) Figure S1A: It looks like the drop in PCA max correlation is larger for larger models, which may suggest to some readers that the same trend observed for ridge max correlation may not hold, contra the authors' claim that all results replicate. Why not include a similar figure as Figure 2B as part of Figure S1?

      PCA is an unsupervised dimensionality reduction technique and may discard model features with small eigenvalues that nonetheless contribute to encoding performance. Ridge regression, a supervised method, can capitalize on these features. We suspect that this is why there appears to be a larger drop in model performance for larger models with PCA than with ridge regression. We replicated the logarithmic relationship between model size and encoding performance using PCA and ordinary least-squares (OLS) regression encoding models. We have updated Supplementary Figure 2.

      (2) Discussion of what might be driving the main result about the influence of model size appears to be missing (cf. the authors aim to provide an explanation of what seems to drive the influence of the layer location in Paragraph 3 of the Discussion section). What explanations have been proposed in the previous functional magnetic resonance imaging studies? Do those explanations also hold in the context of this study?

      We suspect that the increased expressivity of larger models - that is, their improved sensitivity to nuanced structure in natural language - yields improved alignment to brain activity (given large enough samples of brain activity) (Antonello et al., 2023). This effect persists even when dimensionality is tightly controlled in our PCA-based analysis, indicating that the improved alignment with the brain is not a modeling artifact of dimensionality alone, but results from the structural representations learned by these larger models.

      We have added the following text to our Discussion section:

      “We suspect that the improved alignment with brain activity in larger models is driven by their increased expressivity and sensitivity to nuanced linguistic structure present in large-scale naturalistic datasets (Antonello et al., 2023).”

      (3) The GloVe-based selection of language-sensitive electrodes (at least to me) isn't explained/motivated clearly enough (I think a more detailed explanation should be included in the Materials and Methods section). If the electrodes are selected based on GloVe embeddings, then isn't the main experiment just showing that representations from larger language models track more closely with GloVe embeddings? What justifies this methodology?

      We selected electrodes based on previously established methods (Goldstein et al., 2022). Our use of GloVe embeddings for electrode selection does not imply that larger language model representations are simply more closely aligned with GloVe embeddings. On the contrary, contextual embeddings from LLMs, which incorporate the word’s previous context, consistently outperform static embeddings like GloVe or word2vec (Fig. S3). Selecting electrodes using LLM embeddings would likely result in a slightly different, potentially larger set of electrodes (Goldstein et al., 2022), but would be more circular (Kriegeskorte et al., 2009). The GloVe-based electrode selection represents a more conservative approach by identifying words encoding linguistic content without biasing the selection directly toward any LLMs.

      We have added the following text to our Method section:

      “We used GloVe embeddings for electrode selection to avoid biasing our main results toward a particular LLM.”

      (4) (Minor weakness) The main experiments are largely replications of previous functional magnetic resonance imaging studies, with the exception of the one lag-based analysis. Is there anything else that the electrocorticography data can reveal that functional magnetic resonance imaging data can't?

      We thank the reviewer for this thoughtful question. While we agree that a key contribution of our work corroborates previous fMRI findings, we would argue that using ECoG is not merely a replication but a crucial validation and extension of that work. It is important to validate these effects across distinct measurement modalities. In our work, we further observed a novel trend where the peak encoding performance tends to occur in relatively earlier layers for larger models. This is supported by recent studies suggesting that later layers of large LLMs may not significantly contribute to benchmark performance (Csordás et al., 2025). While scaling has been an effective method to improve LLM performance, including in encoding models, future research should explore the potential underutilization of the later layers as models scale.

      Furthermore, ECoG data offers temporal resolution on the order of milliseconds, far superior to fMRI’s. Although we did not observe a relationship between model size and temporal lags in this study, future work should investigate the temporal dynamics of encoding that are accessible with ECoG (Goldstein, Ham, et al., 2025; Goldstein, Wang, et al., 2025).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Thank you to the authors for the fun and personally useful read.

      I see in Supplementary Figure 1 the authors show a comparison of the performance between OLS vs. Ridge regression. Is the OLS model the only one that is working over PC features, or are both models using PC features? The current text is a bit unclear. My current understanding is that the comparison is between (OLS + PCA) and (Ridge with no PCA), but I am not sure.

      The OLS model is the only one that works over PC features, following previous methods (Goldstein et al., 2022).

      We have added the following text to our Results and Methods section for clarity:

      “To control for the different embedding dimensionality across models, we standardized all embeddings to the same size using principal component analysis (PCA) and trained linear encoding models using ordinary least-squares (OLS) regression, replicating the logarithmic relationship but with significantly lower encoding performance overall (Fig. S2). The PC features are used by the OLS models only.”

      Clarification in the text would be appropriate. If this is the correct understanding, the authors should note in the main text that the ridge approach is more effective than the PCA approach, which is still the dominant approach to building linear encoding models in the field for some unjustifiable reason.

      We thank the reviewer for pointing out the confusion. We have updated Supplementary Figure 2.

      How were the alpha values for ridge regression determined? Do you use the same ridge parameter for all electrodes or fit a different parameter for each electrode? This is not mentioned anywhere.

      The alpha values are determined by cross-validation using the “RidgeCV” method from the “himalaya” package (Dupré la Tour et al., 2022). Specifically, we perform a grid search over cross-validation folds in the training data to find the best-performing alpha. The alpha parameter is specific to each ridge regression model, meaning each fold, lag, and electrode has a different alpha parameter.

      We have added the following text to our manuscript:

      “For each ridge regression model (for each fold, lag, and electrode), the alpha parameter is determined by cross-validation using the “RidgeCV” method from the “himalaya” package (Dupré la Tour et al., 2022).”

      It's not entirely clear to me how the authors handle tokens that do not terminate in words (such as the "there" + "'s" example in the text). My current reading of the text is that authors essentially ignore these half-word embeddings, doing one forward pass per word, rather than per token, but the current text is somewhat ambiguous.

      If a word is tokenized into several tokens, like “there” and “‘s”, we average the token embeddings to get a word embedding.

      We have added the following text to our Method section:

      “To facilitate a fair comparison of the encoding effect across different models, we aligned all tokens in the story across all models. We averaged the token embeddings if a word is split into multiple tokens, resulting in one embedding per word for each model.”

      The authors describe the scaling relationship they find as a "log-linear" relationship. I believe this is a misnomer derived from the original paper describing this relationship in fMRI as log-linear (Antonello et al.) The correct term is simply "logarithmic", and for what it's worth, the authors of the original fMRI work have made this correction as well.

      Thank you! We have made this correction.

      Is the data publicly available? If not, there should be some basic justification as to why (consent reasons, etc.).

      We have recently made the data publicly available (Zada et al., 2025). We have also provided tutorials for preprocessing the data and training encoding models: https://hassonlab.github.io/podcast-ecog-tutorials. For this specific project, the analysis code is available at https://github.com/hassonlab/247-pickling/tree/scaling-paper-0 and https://github.com/hassonlab/247-encoding/tree/scaling-paper-1.

      The authors assert that ECoG has "superior spatiotemporal resolution". While this is unquestionably true for temporal resolution, the story is a bit more complicated for spatial resolution, where ECoG has far less cortical coverage than fMRI. Perhaps this sentence should be revised.

      Thank you for pointing out the typo! We have changed it to “superior temporal resolution”.

      Minor Points:

      The bolded title of Figure 3 probably shouldn't be bolded, as this is just actually the title of Figure 3A.

      Fixed.

      Figure 4d is has a typo: "Best Encoidng Layer".

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      The authors could consider adding control regressors such as word rate, word frequency, phonetic features, and syntactic features like node counts, as well as control LLMs of comparable size to serve as baselines. The authors could also include correlation analyses of the embeddings from different layers of the same LLM to further illustrate how distinct the layers are within the models.

      We have added untrained LLM embeddings as a baseline and included a comparison of encoding models between LLM contextual embeddings and classical speech features. We have also performed some preliminary correlation analyses of embeddings. In some models, we found evidence of the “two-phase abstraction process” (Cheng & Antonello, 2024). However, the result is inconclusive across different LLM families. Since each LLM layer accesses and modifies the residual stream (Elhage et al., 2021), the embeddings across layers are inherently correlated. Future work could instead explore the isolated transformations within each layer to illustrate the distinct information across layers (Kumar et al., 2024).

      The analysis codes and data should be made available.

      We have recently made the data publicly available (Zada et al., 2025). We have also provided tutorials for preprocessing the data and training encoding models: https://hassonlab.github.io/podcast-ecog-tutorials. For this specific project, the analysis code is available at https://github.com/hassonlab/247-pickling/tree/scaling-paper-0 and https://github.com/hassonlab/247-encoding/tree/scaling-paper-1.

      Reviewer #3 (Recommendations for the authors):

      Most of my concrete recommendations are in the public review. Below are some additional minor ones:

      (1) Introduction: "Remarkably, these models learn from much the same shared space as humans: from real-world language generated by humans."

      I think this is an extremely strong claim due to e.g. the different nature of child-directed speech vs. written text corpora, the lack of multimodality and grounding in language models, etc. I might suggest re-wording this sentence or removing it entirely.

      We thank the reviewer for their suggestion! We have removed the sentence from the manuscript.

      (2) Introduction: "EleutherAI, n.d." reference for GPT-Neo

      GPT-NeoX-20B has an associated paper, which the authors might cite instead: https://aclanthology.org/2022.bigscience-1.9

      Thank you! We have added the reference for GPT-NeoX-20B (Black et al., 2022).

      (3) Figure 4D: Encoidng -> Encoding

      Fixed.

      (4) Materials and Methods, Contextual embeddings: "except for GPT-Neox-20b, which assigns additional tokens to whitespace characters."

      What do the authors mean by "additional tokens to whitespace characters?" The tokenizer for GPT-NeoX-20B works in much the same way as that of GPT-Neo, just with a different vocabulary set.

      >>> t1 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neo-125M")

      >>> t2 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neox-20b")

      >>> t1.convert_ids_to_tokens(t1("The quick brown fox jumps over the lazy dog.").input_ids) ['The', 'Ġquick', 'Ġbrown', 'Ġfox', 'Ġjumps', 'Ġover', 'Ġthe', 'Ġlazy', 'Ġdog', '.']

      >>> t2.convert_ids_to_tokens(t2("The quick brown fox jumps over the lazy dog.").input_ids)

      ['The', 'Ġquick', 'Ġbrown', 'Ġfox', 'Ġjumps', 'Ġover', 'Ġthe', 'Ġlazy', 'Ġdog', '.']

      If the authors are referring to Ġ as the "additional token to whitespace characters," then these are in all other tokenizers as well (not only that for GPT-Neo, but also those for GPT-2 and OPT).

      We agree that “additional tokens to whitespace characters” is an oversimplification. The GPT-Neo model family, which includes the 125M, 1.3B, and 2.7B models, utilizes the same Byte Pair Encoding (BPE) tokenizer as GPT-2. This common tokenizer has a vocabulary size of 50,257 tokens, providing compatibility and seamless integration across the models.

      The GPT-NeoX-20B model introduces a modified tokenizer to address limitations observed in the GPT-2 tokenizer (Black et al., 2022). As detailed in Section 3.2, this new tokenizer incorporates a few key improvements:

      (1) New BPE tokenizer: A more general-purpose BPE tokenizer was trained using the Pile dataset.

      (2) Space Delimitation: Unlike the GPT-2 tokenizer, which treats tokenization at the start of a string as a non-space-delimited token, the GPT-NeoX-20B tokenizer applies consistent space delimitation regardless. This change resolves inconsistencies related to the presence of prefix spaces in the tokenization input.

      (3) Whitespace Handling: The tokenizer includes tokens for repeated space characters (up to 24 consecutive spaces), enhancing efficiency in tokenizing text with substantial whitespace, such as program source code or LaTeX documents.

      These modifications result in the GPT-NeoX-20B tokenizer representing the Pile validation set with approximately 10% fewer tokens than the GPT-2 tokenizer. This efficiency gain is particularly beneficial for processing texts with extensive whitespace.

      In our analysis, we extracted embeddings by setting `add_prefix_space = True` to all tokenizers, so space delimitation does not result in tokenizer differences. We highlight here examples of the other two tokenizer differences using the Huggingface `AutoTokenizer`:

      >>> t1 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neo-125M")

      >>> t2 = AutoTokenizer.from_pretrained("EleutherAI/gpt-neox-20b")

      >>> t1.convert_ids_to_tokens(t1("The Downing Street.").input_ids) ['The', 'ĠDowning', 'ĠStreet']

      >>> t2.convert_ids_to_tokens(t2("The Downing Street.").input_ids)

      ['The', 'ĠDown', 'ing', 'ĠStreet']

      >>> t1.convert_ids_to_tokens(t1("Hello !").input_ids)

      ['Hello', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ', 'Ġ!']

      >>> t2.convert_ids_to_tokens(t2("Hello !").input_ids)

      ['Hello', ' ', '!']

      More examples showing the differences between the GPT-2 tokenizer and the GPT-NeoX-20B tokenizer can be found in Appendix F: Tokenizer Analysis (Black et al., 2022).

      We have added the following text to our manuscript for simplicity:

      “All models within the same model family adhere to the same tokenizer convention, except for GPT-Neox-20B, which utilizes a different tokenizer (Black et al., 2022).”

      References

      Ameisen, E., Lindsey, J., Pearce, A., Gurnee, W., Turner, N. L., Chen, B., Citro, C., Abrahams, D.,  Carter, S., Hosmer, B., Marcus, J., Sklar, M., Templeton, A., Bricken, T., McDougall, C.,  Cunningham, H., Henighan, T., Jermyn, A., Jones, A., … Batson, J. (2025). Circuit Tracing:  Revealing Computational Graphs in Language Models. Transformer Circuits Thread. https://transformer-circuits.pub/2025/attribution-graphs/methods.html

      Antonello, R., & Huth, A. (2024). Predictive coding or just feature discovery? An alternative account of why language models fit brain data. Neurobiology of Language (Cambridge, Mass.), 5(1), 64–79.

      Antonello, R., Vaidya, A., & Huth, A. G. (2023). Scaling laws for language encoding models in fMRI. NeurIPS 2023. https://doi.org/10.48550/ARXIV.2305.11863

      Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell,  K., Phang, J., Pieler, M., Prashanth, U. S., Purohit, S., Reynolds, L., Tow, J., Wang, B., & Weinbach, S. (2022). GPT-NeoX-20B: An Open-Source Autoregressive Language Model.  Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models. Proceedings of BigScience Episode #5 -- Workshop on Challenges & Perspectives in Creating Large Language Models, virtual+Dublin. https://doi.org/10.18653/v1/2022.bigscience-1.9

      Brants, T., & Franz, A. (2006). Web 1T 5-gram Version 1 [Dataset]. Linguistic Data Consortium. https://doi.org/10.35111/CQPA-A498

      Cantlon, J. F., & Piantadosi, S. T. (2024). Uniquely human intelligence arose from expanded information capacity. Nature Reviews Psychology, 3(4), 275–293.

      Chemla, E., D’Ascoli, S., Diego-Simón, P., King, J.-R., & Lakretz, Y. (2024). A Polar coordinate system represents syntax in large language models. Advances in Neural Information Processing Systems 37, 105375–105396.

      Cheng, E., & Antonello, R. J. (2024). Evidence from fMRI supports a two-phase abstraction process in language models. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2409.05771

      Csordás, R., Manning, C. D., & Potts, C. (2025). Do language models use their depth efficiently?  In arXiv [cs.LG]. https://doi.org/10.48550/ARXIV.2505.13898

      Dupré la Tour, T., Eickenberg, M., Nunez-Elizalde, A. O., & Gallant, J. L. (2022). Feature-space selection with banded ridge regression. In bioRxiv. https://doi.org/10.1101/2022.05.05.490831

      Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., & Olah, C. (2022). Toy Models of Superposition. Transformer Circuits Thread.

      Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A.,  Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A.,  Kernion, J., Lovitt, L., Ndousse, K., … Olah, C. (2021). A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread.

      Fan, S., Jiang, X., Li, X., Meng, X., Han, P., Shang, S., Sun, A., Wang, Y., & Wang, Z. (2024). Not all Layers of LLMs are Necessary during Inference. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2403.02181

      Friederici, A. D., & Becker, Y. (2025). The core language network separated from other networks during primate evolution. Nature Reviews. Neuroscience, 26(2), 131–132.

      Goldstein, A., Ham, E., Schain, M., Nastase, S. A., Aubrey, B., Zada, Z., Grinstein-Dabush, A.,  Gazula, H., Feder, A., Doyle, W., Devore, S., Dugan, P., Friedman, D., Brenner, M., Hassidim, A., Matias, Y., Devinsky, O., Siegelman, N., Flinker, A., … Hasson, U. (2025). Temporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models. Nature Communications, 16(1), 10529.

      Goldstein, A., Wang, H., Niekerken, L., Schain, M., Zada, Z., Aubrey, B., Sheffer, T., Nastase, S. A., Gazula, H., Singh, A., Rao, A., Choe, G., Kim, C., Doyle, W., Friedman, D., Devore, S., Dugan, P., Hassidim, A., Brenner, M., … Hasson, U. (2025). A unified acoustic-to-speech-to-language embedding space captures the neural basis of natural language processing in everyday conversations. Nature Human Behaviour. https://doi.org/10.1038/s41562-025-02105-9

      Goldstein, A., Zada, Z., Buchnik, E., Schain, M., Price, A., Aubrey, B., Nastase, S. A., Feder, A.,  Emanuel, D., Cohen, A., Jansen, A., Gazula, H., Choe, G., Rao, A., Kim, C., Casto, C., Fanda, L., Doyle, W., Friedman, D., … Hasson, U. (2022). Shared computational principles for language processing in humans and deep language models. Nature Neuroscience, 25(3), 369–380.

      Gromov, A., Tirumala, K., Shapourian, H., Glorioso, P., & Roberts, D. A. (2024). The Unreasonable Ineffectiveness of the Deeper Layers. In arXiv [cs.CL]. arXiv. http://arxiv.org/abs/2403.17887

      Herculano-Houzel, S. (2012). The remarkable, yet not extraordinary, human brain as a scaled-up primate brain and its associated cost. Proceedings of the National Academy of Sciences of the United States of America, 109 Suppl 1(supplement_1), 10661–10668.

      Hewitt, J., & Manning, C. D. (2019). A Structural Probe for Finding Syntax in Word Representations. In J. Burstein, C. Doran, & T. Solorio (Eds.), Proceedings of the 2019 Conference of the North (pp. 4129–4138). Association for Computational Linguistics.

      Honnibal, M., Montani, I., Van Landeghem, S., & Boyd, A. (2020). spaCy: Industrial-strength Natural Language Processing in Python.

      Kriegeskorte, N., Simmons, W. K., Bellgowan, P. S. F., & Baker, C. I. (2009). Circular analysis in systems neuroscience: the dangers of double dipping. Nature Neuroscience, 12(5),  535–540.

      Kumar, S., Sumers, T. R., Yamakoshi, T., Goldstein, A., Hasson, U., Norman, K. A., Griffiths, T. L., Hawkins, R. D., & Nastase, S. A. (2024). Shared functional specialization in transformer-based language models and the human brain. Nature Communications, 15(1), 5523.

      Linzen, T., & Baroni, M. (2021). Syntactic Structure from Deep Learning. Annual Review of Linguistics, 7(1), 195–212.

      Manning, C. D., Clark, K., Hewitt, J., Khandelwal, U., & Levy, O. (2020). Emergent linguistic structure in artificial neural networks trained by self-supervision. Proceedings of the National Academy of Sciences of the United States of America, 117(48), 30046–30054.

      Millet, J., Caucheteux, C., Orhan, P., Boubenec, Y., Gramfort, A., Dunbar, E., Pallier, C., & King, J.-R. (2023). Toward a realistic model of speech processing in the brain with self-supervised learning. NeurIPS 2022. https://doi.org/10.48550/ARXIV.2206.01685

      Pavlick, E. (2022). Semantic structure in deep learning. Annual Review of Linguistics, 8(1),  447–471.

      Pennington, J., Socher, R., & Manning, C. (2014). Glove: Global vectors for word representation.  Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar. https://doi.org/10.3115/v1/d14-1162

      Schrimpf, M., Blank, I. A., Tuckute, G., Kauf, C., Hosseini, E. A., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2021). The neural architecture of language: Integrative modeling converges on predictive processing. Proceedings of the National Academy of Sciences of the United States of America, 118(45), e2105646118.

      The CMU Pronouncing Dictionary. (n.d.). Retrieved May 27, 2025, from http://www.speech.cs.cmu.edu/cgi-bin/cmudict

      Vaidya, A. R., Jain, S., & Huth, A. G. (2022). Self-supervised models of audio effectively explain human cortical responses to speech. ICML 2022. https://doi.org/10.48550/ARXIV.2205.14252

      Zada, Z., Goldstein, A., Michelmann, S., Simony, E., Price, A., Hasenfratz, L., Barham, E., Zadbood,  A., Doyle, W., Friedman, D., Dugan, P., Melloni, L., Devore, S., Flinker, A., Devinsky, O., Nastase, S. A., & Hasson, U. (2024). A shared model-based linguistic space for transmitting our thoughts from brain to brain in natural conversations. Neuron, S0896627324004604. Zada, Z., Nastase, S. A., Aubrey, B., Jalon, I., Michelmann, S., Wang, H., Hasenfratz, L., Doyle, W.,  Friedman, D., Dugan, P., Melloni, L., Devore, S., Flinker, A., Devinsky, O., Goldstein, A., & Hasson, U. (2025). The “Podcast” ECoG dataset for modeling neural activity during natural language comprehension. Scientific Data, 12(1), 1135.

    1. Author response:

      The following is the authors’ response to the previous reviews

      We are grateful to you and the reviewers for your careful and positive assessment.  In response to the comment below from reviewer #2, we have expanded Figure 1- figure supplement 1 to show the results (images plus quantification) of ERG and PU.1 immunostaining together with DAPI staining that underly the calculation of the percent on non-endothelial cells that are recombined by Cdh5-CreER.

      However, I found their explanation as to how they arrived at only 18% of the fibroblasts being recombined confusing. Mostly because it looks like there are many GFP+/ERG- cells in panels A & D, these would be the recombined fibroblasts and AB cells. Considering that Cdh5 gene is pretty broadly expressed across LPM, the prediction is the recombination rate would be higher (recognizing mice strains vary and this is inducible Cre).

      Ideally, they would perform recombination analysis with a TF expressed by all fibroblasts in combination with Erg (like Foxc1). However, a potentially simpler approach could be to quantify this using DAPI and Erg in existing images, of the total DAPI+, what are GFP+/DAPI+/Erg- (fibroblasts) vs GFP+/ERG+/DAPI+ (endothelial). The figure would be improved by adding DAPI to one set of panels with GFP/ERG, this would show a lot of DAPI+/GFP- cells, encompassing CD206+ cells and non-recombined fibroblasts.

      Thank you for overseeing this manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This well-conceived manuscript investigates the mechanisms that shape the chromatin landscape following fertilization, using the Drosophila embryo as a model system. Importantly, the authors revisit conflicting data using new approaches and analysis to show that the silent H3K27me3 mark deposited by PRC2 is established de novo in the embryo in coordination with the slowing of the nuclear division cycle and activation of zygotic transcription. Unexpectedly, they demonstrate that the transcription factor GAF is not required for the deposition of this mark, but that the well-studied pioneer factor Zelda, which is required for widespread gene expression, is required for H3K27me3 deposition at a subset of regions. The experiments are rigorously performed, and interpretations are clear. Strengths of this manuscript include the rigor of the experimental design, careful analysis, and well-supported conclusions. Some additional citations, analysis, and broadening of the Discussion section to include additional models and data would further strengthen this manuscript.

      We appreciate the reviewer’s positive assessment and in revision we have revised the Discussion to clarify some of the mechanistic insights of the work as well as including references to related studies in the zebrafish model system.

      Reviewer #2 (Public review):

      Summary:

      Epigenetic silencing of target genes by the Polycomb pathway is central to maintenance of cell fates during development and depends on repressive chromatin states involving Polycomb complexes and histone modifications. However, the mechanisms by which these chromatin states are built at the earliest stages of development are unclear. Here, Gonzaga-Saavedra and colleagues use the premier experimental system for studying Polycomb gene regulation, Drosophila development, to investigate when Polycomb domains emerge and how they are assembled. Using a combination of CRISPR gene editing, imaging, and genomic profiling, they determine that while H3K27me3 is initially present in the first nuclear cycles, it quickly dissipates and does not re-emerge until mid-nuclear cycle 14, during the major wave of zygotic genome activation (ZGA). This finding helps resolve current discrepancies in the field, informs potential mechanisms of transgenerational inheritance, and indicates that repressive Polycomb domains are built de novo on target genes in embryogenesis. The authors then set out to examine how Polycomb domains are built. Through live imaging and immunofluorescence, they determine that the histone H3K27 methyltransferase, E(z), is present in nuclei at high levels throughout cleavage and blastoderm stages. By contrast, they determine that several Polycomb proteins that bind PREs (cis elements that demarcate Polycomb targets in the genome) are absent from early cleavage nuclei and progressively increase following nuclear cycle 10. These findings suggest that the absence of H3K27me3 in early embryos may be due to failure to assemble functional Polycomb complexes at target genes. Lastly, the authors test the requirement of two transcription factors with important roles in ZGA, GAF, and ZLD. Despite binding to many PREs and regulating chromatin accessibility in early embryos, they find that GAF is largely dispensable for the emergence of H3K27me3 domains. On the other hand, they find that the pioneer factor ZLD is required for proper H3K27me3 emergence; in its absence, some Polycomb domains accumulate greater levels of H3K27me3, whereas other Polycomb domains accumulate less H3K27me3.

      Strengths:

      The strengths of this study are manifold. It studies an important topic with broad interest to the chromatin and epigenetics fields. It is well-written with detailed method descriptions. In addition, the experimental design and rigor of execution are exceptional despite working with very small amounts of biological material. Example strengths include that the Polycomb proteins studied were tagged with the same epitope, permitting direct quantitative comparisons in imaging and in genomics experiments. Microscopy studies are quantified and performed both via live imaging and via immunofluorescence. The microscopy studies reinforce and extend conclusions made via ChIP. Sophisticated loss-of-function analyses allow for direct mechanistic tests of Polycomb domain emergence.

      Weaknesses:

      Overall, the study is quite strong already, but it can be further strengthened in several ways. First, several conclusions should be refined based on the data presented. Second, the extent to which ZLD is important for initiating Polycomb domain formation should be made clearer. Third, additional genomic profiling experiments are needed to provide insight into models explaining why H3K27me3 is absent prior to NC14.

      We are grateful for the reviewer’s thorough and supportive comments. We have revised certain assertions and conclusions for objectivity. For the point about providing “insight into models explaining why H3K27me3 is absent prior to NC14,” we have a separate study that addresses this issue directly (Degen, Gonzaga-Saavedra, and Blythe, bioRxiv 2025, in press). In summary, we find evidence that a maternal PcG imprint is indeed maintained through cleavage divisions, albeit through lower-order methylation states (maximally, H3K27me2). We chose not to include these additional results in this manuscript to maintain the focus of this study on ZGA. Our revision of the manuscript includes a reference to this associated study in the Discussion.

      Reviewer #3 (Public review):

      Gonzaga-Saavedra et al report an analysis on genomic binding of Polycomb group proteins, and of H2Aub1 and H3K27me3 domain formation in the early Drosophila embryo. Using carefully staged embryos during the nuclear cycles (NC) leading up to the cellular blastoderm stage, the authors provide compelling evidence that H3K27me3 domains at PcG target genes are only established during NC14 and do not exist in NC13. In contrast, H2Aub1 domains already start to appear during NC13. The authors show that E(z), the catalytic subunit of the H3K27 histone methyltransferase PRC2, is readily detected in interphase nuclei during the rapid nuclear divisions in pre-blastoderm embryos. In contrast, the DNA-binding proteins Pho, Cg, and GAF that are known (Pho) or have been postulated (Cg, GAF) to anchor PRC2 and PRC1 to Polycomb Response Elements (PREs) in Polycomb target genes only start to show nuclear localization from NC10 onwards with gradually increasing nuclear concentrations, reaching a maximum during NC14. These data strongly corroborate the simple, straightforward view that targeting of PRC2 and PRC1 to PREs by sequence-specific DNA-binding proteins is a prerequisite for the formation of H3K27me3 and H2Aub1 domains at Polycomb target genes.

      The authors then explore the potential role of GAF/Trl in this process. They find that in embryos depleted of GAF/Trl, H3K27me3 domain formation is largely unperturbed.

      The authors also depleted the pioneer factor Zelda (Zld) and found that removal of Zld results in a more complex outcome. Zelda appears to counteract the accumulation of H3K27me3 at the Polycomb targets eve and zen, but also appears to be required for effective H3K27me3 domain formation at Polycomb targets such as amos or atonal.

      This is a very thorough study that reports data of superior technical quality that are highly relevant for the field. The study by Gonzaga-Saavedra et al extends and strengthens previous work from the labs of Eisen (Li et al, eLife 2014) and Zeitlinger (Chen et al, eLife 2013) to convincingly demonstrate that Polycomb domain formation in the early embryo occurs during ZGA but that such domains do not exist prior to ZGA. This should now finally put to rest earlier claims by the Iovino lab (Zenk et al, Science 2017) that H3K27me3 domains present in the zygote nucleus would be propagated and partially maintained during the rapid nuclear cleavage cycles and serve as seeds for H3K27me3 domain formation during ZGA.

      The experiments analyzing H3K27me3 domain formation in embryos depleted of GAF/Trl or Zelda will be of great interest to the field.

      We thank the reviewer for recognizing the strength of our data and conclusions, and we agree that our results help settle conflicting claims in the field. We have emphasized Zelda’s context-dependent effects more clearly in the revised manuscript.

      Recommendations for the authors:

      Reviewing Editor Comments:

      It would strengthen the manuscript to more fully acknowledge and discuss related work in the field. Addressing the caveats in the functional analyses, either through editorial clarification or additional experiments, would also improve the study.

      For recommendations to the authors, comments from each reviewer are listed below.

      Reviewer #1 (Recommendations for the authors):

      Some additional analysis would clarify the relationship between pPREs and transcription factors.

      Because of the limitations of depleting Pho, the model that nuclear levels of Pho, Cg, and GAF regulated E(z) activity is purely correlative, as GAF depletion did not change H3K27me3 distribution. As it stands, it is possible that this NC14 nuclear enrichment of these factors is not relevant. The statement on lines 295-297 regarding the correlation between re-establishment of the modification state and nuclear localization of Pho, Cg, and GAF is true, but a bit misleading since there is no evidence to support the necessity of these factors for H3K27me3 establishment. Other models remain possible, and the discussion should be toned down to account for this. Furthermore, the only factor that is shown to influence H3K27me3 is Zelda, which does not show an increase in nuclear localization at NC14.

      We thank the reviewer for highlighting this issue. As the reviewer indicates, we lack definitive mechanistic evidence that limited nuclear localization of any recruitment/nucleating factor is limiting for H3K27me3 deposition. However, our data are, we feel, definitive in terms of demonstrating that nucleation from PREs and PRE-like regions arises at mid-NC14 for H3K27me3, and in late cleavages for H2Aub. Ideally, we would test each known nucleating factor for a necessary role in mediating this activity. We have chosen not to include such measurements in this manuscript because a proper mechanistic treatment of each factor (Pho, Cg, and others) would need to be extensive, and we feel better suited for an independent study. We have added language to the “Limitations of the Study” section to reflect the remaining need to identify the key nucleating factors responsible for establishment of the zygotic PcG landscape.

      We re-read the lines the reviewer suggested were misleading and we respectfully disagree. The lines (“Taken together, these observations are consistent with a model where…re-establishment of this modification state is restricted to late cleavage divisions by limiting nuclear localization of nucleating factors such as Pho, Cg, and GAF.”) are expressing a hypothesis/model, and in our opinion these lines are suitably framed as to not be misleading.

      Additional analysis/discussion regarding the relationship between Zelda, H3K27me3, CBP, and paused polymerase would provide further clarity into how Zelda might promote this methylation. It would be useful to discuss how the various classes defined correlate with enhancers versus promoters, and also the gene expression of the underlying gene. It would be clarifying to discuss how the H3K27me3 at these Zelda-dependent regions relates to gene silencing since many of the genes depend on Zelda for expression. Are these genes expressed prior to NC14 and then silenced at this time point? How do H3K27ac levels, which Zelda promotes through recruitment of CBP, relate to the various pPRE classes? The authors should also consider the report from the Mannervik lab that CBP is instrumental in promoting H3K27me3 (Hunt, Boija, Mannervik et al. Mol Cell 2022 82:3580-3597) and the relationship with paused RNA Pol II. Given this possible connection, it could be useful to consider that paused polymerase is also first evident at NC13/14. Overlaying the pPREs with paused polymerase from Chen et al. 2013 eLife (current citation 30) could be informative. At a minimum, a discussion of these additional mechanisms and the implications of the data in Hunt et al. would strengthen the Discussion section.

      We thank the reviewer for this comment. We agree that additional analysis of the relationship between Zelda, H3K27me3 and CBP would be essential for providing mechanistic clarity on the role of Zelda for putting these loci into play for apparently either positive or negative regulation. At present, we feel that extensive additional analysis, including work with CBP, PRC1, and nucleating factors such as Pho would be better suited for a future study.

      H2Aub is clear earlier in development than H3K27me3, and in mice and zebrafish, it promotes PRC2-mediated H3K27me3 (Hickey et al. eLife 2022 doi: 10.7554/eLife.67738, citations 86, 89). As such, it remains possible that this mark is instructive for the H3K27me3 deposition observed. As such, a bit more analysis of where H2Aub is deposited and how it overlaps with pPREs might help determine whether similar mechanisms could be important in Drosophila.

      We suspect that similar mechanisms are important in Drosophila as well. The omission of the Hickey…Cairns reference was an oversight in the original document. We have revised this sentence to refer to both mouse and zebrafish and have added the citation.

      Prior work has noted the sudden increase in GAF nuclear concentration. Please cite Dima and Reeves. Development 2025 152:dev204460 in support of the observations shown in Figure 3D.

      Thank you. Yes, this article was published shortly after we submitted this manuscript for review and we have now added it to reflect its independent replication of the GAF nuclear concentration result.

      It is not clear that Figure 4 warrants an entirely new figure, since the conclusions drawn are similar to/the same as Figure 3. Perhaps change to a supporting figure?

      We agree that this figure was repetitive. We have now moved it to a figure supplement of Figure 3.

      The authors have developed a powerful modified ChIP protocol that enables them to perform the experiment on the equivalent of 10 embryos! This is not highlighted in the manuscript, despite the vast improvement this provides. The authors should highlight this in the manuscript, unless a separate manuscript describing this technique is being written/published. Regardless, this is a very exciting protocol.

      Thank you. A methods paper describing this approach is in preparation.

      Minor:

      (1) Line 49-53: clarify that, as opposed to the mechanisms described earlier in the paragraph, these mechanisms are specific to Drosophila.

      Done.

      (2) Line 61: cite Sun et al. and Schulz et al. (current citations 60 and 61) since these papers demonstrated the pioneering function of Zelda.

      Done.

      (3) Line 95: a word seems to be missing. Perhaps "approach (STAN) to identify a set of PcG "domains" from our"?

      We have made this revision.

      (4) Line 217: Calling an embryo a "specimen" is odd. Can you just say all embryos?

      Ok.

      (5) Line 316: GAF is encoded by Trithorax-like, not Trithorax-related.

      Revised.

      (6) Figure 5: The arrowheads and asterisk are so small that they are nearly impossible to see when the figure is printed. Please make it larger and perhaps use a color that stands out more.

      We have made the arrowheads and asterisk larger and changed the color to yellow to improve visibility.

      Reviewer #2 (Recommendations for the authors):

      (1) Regarding weakness 1, the title states "nucleation-dependent propagation". This is an overstatement of the study's conclusions because these features are inferred and not directly tested here. This is an easy fix with text revisions.

      We acknowledge that we have not directly tested the role of specific nucleating factors for the process of nucleation-dependent propagation. This is reflected in the Discussion text. We have added to the “Limitations of the Study” section the following text: “Finally, we acknowledge that further loss-of-function analysis will be necessary to determine the key nucleating factors responsible for the initial establishment of the zygotic H2Aub and H3K27me3 landscape.” However, we feel that we have demonstrated definitively that, by mid-NC14, nucleation of H3K27me3 sites is first detected genome-wide. As such, we have left the title as-is.

      (2) Also, regarding weakness 1, the abstract makes additional overstatements. These are easy fixes with text revisions.

      (a) "A large subset of targets requires ZLD..." (line 24). This phrase makes it sound like the majority of Polycomb domains depend on ZLD, which I am not sure is accurate.

      Thank you for pointing this out. We agree with this assessment and have revised the abstract to remove the word “large” so that now the sentence reads “; a subset of targets requires Zelda…”

      (b) "...requires ZLD not for PcG factor recruitment" (same sentence). This sentence should state "E(z)" instead of "PcG factor" because E(z) was the only one tested in ZLD mutants.

      We agree as well and have made the requested change to exchange “PcG factor” to “E(z)”.

      (c) "to license a loaded PRE" (same sentence). Whether PREs are fully loaded in ZLD mutants was not tested. In addition, see comments below for feedback on the use of the term "license."

      We have revised this to read “to license an E(z)-loaded PRE,” to keep consistent with the above suggested change. We also removed the mention of “H2Aub” because now the PRE-loading acknowledges we have only measured E(z), although we see effects on both modifications. We acknowledge, in light of Reviewer 3’s comment, that we have not directly measured whether PRC1 still can bind to certain PREs in the absence of Zelda like we see with E(z)/PRC2.

      (3) Also regarding weakness 1, the authors set up a dichotomy for ZLD's role in initiating Polycomb domain formation, either acting as a pioneer or as a licensing factor. However, based on the data presented (e.g. the atonal browser shot), it appears that chromatin accessibility is lost in both sub-classes of H3K27me3 domains that depend on ZLD, meaning that ZLD's role as a pioneer would explain both sub-classes, and that a licensing role is not supported by the data. To assess whether there are pioneer-independent roles of ZLD in the emergence of Polycomb domains, it may help to examine the role of chromatin accessibility changes directly (e.g., is there a substantial fraction of changes in H3K27me3 domains or E(z) peaks that cannot be attributed to changes in chromatin accessibility in ZLD mutants?). The authors should refine their language or provide additional support for the non-pioneering role of ZLD.

      We acknowledge that the pioneer/licensing distinction is still not clear from a mechanistic perspective and that future work will be needed to address this issue. We do not mean to imply that the licensing is necessarily independent of pioneering, rather that at sites like atonal, clearly chromatin accessibility is not the only job that Zelda performs. There are at least two possible mechanisms: 1) it is all about pioneering, and in the absence of accessible chromatin, some other factor required for stimulating H3K27me3 deposition is unable to bind; or 2) it is all about Zelda, and in mutants Zelda both does not confer accessible chromatin, and also does not stimulate H3K27me3 deposition through whatever means. For either mechanism, the remarkable feature is that E(z) still localizes to its genomic target and requires additional information to deposit H3K27me3. This is distinct from the other class (represented by amos in Figure 6 in the final manuscript) where loss of accessibility correlates with loss of E(z) –and presumably PRC2– binding. To explain this activity, we have invoked the term “licensing” which we feel captures the effect of Zelda (whether it be direct or indirect). The possibility of indirectness is a significant caveat, so we have added clarifying text to indicate this possibility. We have added to paragraph 2 of the Discussion the sentence, “As such, we emphasize that Zelda-dependent licensing could stem either directly or indirectly from Zelda function.”

      (4) Regarding weakness 2, as currently written, it seems like a small fraction of H3K27me3 domains depend on ZLD. Is this accurate? Some of my confusion may stem from alternating use of bins and runs. Can the fraction of domains that depend on ZLD be made more explicit through text revisions and additional bioinformatics? More comprehensive bioinformatics analyses can be performed with the ZLD mutant datasets by incorporating ATAC and E(z) peaks. For instance, what fraction of E(z) peaks inside and outside of Polycomb domains are affected in ZLD mutants, and do these E(z) changes correlate with H3K27me3 changes? Similarly, what fraction of PREs change in accessibility in ZLD mutants, and are these accessibility changes correlated with H3K27me3 changes? When do PREs become accessible during wild-type embryogenesis? Are there unique features of ZLD-dependent H3K27me3 domains or E(z) peaks? The authors' perspective on the extent to which ZLD is required for Polycomb domain initiation should also be added to the Discussion.

      We find that 38 PcG domains are sensitive to Zelda (9 have increases, 29 have decreases in H3K27me3), and this is stated in the Results. This accounts for 16% of domains, which is a small fraction of the total. We have added a sentence to the results reporting this fraction. An earlier draft of this manuscript included an analysis of accessibility: both in terms of the timing of when PREs gain accessibility and their dependency on Zelda function. This section was omitted from the submitted manuscript because it did not add clarity the distinction between classes that we report. As for the final request that we add to the Discussion our perspective on the extent to which Zelda is required for PcG domain initiation, we now address this in the second paragraph of the Discussion.

      (5) Regarding weakness 3 (why is H3K27me3 missing pre NC14?), it is suggested that the absence of H3K27me3 is due to the short duration of nuclear cycles relative to the rate of me2->me3 catalysis. And although the authors suggest that Polycomb complex assembly occurs at pPREs without H3K27me3, their microscopy studies indicate that binding of Polycomb proteins to PREs may be regulated via nuclear accumulation. Therefore, it remains unclear whether (and when) the lack of H3K27me3 is due to incomplete Polycomb complex assembly. Additional genomic profiling experiments are needed to directly test when Polycomb complexes assemble on chromatin relative to the emergence of H3K27me3 domains. These experiments would also support claims of nucleation.

      (a) ChIP of E(z) and a PRE binding protein (Pho, Gc, GAF) should be performed at NC13 to test whether Polycomb proteins (especially PRC2) are bound at PREs prior to H3K27me3 emergence.

      (b) ChIP of E(z) should also be performed at NC10 when the PRE binding proteins appear to be absent but when E(z) is hyperabundant in the nucleus. Can E(z) bind PREs in the absence of these "nucleators"? Or does E(z) promiscuously interact with chromatin, helping to explain the broad, low-level H3K27me1 enrichment profile?

      We thank the reviewer for this comment. We have addressed this issue in a separate study (preprinted and accepted for publication as of this writing). Although H3K27me3 is not detectable on cleavage-stage chromatin, lower-order H3K27me2 is maintained on chromatin and detected throughout cleavages by immunostaining and by ChIP. The maintenance of cleavage-stage H3K27me2 depends on both E(z) and Esc. Notably, this lower-order state reflects maintenance of a maternally supplied H3K27 methylation state: H3K27me2 is only detected on maternal (not paternal) chromatin during the period (prior to NC10) when Pho/Cg/GAF do not localize to nuclei. Overall, these observations underscore the limiting nature of early cleavages to support de novo establishment of H3K27 methyl states, but also demonstrate the competency of the PcG system during this time to engage in some degree of H3K27 maintenance.

      Minor Comments:

      (1) It is interesting that the great majority of E(z) peaks (72%) are outside of H3K27me3 domains at NC14. What are these sites? Do these sites correspond to Polycomb domains later in embryogenesis (ie, do they become marked by H3K27me3 later)? Or, do these E(z) peaks disappear at later stages of embryogenesis?

      We agree that this observation is interesting. Figure 2 shows that the majority of these sites are concurrent with transcription start sites. Visual comparison between ChIP datasets generated here and any of the various publicly available datasets generated at later stages (e.g., stage 16 embryos) reveals that the set of PcG domains observed at ZGA is fairly consistent with the set of PcG domains observed later on. Therefore, on a bulk level, these extra-domain E(z) sites observed at ZGA do not represent later-onset PcG domains. However, at this level of resolution, we cannot rule out that in some cell type these sites are converted to a cell-type-specific PcG domain. We have not addressed the perdurance of these E(z) peaks at later stages. The function and the fate of these extra-domain binding sites remains an open question for future investigation.

      (2) Figure 3A live imaging indicates that E(z) nuclear signal intensity diminishes over subsequent nuclear cycles. Can the authors expand on whether photobleaching may contribute to this decrease? The E(z) signal in IF experiments should be quantified to test whether it also diminishes over successive nuclear cycles.

      The imaging conditions for this experiment were controlled to minimize photobleaching in the EGFP channel. While we cannot rule out some small contribution of photobleaching to the overall signal intensity, the overall trend reported in the quantification of Figure 3A reflects, to the best of our abilities to measure, a biological effect. This effect is also evident in the immunofluorescence imaging without need for additional quantification: From NC10 to NC14, when nuclei are presented on the embryo surface, a clear decrease in staining intensity is observed (Figure 3-figure supplement 1A in the final manuscript). Our observations indicate that E(z) decreases in concentration with increasing nuclear content during the cleavage divisions.

      (3) Can the authors provide further interpretation of the H3K27me1 signal profile? There is very little difference in signal amplitude inside relative to outside domains in Figure 1C, and across a 20kb window surrounding E(z) peaks in Figure 1B. Is all this chromatin considered to be H3K27me1-enriched, or is there a high level of noise? The anticorrelation with H3K27me3 at NC14 is compelling and seems to argue that the broad H3K27me1 signal is real.

      The H3K27me1 signal profile reflects a likely broad distribution of H3K27me1 that includes not only canonical “PcG domains” (i.e., regions that ultimately gain high-level H3K27me3) but also non-canonical domains. Consistent with observations in other systems (e.g., PMID: 24289921), H3K27me1 is broadly distributed across the Drosophila genome. We have added the following sentence to the Results section to contextualize our description of the H3K27me1 domains: “The broad genome-wide distribution of H3K27me1, including in regions outside of canonical PcG domains, resembles profiles previously measured in mammalian tissue culture.” (citing the Ferrari et al study).

      Reviewer #3 (Recommendations for the authors):

      My suggestions for changes are mainly of an editorial nature.

      My comments below are not in order of priority, but grouped into different types of suggestions for changes.

      Comments on the presentation of results:

      (1) The relationship between the genomic regions shown in the heat map in Figure 1A and Figure 2A is not clear. In Figure 2A, 1264 regions with E(z) peaks in PcG/H3K27me3 domains are shown. In Figure 1A, the text (line 97/98) states that 237 PcG/H3K27me3 domains contain 1 or more E(z) peaks, and on the left of the H3K27me3 heat map panel, it says "PcG domains". Are these then the 237 regions that are shown in Figure 1A? Only the top ones seem to have high levels of H3K27me3. Please clarify.

      We could have been more clear about this. As the reviewer points out, there are different numbers of peaks shown in the heatmaps in Figures 1A and 2A. Figures 1A and B plot one E(z) peak per domain, determined by finding the one maximal E(z) peak per and plotting it. This is stated in the figure legend: “One representative E(z) peak per domain was selected for plotting.” Figure 2A plots all E(z) peaks within Domains (n = 1264). For the heatmaps and the average plots in Figures 1A and 1B, we found that selecting the maximal E(z) peak yielded a more accurate representation of the accumulation of H3K27me1/3, compared with plotting this for all E(z) peaks within domains, presumably because some called peaks (e.g., many of the minor E(z) peaks shown in Figure 1E) do not appear to be the primary sites of nucleation within the domain. For Figure 2, we wished to compare all E(z) peaks, inside and outside of domains, and for this we plotted heatmaps over all of the peaks.

      (2) In general, in the text, the figure panels should be discussed in order of appearance. For example, it is not ideal that after describing the data in Figure 1A, the text jumps to discuss results shown in Figure 2A and only then goes back to discuss Figure 1B and C. This should be easy to resolve by reorganizing the text or perhaps re-arranging figure panels (e.g., perhaps moving elements to additional supplemental figures?).

      We acknowledge that this is inconvenient and we apologize for this. However, we feel the flow of the manuscript and the figures themselves benefit from this somewhat awkward order of discussion and we have chosen to leave them as-is.

      (3) In general, I felt that the text could be improved by putting some of the more detailed technical procedures and result descriptions into Materials and Methods or Figure legends in order not to disrupt the flow of the text.

      Comments for improving discussion:

      (4) When citing and discussing the previous literature that claimed a function for GAGA factor (GAF/ Trl) in Polycomb repression (i.e., refs. 12, 14, 49-52) on page 15/16 and again on page 25, the authors may also want to cite earlier studies that failed to observe a function of GAF/Trl in Polycomb repression. Specifically, previous studies (Brown et al, Development 2003) had investigated a possible role of GAF/Trl in Polycomb repression in larvae using stringent tests (analysis of HOX gene expression in Trl null mutant cell clones in imaginal discs, removing Trl in a sensitized pho null mutant background, mutation of GAF/Trl binding sites in HOX-LacZ reporter genes) and had found no evidence for a role of GAF/Trl in Polycomb repression in larval tissues. It is fair to say that in most of the studies cited (i.e., references 12, 14, 49-52), the effect of GAF/Trl had been analyzed using PRE-miniwhite reporter gene activity as read-out, a much less stringent assay.

      Thank you for pointing this out. Brown et al (reference 15) was overlooked when entering citations to these sections. We have added a sentence to the discussion to highlight the lack of necessity for Trl/GAF for imaginal disc silencing of Hox targets.

      (5) The authors report that removal of GAF/Trl had almost no impact on H3K27me3 domain formation. Although this supports a model where PRC2 recruitment to PREs is unaltered after GAF/Trl depletion, it does not eliminate the potential caveat that PRC1 binding may be affected. Do the authors have binding profiles of PRC1 subunits or the H2Aub1 profile in embryos lacking GAF/Trl? It seems that the current paper would be a great opportunity to include such data and get them published. There is no need to generate PRC1 or H2Aub1 profiles if they don't already exist.

      Unfortunately, we do not as of yet have any PRC1 reagents that are compatible with ChIP. Following the lack of effect of GAF on H3K27me3, we did not pursue the measurements of H2Aub in these knockdown conditions, given the likely dependence of H3K27me3 on H2Aub deposition at ZGA.

      (6) Previous studies found that Polycomb group protein complexes bind to PREs in HOX genes both in cells where genes are OFF but also in cells where genes are ON but that the H3K27me3 profile is different in the two states, with H3K27me3 decorating the gene in the OFF but not in the ON state (Papp and Müller, Genes Dev 2006; Bowman et al, eLife 2014). In this study, the H3K27me3 profiles at target genes represent the sum of ChIP signals coming from cells where the gene is OFF and cells where the gene is ON. Is it known how eve and zen are deregulated in Zelda-depleted embryos? Could it be that the increased H3K27me3 enrichment at eve and zen is an indirect effect caused by loss of eve and zen expression in a large fraction of cells and consequently a gain of H3K27me3 ChIP signal at the gene in those cells? It may be worth at least discussing such scenarios in light of the studies mentioned above, and referring to them.

      Thank you for raising this question. We have a manuscript in preparation that specifically addresses this issue, namely the relationship of on/off states to H3K27me3 deposition in the early embryo. In short, it is likely that the increase in H3K27me3 at eve reflects significantly reduced eve expression in zelda mutant embryos. We have added a clarifying sentence to the Discussion and added the indicated references.

    1. Author response:

      We would like to thank the editorial team and the reviewers for their thoughtful assessment of our manuscript. We are highly encouraged that the reviewers found our closed-loop OMR virtual reality assay to be a valuable new behavioural paradigm, and that our findings regarding the early behavioural phenotypes in adgrl3.1 mutant zebrafish provide solid, high-quality data of broad interest to the community. We also appreciate your constructive feedback regarding the structural flow of our manuscript, the need for tighter conceptual framing, and the request for more rigorous methodological explanation and improved presentation. In our upcoming revision, we plan to directly address the suggestions raised to ensure clarity of our work.

      For our revised manuscript, we will ensure to focus on fully contextualising our conceptual rationale and unique utility of our behavioural assay, ensuring a clear link to clinical relevance. This means we will provide further clarity on the rationale and relevance of the paradigm and interpretations of behaviours, while ensuring we are not overstating the absolute separation of anxiety-like states from hyperactivity and contextualising our findings within the broader literature. In addition, as suggested by the reviewers, we will improve the Abstract to explicitly articulate the unique advantages of the closed-loop OMR virtual reality paradigm over standard static assays and provide definitions. In addition, we will also revise our interpretations in the Discussion, such as around swim speed, exploration, and visual sensitivity, to ensure they are fully grounded and supported with clear flow.

      Equally important to our revision is to clarify genetic validation, methods and statistical rigour as recommended by the reviewers. Therefore, we will update the Methods to explicitly detail further husbandry details, such as the embryo medium used, breeding protocols, and the exact timings in hours post-fertilisation. To resolve any ambiguity surrounding our genetic model, we will provide further explanations and include details on genotype and sequencing protocols. Where needed, we will also update the statistical analysis and expand the Methods accordingly to detail all statistical reporting for clarity.

      Finally, as suggested, we will also improve the overall structure, flow, and accessibility. In the Results, we will ensure the core behavioural phenotypes take primary focus throughout the writing. Furthermore, we will also improve the Discussion to ensure it is a unified, cohesive narrative that smoothly integrates our main behavioural findings with genetic model validation, limitations, and future directions. Additionally, we will update all visual presentations, figure formatting, and labelling to ensure accessibility and improved readability.

      We are incredibly grateful to the reviewers for their insightful recommendations. We are confident that by integrating their revisions, we will substantially strengthen the clarity of the work and better demonstrate its impact.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors report the results of a tDCS brain stimulation study (verum vs sham stimulation of left DLPFC; between-subjects) in 46 participants, using an intense stimulation protocol over 2 weeks, combined with an experience-sampling approach, plus follow-up measures after 6 months.

      Strengths:

      The authors are studying a relevant and interesting research question using an intriguing design, following participants quite intensely over time and even at a follow-up time point. The use of an experience-sampling approach is another strength of the work.

      Comments on revisions:

      Overall, I think the authors made many improvements to their manuscript. There are, however, still a number of concerns that first need to be addressed, since it is still not currently possible to fully evaluate the analyses, results, and conclusions presented in the paper. I list these points below:

      (1) The authors still use causal language where they must not use causal language. This is true for many places in the manuscript; I am highlighting here just a few places, but the authors nevertheless have to go carefully through the whole manuscript to change these instances.

      We sincerely thank the reviewer for this critical and well-taken point. We fully agree that our design (manipulating DLPFC excitability while measuring procrastination, task value, and aversiveness) does not directly measure or manipulate self-control, nor does it rule out alternative neurocognitive mechanisms. Accordingly, we have conducted a comprehensive, line-by-line revision of the entire manuscript to systematically replace causal claims with cautious, hypothesis-consistent language. In response, we have replaced all the wordings that may imply causal inferences, such as impair, cause, boost, by association-consistent phrasing. Furthermore, as you clearly raised below, we explicitly reframed self-control as a hypothesized theoretical construct rather than an empirically verified mediator throughout the whole re-revised manuscript. Please see specific revisions below:

      Abstract Section (Page 2, Line 53-59)

      “... a mediation analysis indicated a disassociable mechanism: the increase in task outcome value (but not task aversiveness) showed a statistical pattern consistent with accounting for the observed behavioral improvement. In conclusion, these findings are consistent with the hypothesis that enhancing DLPFC function may reduce procrastination by selectively amplifying the valuation of future rewards, not by simply reducing negative feelings about the task.”

      Introduction Section (Page 3, Line 83-86)

      “... Even worse, chronic procrastination has been consistently associated with poor general health conditions, such as immune system disruption, gastrointestinal disturbance, hypertension and cardiovascular disease (Sirois, 2015; Sirois, 2016).”

      Introduction Section (Page 4, Line 143-145)

      “Consistent with this framework, the left dorsolateral prefrontal cortex (DLPFC)—a region frequently implicated in value-based decision-making and top-down regulation—has been associated with procrastination. ...”

      Introduction Section (Page 4, Line 135-139)

      “... Also, given the integrative nature of prefrontal regulatory functions, we hypothesize a third pathway whereby both decreased task aversiveness and increased task-outcome value may jointly contribute to reduced procrastination, potentially reflecting coordinated downstream effects on valuation and affective processing.”

      Introduction Section (Page 5, Line 183-186)

      “... Thus, this study aims to clarify the brain-behavior association between DLPFC neuromodulation and procrastination, and to test whether observed changes in task valuation and aversiveness are consistent with theoretical models of top-down regulation.”

      Results Section (Page 11, Line 534-536)

      “... Thus, these findings are consistent with the view that neuromodulation of the left DLPFC is associated with reduced task aversiveness and increased task-outcome value.”

      Results Section (Page 11, Line 579-584)

      “... In summary, these findings identified a statistical pathway consistent with the theoretical model: neuromodulation of the left DLPFC was associated with increased task-outcome value, which in turn was associated with reduced procrastination.”

      Discussion Section (Page 12, Line 617-620)

      “On balance, our findings provide evidence consistent with the hypothesis that neuromodulation of the left DLPFC is associated with reduced procrastination, primarily through increasing task-outcome value rather than merely reducing task aversiveness. ...”

      Discussion Section (Page 13, Line 664-669)

      “... Building on this foundation, among several theoretical interpretations and cognitive pathways, our study showed the one plausible neurocognitive mechanism of procrastination: the cortical excitability of the DLPFC produced by active neuromodulation may engage prefrontal regulatory networks to increase task outcome value, which in turn is associated with reduced procrastination behavior, statistically supporting the theoretical accounts of temporal decision model (TDM, Zhang et al., 2019).”

      Discussion Section (Page 14, Line 755-762)

      “... Moreover, this study did not collect data for assessing participants' self-control at either baseline or post-neuromodulation. Accordingly, we explicitly note that self-control was not directly measured or manipulated in this study; the observed associations between DLPFC neuromodulation, task-outcome value, and procrastination are consistent with theoretical models positing a role for top-down regulatory processes, but do not constitute direct evidence that self-control mechanisms were engaged. This limitation precludes definitive conclusions about the unique contribution of self-control-related pathways versus alternative neurocognitive mechanisms.”

      Some examples:

      (a) In response to my comment (1) in the previous round, where the authors adjusted their text, the authors still use causal language in their last sentence "... procrastination behavior has been observed to impair general health..." Unless the cited study truly allowed causal conclusions, the causal language should be removed here as well.

      Thank you for pointing out this inappropriate phrasing. As you kindly suggested, we have reworded it as “Even worse, chronic procrastination has been consistently associated with poor general health conditions, such as immune system disruption, gastrointestinal disturbance, hypertension and cardiovascular disease” (Introduction Section, Page 3, Line 83-86).

      (b) The authors still make (causal) claims about the involvement of self-control in their observed results. To reiterate from the previous round of revisions: The authors cannot make any strong claims about the role of self-control processes because they do not directly measure self-control nor do they directly manipulate self-control or have a design that would rule out alternative mechanisms other than self-control. Therefore, their claims about self-control have to be toned down. It is laudable that the authors have added a statement towards the end of their discussion about not being able to make strong conclusions about the role of self-control. But the authors need to use similar careful wording not just at the end of the discussion but throughout the manuscript.

      We appreciate you reiterating this concern. In the re-revised manuscript, we have thoroughly removed or rewritten all the statements implying causal inferences, and have substantially toned-down claims for the roles of self-control in procrastination reduction from this neuromodulation. Please see instances for what we have replied to the Comment #1.

      (i) In the abstract, the authors use the formulation "...conceptualized roles of self-control on procrastination..." -- this wording is still too strong, suggesting that you actually studied self-control.

      Thank you for providing this specific instance. This inappropriate sentence has been removed.

      (ii) In the introduction (page 4, lines162-169), the way the authors formulate these sentences suggests that they directly measured self-control. Again, the authors need to make it explicit that they are not directly measuring self-control but its hypothesized down-stream consequences on valuations/behavior.

      Many thanks. This statement has been removed, and we have reworded it as “... Thus, this study aims to clarify the brain-behavior association between DLPFC neuromodulation and procrastination, and to test whether observed changes in task valuation and aversiveness are consistent with theoretical models of top-down regulation.” (Introduction Section, Page 5, Line 183-186).

      (iii) In the discussion, for example, on page 11, lines 555 and following, the authors write: "One major contribution this study has made is to disentangle the neurocognitive mechanism of procrastination by demonstrating that self-control could increase task-outcome value so as to reduce procrastination."

      As you kindly instructed, we have rewritten this statement as “One contribution of this study is to provide empirical evidence partially consistent with the temporal decision model (TDM), showing that increased task-outcome value—rather than decreased task aversiveness—was statistically associated with reduced procrastination following DLPFC neuromodulation.” (Introduction Section, Page 12, Line 625-628), which no longer implies any conclusions for the role of self-control in this study.

      Again, please be aware that you are NOT demonstrating that self-control does anything, since you only measure procrastination rates, outcome values, and task aversiveness. It is possible that mechanisms other than self-control might be relevant for this. Perhaps neuromodulation directly increases outcome values, without involvement of self-control processes. You simply cannot know that and therefore you cannot make those claims in the form that you are making them. You can write that the observed results are consistent with the idea that neuromodulation might have had an effect on self-control and this in turn might have affected outcome values. But you also need to make it explicit that, to substantiate these claims, you would need more direct evidence that indeed self-control was involved. These more careful formulations would not at all reduce the value of your work, but indeed they would rather demonstrate your carefulness in interpreting the results you obtained.

      We sincerely thank the reviewer for this exceptionally clear and constructive guidance. We fully agree that our study design does not measure or manipulate self-control, and therefore we cannot demonstrate that self-control processes are causally involved in the observed effects. As you correctly note, it is entirely possible that neuromodulation directly modulates outcome valuation or engages alternative neurocognitive pathways (e.g., attentional allocation, feedback learning, or affective processing) without invoking self-control mechanisms.

      In direct response, as we replied above, we have completely rewritten the whole revised manuscript to remove any assertions that we “identified” a role of self-control. The revised text now explicitly states as follow: (1) our findings merely are consistent with the theoretical hypothesis that DLPFC neuromodulation might engage prefrontal self-regulatory functions, which in turn influence outcome valuation; (2) we explicitly acknowledge that substantiating this specific pathway would require more direct evidence. Rather than single sentence, we have applied this careful, hypothesis-consistent framing systematically across the Abstract, Introduction, Results, and Discussion. As you suggested, these revisions more accurately reflect the interpretative boundaries of our data and demonstrate our commitment to rigorous, transparent scientific reporting. Please see specific cases for this revision above.

      (2) I am still puzzled by the power analysis. In the text, you write that a sample size of 18 participants (i.e., 9 per group) would be sufficient to achieve 80% power. I still feel this seems far too optimistic and hard to believe, but that is not my point here. While in the text, you write that you need 18 participants, the G*power output seems to suggest a sample size of 34, not 18. Why this contradiction? Or is it not contradictory? If it is not, then please explain it more fully.

      We appreciate you pointing out this critical typo. In the last round of revision, we mean that 18 participants per group are required to achieve at least 80% statistical power, rather than a total sample size, as shown by the GPower software. We are sorry for this critical typo to confuse you. As you correctly pointed out, the GPower indicated that the minimum sample size to reach 80% power is 34 (i.e., 17 per group). Thus, we selected 36 (i.e., 18 per group) participants as minimum sample size in case of potential drop-out. We have thoroughly corrected this typo, and double-checked no such numeric issues:

      Methods Section (Page 5, Line 234-237)

      “... statistical power was predetermined by G*Power at a relatively medium effect size (1-β err prob = 0.80, f = 0.25), indicating the total sample size at 34 (17 per group) to reach acceptable power. To account for potential attrition, we determined to recruit 36 participants, at least.”.

      (3) I have several comments about the mixed-effects analysis.

      First of all, I want to thank the authors for adding more details, things have become much clearer now. However, I still have a few questions and comments related to these analyses:

      (a) The variable Emotions was within-subjects, as far as I understood. Accordingly, Emotions should most likely be modelled with random slopes varying over participants (in addition to being modelled as a fixed effect).

      We thank you raising this reasonable concern on the mixed-effect linear modeling. Yes, the Emotions reflect daily baseline affect, which is modeled as covariates of no interests to adjust for daily emotional fluctuation (if any). In this vein, this baseline emotion score is included for each participant across all the sessions. Therefore, it should be modeled with random slopes as you assumed indeed.

      In response, we have remodeled this mixed-effects analysis by including the daily baseline emotion as random slopes varying over participants. After centering the variables, we estimated the revised models as “Procrastination Rate ~ Group * Treatment day + Age + Gender + SES + Emotions + (1 + Treatment day + Emotions || SubjectID)” and “Task execution willingness ~ Group * Treatment day + Age + Gender + SES + Emotions + (1 + Treatment day + Emotions || SubjectID)”. Notably, as you correctly assumed, fitting this complicated random-effect structure is likely to result in convergence failure, given the limited sample size in the present study. Therefore, we hypothesized the independence among random effects for model simplification. Consistent with this assumption, model comparisons indicated that the simplified models fit better than original ones (Procrastination Rate model, ∆AIC = -1.0, ∆BIC = -11.8, LRT, χ<sup>2</sup>(3) = 5.04, p = .17; Task-execution willingness model, ∆AIC = -5.5, ∆BIC = -16.4, LRT, χ<sup>2</sup>(3) = 0.51, p = .91).

      Taken together, as you kindly suggested, we have rebuilt the mixed-effect models by adding daily baseline emotion as random slopes varying over participants, and have demonstrated the consistent findings with the original one:

      Methods Section (Page 8-9, Line 412-419)

      “... Given the risks of convergence failure with the two correlated random-effects structure (i.e., treatment days and self-reported emotions), we hypothesized that the random effects are independent, leading to model simplification. Consistent with this assumption, model comparisons favored the simplified independent structure over the full correlated structure for both outcomes. For the actual procrastination model, the simplified model showed lower AIC (∆ = -1.0) and BIC (∆ = -11.8), with a non-significant likelihood ratio test (χ<sup>2</sup> (3) = 5.04, p = .17). For the task-execution willingness model, the simplified model was also preferred (∆AIC = -5.5, ∆BIC = -16.4; LRT: χ<sup>2</sup> (3) = 0.51, p = .91).”

      Results Section (Page 9-10, Line 469-489)

      “For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001; Fig. 3A). In the post-hoc simple effect analysis, it demonstrated a significantly increased task-execution willingness (i.e., decreased procrastination willingness) after neuromodulation in the active neuromodulation group (NM-before: 35.65 ± 30.21, NM-after: 80.43 ± 19.92, Mean Diff = 41.79, SE = 7.58, DF = 103.4, t.ratio = 5.51, p < .0001, Tukey correction), but no such effects were identified in the sham control group (SC-before: 37.57 ± 26.46, SC-after: 47.35 ± 30.49, Mean Diff = 2.58, SE = 7.56, DF = 96.8, t.ratio = 0.34, p = .73, Tukey correction) (Fig. 3B-C). A linear uptrend for task-execution willingness was further observed across multiple sessions in the active NM group, indicating gradually increasing neuromodulation effects (Fig. 3D; p < .01, Mann-Kendall test). For actual procrastination behavior, changes to actual procrastination rates across all the sessions have been detailed in the Fig. 3E. Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group (NM-before: 56.74 ± 39.10, NM-after: 0.00 ± 0.00, Mean Diff = 44.40, SE = 9.36, DF = 110.0, t.ratio = 4.74, p < .0001, Tukey correction), but no such prominent changes found in the sham control group (SC-before: 46.47 ± 40.76, SC-after: 33.35 ± 37.82, Mean Diff = 7.53, SE = 9.28, DF = 102.0, t.ratio = 0.81, p = .42, Tukey correction) (Fig. 3F-G).”

      (b) The analyses still cannot fully be evaluated as I cannot access the scripts and data. The authors mention that the scripts and data should be available via a link they provide (https://doi.org/10.57760/sciencedb.35140). However, when I try to access these materials via this link, no page opens; it seems the link is dead?

      Thank you very much for bringing this case to us. We checked this link and found it to be still active.

      To ensure accessibility for your evaluation, we have uploaded scripts and data into this online submission system. Please do let us know if you are still unable to access them. We are glad to send them to you by other available pathways. This link is a private access to you, and the repository would be openly available for other users upon the final publication.

      (c) What are the results and conclusions if you do not include the covariates of no interest? I.e., please re-run your main models without age, gender, SES, Emotions.

      Thank you for raising this question. As you clearly instructed, we have rerun main models without all those covariates. As shown in the table below, the results for the key predictors of interest (Group, Treatment day, and their interaction) remained largely unchanged in terms of effect size, direction, and statistical significance:

      Author response table 1.

      Comparison to statistics derived from model with covariates (i.e., age, gender, SES, Emotions) and without covariates

      (d) The authors mention that they use GLMMs, which would suggest generalized mixed-effects models, but they do not describe what family/distribution they used. Since they mention lmerTest and seem to report F-tests, my guess is that they used Gaussian models. However, both their DVs (procrastination rates and their ratings) are bounded variables and at least procrastination rates hit the lower boundary. That can mean that their analyses suffer from inflated Type 1 and/or Type 2 rates. Therefore, please repeat the analyses with an appropriate generalized mixed-effects model (perhaps a beta regression type of model?).

      We are very grateful to you for raising this crucial statistical point. As you correctly pointed out, we used the Gaussian distribution in estimating this model. We are sorry to confuse you due to the absence of reporting family/distribution we used. In the original manuscript, we meant “general” linear mixed-effect model, rather than “generalized” one. As you clearly and correctly raised, procrastination rates and willingness are technically bounded, and that procrastination rates frequently reached the lower boundary (0%) in the present study, which are in high risks to be inflated for Type 1 and/or Type 2 error.

      Thus, as you kindly suggested, a beta family distribution with logit function is used to reanalyze those main effects of interest. Results are tabulated in Author response table 2.

      Author response table 2.

      These convergent results confirm that the critical main effect (i.e., Group and Treatment day) and their interaction remain statistically significant across distributional specifications, and that our primary conclusions are not artifacts of the Gaussian assumption. Taken them together, as you kindly suggested, we have repeated the analyses with beta regression family distribution, which replicated our main findings, potentially supporting their statistical robustness.

      Following your suggestion, we have added those results derived from such sensitivity analyses into the revised manuscript:

      Methods Section (Page 9, Line 430--436)

      “... To examine whether our findings were sensitive to the distributional assumptions of the dependent variables, we re-analyzed the main models using an alternative distributional specification. Given that both procrastination rates (ranging from 0% to 100%) and task-execution willingness (measured on a 0-100 visual analog scale) are bounded continuous outcomes, and that procrastination rates frequently reached the lower boundary (0%) in the present study, a Beta regression model with a logit link function was employed for a sensitivity analysis.”

      Results Section (Page 10, Line 506-512)

      “... Furthermore, as a sensitivity analysis, we reran the main LMMs using Beta regression distribution with a logit link function, which is appropriate for the both bounded outcomes mentioned above (i.e., procrastination rate and procrastination willingness). The main effects (i.e., Group and Treatment day) and their interaction remained significant for both procrastination rate and willingness (see SI Results and Tab. S5), confirming that our findings are robust to alternative distributional assumptions.”

      (e) When reporting the results of the mixed-effects models, the authors report the regression coefficient, standard error, DFs and p value, but not the actual test statistic. Please add the information about the test statistic and report all degrees of freedom (in case of F tests that would be the degrees of freedom of the test and the residual degrees of freedom).

      We truly thank you for this nuanced reminder. As you suggested, we have added actual test statistics, including t-values and all degrees of freedom (DF). Please see specific instances below:

      Results Section (Page 9-10, Line 469-489)

      “For procrastination willingness, results showed a statistically significant interaction effect between multi-session neuromodulations and groups (β = -7.84, SE = 1.80, t = -4.36, DF = 45.6, p < .001; Fig. 3A). In the post-hoc simple effect analysis, it demonstrated a significantly increased task-execution willingness (i.e., decreased procrastination willingness) after neuromodulation in the active neuromodulation group (NM-before: 35.65 ± 30.21, NM-after: 80.43 ± 19.92, Mean Diff = 41.79, SE = 7.58, DF = 103.4, t.ratio = 5.51, p < .0001, Tukey correction), but no such effects were identified in the sham control group (SC-before: 37.57 ± 26.46, SC-after: 47.35 ± 30.49, Mean Diff = 2.58, SE = 7.56, DF = 96.8, t.ratio = 0.34, p = .73, Tukey correction) (Fig. 3B-C). A linear uptrend for task-execution willingness was further observed across multiple sessions in the active NM group, indicating gradually increasing neuromodulation effects (Fig. 3D; p < .01, Mann-Kendall test). For actual procrastination behavior, changes to actual procrastination rates across all the sessions have been detailed in the Fig. 3E. Similarly, a statistically significant interaction effect was identified here (β = -7.37, SE = 2.40, t = -3.02, DF = 46.6, p = .004), and the simple effect analysis further revealed decreased actual procrastination rates after ms-tDCS in the active neuromodulation group (NM-before: 56.74 ± 39.10, NM-after: 0.00 ± 0.00, Mean Diff = 44.40, SE = 9.36, DF = 110.0, t.ratio = 4.74, p < .0001, Tukey correction), but no such prominent changes found in the sham control group (SC-before: 46.47 ± 40.76, SC-after: 33.35 ± 37.82, Mean Diff = 7.53, SE = 9.28, DF = 102.0, t.ratio = 0.81, p = .42, Tukey correction) (Fig. 3F-G).”

      (f) Thank you for adding the analysis where you remove the last two sessions. But currently you present them in the manuscript without explaining/motivating why you do this. Please add this motivation, as otherwise it will be puzzling for the reader why you conduct these analyses.

      Thank you for this very practical requirement to clarify the motivation of reanalyzing main models from removing the last two sessions. Please see specific explanation as follow:

      Results Section (Page, Line 499-506)

      “... To systematically test whether such effects are biased by extreme data points or patterns, we reran the main LMMs by iteratively removing data from the last two sessions, which showed extraordinarily high effectiveness from neuromodulation (e.g., all the participants in the neuromodulation group had no actual procrastination behavior in session #6 and #7). Results showed the significant group*neuromodulation sessions interaction effects across all those nested models (removing session #6, #7 or both, all p < .05; see SI Results and Tab. S3-4), potentially indicating a statistical robustness from the data pattern.”

      (4) Mediation analysis

      In your manuscript, you present some mediation analyses. Please be aware that such mediation analyses cannot establish causality and they suffer from extremely high Type 1 error rates (see, e.g., https://datacolada.org/103). My suggestion would be to completely remove all mediation analyses. However, if you want to keep them, then you need to be extremely careful in how you present the results. You need to explicitly mention that you cannot derive any causal conclusions from them and that simulation studies have shown that such mediation analyses suffer from extremely high Type 1 errors.

      We sincerely thank you for this exceptionally important methodological guidance. We fully agree that mediation analyses, especially those based on observational measures rather than experimentally manipulated mediators, cannot establish causal pathways and are susceptible to inflated Type 1 error rates, as rigorously demonstrated in recent simulation studies (https://datacolada.org/103).

      As you kindly suggested, please allow us to retain those mediation analyses upon explicitly highlighting that this mediation statistical model cannot generate any causal conclusions. In response, we have systematically replaced all instances of “causal mediation” with “statistical mediation” or “exploratory mediation analysis”, and removed causal verbs (e.g., “depends on”, "drives", "explains") in favor of association-consistent phrasing (e.g., “is statistically mediated”, “aligns with the hypothesis that”) throughout the abstract, introduction, methods, results and discussion sections. Furthermore, in the Discussion section, we explicitly reiterated the limitations of extending this mediation associations to causal conclusions. Please see specific modifications underneath:

      Abstract Section (Page 2, Line 52-56)

      “... While the intervention is significantly associated with both decreased task aversiveness and increased perceived task outcome value, a mediation analysis indicated a disassociable mechanism: the increase in task outcome value (but not task aversiveness) showed a statistical pattern consistent with accounting for the observed behavioral improvement.”

      Methods Section (Page 9, Line 448-453)

      “... As these mediation analyses are based on observational measures rather than experimentally manipulated mediators, they do not establish causal pathways. Simulation studies have shown that such analyses can suffer from inflated Type 1 error rates. Results should therefore be interpreted as hypothesis-generating and statistically consistent with the proposed theoretical model, rather than as confirmatory evidence of causal mechanisms.”

      Methods Section (Page 9, Line 438-440)

      “... the Quasi-Bayesian mediation analysis was used to model the association between the effects of tDCS, task aversiveness/outcome and decreased procrastination.”

      Results Section (Page 11, Line 568-573)

      “As an exploratory analysis, results indicated that increased task outcome value was associated with changes in the task-execution willingness (δ = 21.73, p < .01; ζ = 11.25, p = .07, ρ = 32.99, p < .01, simulation = 1,000; see Fig. 5A) and real-world procrastination (δ = 30.75, p < .01; ζ = 3.05, p = .52, ρ = 33.81, p < .01, simulation = 1,000; see Fig. 5B), in the context of ms-tDCS neuromodulation. ...”

      Results Section (Page 11, Line 582-584)

      “... Nevertheless, all mediation findings are now explicitly labeled as “exploratory” and framed as quantifying statistical associations consistent with the TDM pathway, not causal mediation.”

      Discussion Section (Page 14, Line 747-755)

      “Notably, we explicitly acknowledge that exploratory Quasi-Bayesian mediation analyses, based on observational measures rather than experimentally manipulated mediators, cannot establish causal pathways and are susceptible to inflated Type 1 error rates as demonstrated in recent simulation studies. These findings should be interpreted strictly as hypothesis-generating and statistically consistent with the proposed theoretical model, rather than as confirmatory evidence of causal mechanisms. Substantiating the precise neurocognitive pathway will require future studies employing stronger causal designs, such as experimental manipulation of task valuation or longitudinal cross-lagged modeling. ...”

      As an example (but the mediation results are mentioned in several places, for example, also in the abstract): On page 10, lines 501-503: What you can causally conclude is that neuromodulation affects your measured variables (outcome values, procrastination rates, task aversiveness), but you cannot conclude that the effect of neuromodulation on procrastination rates causally operates via outcome values. Thus, please adjust the formulation accordingly. The same applies to the mediation section that follows right afterwards (page 10, lines 505-522).

      Thank you for offering those specific instances. As we replied above, they have been revised accordingly:

      Results Section (Page 11, Line 557-559)

      “... Collectively, these findings provide statistical evidence consistent with the hypothesis that the outcome-value pathway may contribute to procrastination reduction.”

      Results Section (Page 11, Line 563-582)

      “Increased task outcome value is specifically associated with reduced procrastination in the context of neuromodulation

      To explore the potential neurocognitive pathways of procrastination, the Quasi-Bayesian mediation analysis was undertaken, with increased task outcome value as a statistically mediated variable. As an exploratory analysis, results indicated that increased task outcome value was associated with changes in the task-execution willingness (δ = 21.73, p < .01; ζ = 11.25, p = .07, ρ = 32.99, p < .01, simulation = 1,000; see Fig. 5A) and real-world procrastination (δ = 30.75, p < .01; ζ = 3.05, p = .52, ρ = 33.81, p < .01, simulation = 1,000; see Fig. 5B), in the context of ms-tDCS neuromodulation. To ensure the statistical robustness and specificity of these findings, the sensitivity analysis was implemented by changing sampling parameters and outcome variables. By doing so, those findings were validated statistically robust, as shown by replicated observations across bootstrapping sampling subsets (see SI Results and Tab. S6-7). Moreover, the results of the control analysis further validated the specificity of these findings by showing a null statistically mediated effect of this model to predict one’s task aversiveness (see SI Results and Tab. S8). In summary, these findings identified a statistical pathway consistent with the theoretical model: neuromodulation of the left DLPFC was associated with increased task-outcome value, which in turn was associated with reduced procrastination.”

      (5) In the introduction, the authors introduce several theoretical procrastination frameworks (TMT, mood repair, TDM). Do the results of the current paper help to decide which framework might be the most appropriate, at least for the authors data set? It might be of interest to address this explicitly.

      We do thank the reviewer for this insightful theoretical question. We agree that explicitly positioning our findings within the broader theoretical landscape strengthens the conceptual contribution of our work. Upon careful consideration, we believe that our results provide the strongest empirical support for the TDM over alternative frameworks (TMT, mood repair). TDM uniquely posits procrastination as contingent on the dynamic trade-off between task aversiveness and task-outcome value. In the present study, neuromodulation was identified to be associated with both pathways but only increased outcome value statistically predicted reduced procrastination. This aligns precisely with TDM’s hypothesis that value-based processes may dominate aversiveness-avoidance processes in driving behavioral change. Neither TMT (which emphasizes temporal discounting of utility per se) nor the mood repair perspective (which prioritizes short-term affect regulation) explicitly predicts this dissociable pattern. As you kindly suggested, we have extended the Discussion section to explicitly contextualize our findings into this theoretical landscape (Discussion Section, Page 13, Line 669-687).

      (6) The language is sometimes hard to understand and seems in quite some places grammatically incorrect. Thus, I think the paper would profit very much from thorough English proofreading.

      We sincerely thank you for this practical suggestion. We fully agree that the original manuscript contained grammatical inaccuracies and awkward phrasing that could hinder readability. In response, we have engaged a professional academic editing service to thoroughly proofread and polish the entire manuscript. All sentences have been revised for grammatical correctness, syntactic clarity, and academic tone, while strictly preserving the original scientific meaning and technical terminology. We believe this language improvements have substantially enhanced the readability and overall quality of the paper.

      Reviewer #2 (Public review):

      Summary:

      Chen and colleagues conducted a cross-sectional longitudinal study, administering high-definition transcranial direct stimulation (HD-tDCS) targeting the left DLPFC to examine the effect of HD-tDCS on real-world procrastination behavior. They find that seven sessions of active neuromodulation to the left DLPFC elicited greater modulation of procrastination measures (e.g., task-execution willingness, procrastination rates, task aversiveness, outcome value) relative to sham. They show that HD-tDCS reduces task aversiveness and increases task-execution willingness on real-world tasks as quantified by intensive experience sampling methods, providing causal evidence for the role of DLPFC in modulating contextual features to delaying or completing one's goals.

      Strengths:

      • This is a well-designed protocol with rigorous administration of high-definition transcranial direct current stimulation across multiple sessions. The intensive experience sampling approach which probes and assesses self-relevant task goals is innovative and aims to address an important question regarding the specific role of DLPFC in modulating specific features of chronic procrastination behavior (e.g., task-execution willingness, task aversiveness).

      • The quantification of task aversiveness through AUC metrics is a clever approach to account for the temporal dynamics of task aversiveness, which is notoriously difficult to quantify.

      Weaknesses:

      • While the findings that neurostimulation reduces procrastination behavior is compelling, there remain several alternative interpretations for these effects. For example, it could be that the task-execution willingness isn't increased per se, but rather that the goal completion becomes more valuable as participants learn from feedback or become more aware of their successful attainment of or failure to complete task goals. It is unclear whether the effects could be driven by improved working memory or attention to the reported tasks (and this limitation is addressed by the authors). In short, it is also difficult to examine the temporal dynamics of how these goals are selected across time.

      We sincerely thank you for raising these thoughtful and methodologically important points. We fully agree that the observed reductions in procrastination could reflect multiple neurocognitive pathways beyond the value-based mechanism emphasized in our primary analysis.

      In response, we have thoroughly removed claims on the “unique mechanism” of value-based pathways to procrastination reduction, and fully substituted language implying exclusive mediation by “value amplification” with more cautious phrasing (e.g., “statistically consistent with a value-based pathway”; “one plausible mechanism among several processes”). In the revised manuscript, we reiterated that the pattern of results, showing increased outcome value predicting reduced procrastination while decreased aversiveness did not, aligned with the Temporal Decision Model, yet does not rule out concurrent contributions from attention, learning, or executive processes. To clearly bring this interpretative boundary of our primary findings for audiences, we have explicitly warranted such cautions in the Discussion Section. Please see specific modifications underneath:

      Discussion Section (Page 13, Line 664-669)

      “... Building on this foundation, among several theoretical interpretations and cognitive pathways, our study showed the one plausible neurocognitive mechanism of procrastination: the cortical excitability of the DLPFC produced by active neuromodulation may engage prefrontal regulatory networks to increase task outcome value, which in turn is associated with reduced procrastination behavior, statistically supporting the theoretical accounts of temporal decision model (TDM, Zhang et al., 2019). ”

      Discussion Section (Page 13, Line 676-687)

      “... Despite statistically supporting the TDM, we acknowledge that alternative neurocognitive mechanisms could contribute to the observed reductions in procrastination. For instance, repeated exposure to the experience-sampling protocol may have enhanced participants’ awareness of task progress or facilitated feedback-based learning, thereby increasing the subjective value of goal completion independent of DLPFC neuromodulation. Similarly, improvements in working memory for task maintenance, attentional allocation to reported goals, or strategic shifts in goal selection across sessions could plausibly mediate the intervention effects. While our double-blind, sham-controlled design and inclusion of daily emotional covariates help mitigate some non-specific confounds, the present study did not incorporate direct measures of these alternative processes. Consequently, we cannot definitively isolate the value-based pathway posited by the TDM from concurrent contributions of attention, learning, or executive functions.”

      • It is unclear whether the current evidence support long-retention of this neurostimulation intervention. The study includes one 6-month timepoint after the study to examine the long-term retention of the neural stimulation effect. Future studies that evaluate the long-term effects across multiple time points would strengthen the evidence for the robustness of this intervention.

      We genuinely appreciate you for this insightful and methodologically reasonable point. We fully agree that a single 6-month follow-up assessment, while valuable, provides only preliminary evidence for long-term retention, and that multiple follow-up timepoints would substantially strengthen claims about the durability of neuromodulation effects. To carefully address this point, we have rephrased the “long-term retention” as “long-term after-effects” throughout the whole revised manuscript, and overall toned down the claims on the retention effects. Moreover, this limitation has been explicitly elaborated in the Discussion section:

      Abstract Section (Page 2, Line 49-50)

      “... we assessed the effect of anodal HD-tDCS on real-world procrastination behavior at offline after-effect (2-day interval) and long-term after-effect (6-month follow-up).”

      Results Section (Page 12, Line 606-608)

      “... Therefore, beyond short-term effects, the benefits of ms-tDCS neuromodulation on reducing procrastination were still detectable at a 6-month follow-up, providing preliminary evidence consistent with long-term after-effects.”

      Discussion Section (Page 14, Line 721-725)

      “... Thus, the detectable effects at 6 months are consistent with the hypothesis that repeated neuromodulation may induce neuroplastic changes in the DLPFC that support sustained behavioral change. However, we explicitly note that a single follow-up timepoint cannot establish the stability or trajectory of these effects; future studies with multiple longitudinal assessments are required to substantiate claims about long-term retention.”

      Discussion Section (Page 15, Line 771-775)

      “... Finally, while our 6-month follow-up provides preliminary evidence for sustained effects, the use of a single follow-up timepoint limits our ability to characterize the temporal trajectory of retention. Future studies incorporating multiple follow-up assessments (e.g., 1-month, 3-month, 6-month, 12-month) would strengthen evidence for the robustness and durability of this intervention.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please see my detailed comments above (6 points; several of them with subpoints a, b, c, etc).

      Thank you for listing those specific and helpful recommendations above. Please see our detailed response posed above, point-by-point.

    1. Author response:

      The following is the authors’ response to the original reviews

      Reviewer #1 (Public review):

      In the revised manuscript, the authors have addressed most of my concerns. In the text of this manuscript, the authors should still include more discussion on why osm-5 and daf-2 are categorized into two different groups. 

      According to this suggestion, we have expanded our discussion to discuss why osm-5 and daf-2 worms fall into different longevity groups despite the fact that disruption of DAF-16 decreases both of the their lifespans. Please see lines 363-376.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study uses an elegant visual-anagram approach to test whether perceived animacy structures visual working memory and attention while controlling for many low-level image properties. The evidence is solid, with converging results across seven preregistered experiments, but the central claim that animacy itself is represented independently of visual features should be tempered, as residual mid-level configural cues, ensemble or category structure, and broader semantic differences may also contribute to the effects. The work will be of interest to researchers studying high-level visual representation, attention, and working memory.

      We thank the Editors and Reviewers for this careful and informed assessment. We appreciate that every Reviewer found our approach to be elegant, our findings to be solid, and our question to be of broad interest. We respond to each Reviewer’s specific comments in more detail below; but we thought to summarize some of the highlights - especially the specific comments that come up in this Assessment - here.

      (1) The Reviewers make the insightful point that, even if our stimuli effectively control for many low-level features, there may be other high-level features that explain performance in our experiments (R2: “Although the anagram paradigm effectively controls low-level visual features […] these stimuli differ not only in animacy but also along other semantic dimensions such as natural versus manmade categories.”). We are happy to embrace this possibility. If our results were explained by high-level visual representation of the natural vs. manmade distinction, rather than the animate vs. inanimate distinction, this would still be an appeal to a (not altogether unrelated) high-level property being represented independently from its lower-level features, which was the primary motivation for our study. We framed our work specifically around animacy given the persistent debates regarding perceived animacy, as well as the fact that our stimuli do quite saliently vary along that dimension; but we are certainly open to other nearby high-level categories being at play. We also think this is an empirical question that could be tested in future work. For example, objects like rocks and lakes are natural but inanimate. If they behave more like dogs than like boots in our paradigms, then Reviewer #2 may be right that naturalness was the relevant property all along; but if they behave more like boots than like dogs, then perhaps it really was animacy doing the work. We now discuss this explicitly in our paper, and we appreciate the opportunity to not only clarify our claims but also spur discussion for future work.

      (2) Multiple Reviewers raise the question of whether semantic factors that go beyond the images themselves may be driving our effects. Reviewer #3 raises a particularly interesting question along these lines: “if all the stimuli in the experiments were replaced with the verbal names of the depicted objects instead of pictures, would we expect different results?” We have now taken this question quite literally and run this experiment exactly as described. Of course, much research already explores cognitive processing of animate/inanimate words, finding (for example) stronger memory for animate objects than inanimate ones (e.g., Nairne et al., 2013; Nairne et al., 2017). However, such tasks do not invoke effects of visual processing, whereas the question at issue here is specifically whether the visual system prioritizes animacy independent of its lower-level features. To this end, we conducted a new, pre-registered experiment (now Experiment 8) where participants search for animate/inanimate words on some trials, and animate/inanimate pictures on others. Given the nature of visual search tasks, we should expect to find no search advantage for words (as their meanings are not processed in vision per se) — and we should also expect to replicate (once again) our search advantage for pictures. This is exactly what we found. In other words, linguistic stimuli alone failed to produce the effect, while anagrams did produce the effect. We believe this rules out the strongest form of the semantic labeling account.

      (3) Finally, Reviewers #2 and #4 raise some concerns regarding residual mid-level features such as configural shape and ensemble statistics, which lie somewhere between animacy itself and more basic properties like contrast or spatial frequency. In our paper, we now clarify each of these concerns in greater detail. In short: We think that our stimuli and experiments indeed control for these residual cues. For example, rotating an image preserves its configural shape; and, as we argue below, the specific ensemble statistics argument fails to get off the ground without appeal to animacy itself.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Evidence for visual representation of animacy.

      Strengths:

      This is a very cool paper that casts light on a persistent problem in the psychology and philosophy of visual representation: is there high-level perception? Every vision scientist agrees that low-level features such as shape, color, texture, motion and spatial frequency are represented in visual perception, but there is a great deal of controversy about the representation of high-level properties such as causation, faces, agency and animacy. Animacy is especially problematic because there are large differences in line curvature between stimuli that represent animate and inanimate items.

      This article uses a novel approach-visual "anagrams" that are exactly the same image, except one is rotated 90 degrees relative to the other. They found persistent differences in visual processing between animate and inanimate stimuli. (Of course, the stimuli aren't animate-they represent animate items). For example, there were processing differences between changes between animate and inanimate items (rabbit to boot) that were not present in rabbit to dog. They also showed such differences in two kinds of visual search tasks.

      Of course, there are feature differences that exploit orientation. A classic example is the difference between a square and a diamond that is produced from the square by rotating it 45 degrees.

      They addressed an aspect of this challenge having to do with some features using silhouettes. There was no search advantage for silhouetted stimuli.

      Weaknesses:

      I thought this was an excellent submission. I have two suggestions for revision:

      We are glad to hear this Reviewer recognizes the broad challenge we are tackling in this work (separating high-level from low-level visual features) and found our submission to be “excellent”.

      (1) I thought that experiment 7 should have been described in more detail, with the upshot explained better. What exactly do the authors take it to show?

      Sorry for the lack of clarity here. We think the Reviewer actually gets this right earlier in their review; many feature differences exploit orientation, and our silhouettes control (Experiment 7) shows that those differences alone fail to explain our effects. For example, one might worry that our search effects merely reflect oddities in the aspect ratio or center of mass of the images. Converting the anagrams into silhouettes preserves these features. Thus, the fact that we found no search advantage with silhouettes suggests that these features on their own fail to produce the relevant effects; put the other way around, the effects we observed earlier must go beyond those features. We now discuss this in greater depth in our paper.

      (2) There should be a candid discussion of what the loose ends are and how they might be addressed. It would be good to have some examples like the square/diamond case with some indication of what would address such challenges.

      We agree with this, though we are somewhat limited by the space constraints of the Short Report format. A primary loose end we see is the possibility that high-level properties other than animacy explain our results (as raised by other Reviewers). We have added some discussion of this possibility to the paper.

      We would like to thank this Reviewer for their thoughtful feedback.

      Reviewer #2 (Public review):

      Summary:

      The authors present a creative approach using visual anagrams matched on low-level image statistics to isolate animacy from low-level visual features and report consistent effects of animacy on visual working memory and attention. While this is a thoughtful design and is well executed across seven pre-registered experiments, it remains unclear whether the reported effect is truly driven by animacy, as opposed to broader differences in ensemble statistics or semantic structure across the "mixed animacy" versus "uniform animacy" conditions. As such, the interpretation of a "pure" animacy effect may be overstated.

      Strengths:

      (1) An important methodological advance in controlling low-level confounds that have historically complicated the study of animacy.

      (2) The converging effects across multiple experiments, together with the pre-registered design, strengthen the reliability of the reported findings.

      We are glad to hear this Reviewer found our work to be “creative” and believes it offers an “important methodological advance”.

      Weaknesses:

      (1) Specificity of the animacy effect vs. category-level ensemble structure

      The central claim is that animacy itself drives the observed effects. However, the key manipulation ("mixed animacy" versus "uniform animacy") also introduces differences in category-level ensemble structure. For example, in Experiments 1-2, cross-category change detection (e.g., dog to chair) may be easier not because of animacy per se, but because of a change in overall ensemble statistics (Brady & Alvarez, 2011, 2015). In addition, since each display contains five objects (two in one category and three in the other category), cross-category changes may also alter category balance in a way that further facilitates detection. In contrast, within-category changes preserve both ensemble structure and category composition, making them more difficult to detect.

      Brady, T. F., & Alvarez, G. A. (2011). Hierarchical encoding in visual working memory: Ensemble statistics bias memory for individual items. Psychological Science.

      Brady, T. F., & Alvarez, G. A. (2015). Contextual effects in visual working memory reveal hierarchically structured memory representations. Journal of Vision.

      We appreciate the opportunity to clarify our claims and the support for them. Our claim is indeed that animacy (or a closely related high-level property; see below) drives our effects, over and above its lower-level correlates — i.e., that the explanation for differences in change detection or search across conditions will invoke a high-level property of the images. As we understand the Reviewer’s concern(s), they either (a) are already addressed by our novel methodology, or (b) would still fall perfectly in line with our claim as stated above.

      Consider the Reviewer’s concern that cross-category change detection “may be easier not because of animacy per se, but because of a change in overall ensemble statistics”. Which ensemble statistics change across categories in our stimulus set? Take as an example the case depicted in our figure, where a rabbit changes into either a dog (within-category) or a boot (cross-category). The dog and the boot are the very same image, just rotated; thus, they have the same luminance, curvature, area, spatial frequency, and so on. So if the change from rabbit to dog changes the array’s ensemble statistics with respect to any of those properties, it does so in the very same way as the change from rabbit to boot — and yet detection is still better for rabbit → boot than for rabbit → dog. Indeed, for nearly any ensemble statistic, the difference between the rabbit-display and the dog-display will be identical to the difference between the rabbit-display and the boot-display. To engage with the specific cases discussed in the two cited papers (Brady & Alvarez, 2011, 2015): The dog and the boot are the same size (because they are the same image), so average size is identical (just as average luminance, curvature, area, and spatial frequency are identical). And the very few properties left over (e.g., aspect-ratio) are addressed by later experiments.

      To be clear: We are not saying that there are no differences in ensemble statistics between the rabbit-display and the dog-display; across those displays, we replace one image with a different image, so there are likely all kinds of corresponding differences in ensemble statistics. The key question is whether that change in ensemble statistics differs across trial types in ways that might explain our effect - i.e., whether there is any difference between the rabbit-display and dog-display that is not also present between the rabbit-display and the boot-display. We don’t see how the answer could be yes, at least with respect to the statistics typically considered. A similar logic applies to the search tasks, with the silhouette control (Experiment 7) providing especially strong evidence that certain ensemble statistics or lower-level features cannot explain our effect.

      Now, it’s possible the Reviewer is referring to properties other than the low-/mid-level properties we mention above. Perhaps, for example, many animate stimuli on a display at one time have a striking collective appearance (all these animals are looking at me!) that lots of inanimate stimuli do not (this might be related to the Reviewer’s concern about “category balance”). But as we see it, this explanation just invokes animacy all over again, and so is the sort of explanation we would embrace.

      We now say more about this concern in the paper to be as clear as possible about our claims.

      (2) Limited stimulus set and potential learning effects

      The relatively small stimulus set (six anagram pairs) and repeated exposure raise the possibility of learning or familiarity effects. Does performance change over time? e.g., are there meaningful differences between early and late trials (e.g., first 10% vs. last 10%)? If such differences are present, they could suggest the development of task-specific strategies or increased efficiency with repeated exposure, rather than stable effects driven by the experimental manipulation itself.

      This is an interesting question, and we recognize this analysis absent from our initial submission. To be fair, stimulus sets of this size are not unusual in change-detection and search tasks, which often involve red, green, and blue squares repeated over the course of several hundred trials. Still, we certainly take the Reviewer’s point here and also embrace their analytical approach to addressing it. We’ve now run the “familiarity effects” analyses the Reviewer suggests (as well as some they did not suggest). The top-level headline is that learning or familiarity effects cannot explain our results, and if anything most of these analyses not only fail to support this alternative account but actively point against it. Below are more details.

      First, we worry that the Reviewer’s concern about “the development of task-specific strategies … rather than stable effects driven by the experimental manipulation itself” isn’t actually addressed by the suggested analysis of comparing the last 10% of trials to the first 10%. One reason for this is simply that it’s possible that both mechanisms are at play - i.e., that there is a baseline difference even without any familiarity that is then enhanced by some learning mechanism. (There are other issues as well: For example, one might imagine that participants get quite good at the task during the middle 80% of trials, but then get fatigued at the end. If this were true, then comparing the first 10% to the last 10% of trials could make it seem like there is no learning or familiarity, even if there were such effects. And on top of all this there is just the issue of statistical power, since far fewer trials go into these analyses than into our primary, pre-registered analyses). Nevertheless, we ran the Reviewer’s proposed analyses (using the first and last 10 trials of each type, which offers the best chance to find the pattern the Reviewer is concerned about). If anything, this analysis points in the opposite direction to the Reviewer’s prediction: 4/6 experiments (Experiments 1, 3, 4, and 5) revealed numerically weaker effects at the end of the task than the start, while only 2/6 experiments (Experiments 2 and 6) revealed numerically stronger effects at the end of the task than the start. Moreover, most of these results were non-significant, with only one marginal result (Experiment 5, p< = 0.08) and one significant result (Experiment 2, p = 0.01), and this is before any correction for multiple comparisons, which would make all of these results non-significant. So even though our account could easily accommodate learning effects, it’s not clear that they even exist here in any consistent or reliable way.

      Second, however, we think a more informative way to answer the Reviewer’s question is to ask not about learning over the course of the experiment but rather whether the key effects arise very early in the task. If they do, then any learning effects arising later couldn’t fully account for our results. Now, again, these tests are underpowered and only exploratory (to do this analysis properly, we would want to run entirely new experiments designed for this purpose), but we in fact did find evidence that our key effects arise early. In 5/6 experiments (Experiments 1, 3, 4, 5, and 6), the key effect was significantly (or in one case marginally) present even at the beginning of the experiment (Experiment 1, p = 0.07; Experiment 3, p = 0.01; Experiment 4, p < 0.001; Experiment 5, p < 0.001; Experiment 6, p < 0.01), and most of these results would survive correction for multiple comparisons. (In only one experiment, Experiment 2, was there a numerical disadvantage, but it was not significant; p = 0.34.) So even though our experiments were not designed or powered for this purpose, they do seem to suggest that the effects arise even without much familiarity at all.

      All told, we think these analyses suggest quite strongly that learning alone fails to explain our key effects. There is no evidence that the effects in general are stronger at the end of the experiment than the beginning (if anything it is the opposite); and there is evidence that most of the effects we investigated can be detected even very early in the experimental sessions. We have added discussion of these new analyses to our manuscript.

      (3) Role of semantics

      Although the anagram paradigm effectively controls low-level visual features, it still relies on high-level semantics (e.g., "dog" vs. "boot"). These stimuli differ not only in animacy but also along other semantic dimensions such as natural versus manmade categories. From a semantic standpoint, it remains unclear whether the observed effects can be uniquely attributed to animacy or whether they reflect broader conceptual distinctions.

      We agree with the Reviewer here. While we feel comfortable interpreting our effects in terms of a high-level property like animacy as opposed to a lower-level property like curvature, it remains possible that the observed effects reflect some other, closely related high-level distinction (like natural vs. manmade). Our primary concern was to tease apart high-level properties from low-level features, which the Reviewer’s question does not threaten — if attention and memory are sensitive to the natural/artificial distinction, that’s interesting too, and a near neighbor of our actual claim. Still, we agree that this could be addressed, and we even see it as an empirical question testable in future work. Perhaps the most relevant departures between animate/inanimate and natural/manmade include objects like clouds, plants, and rocks — objects that are natural but not “animate” in the sense often used in this literature. If something like our paradigm revealed that rocks behave more like dogs than like boots, that would suggest that naturalness, rather than animacy, was driving the effects; but if rocks behave more like boots than like dogs, that would point to animacy even more strongly. We remain open-minded about this possibility, but it would of course require multiple new experiments with a brand new stimulus set and so goes beyond the present contribution. In any case, we have added a discussion of this issue to the paper and have adjusted our claims accordingly.

      Reviewer #3 (Public review):

      Summary:

      This study makes clever use of generative AI to create stimuli that are pixel-for-pixel identical but which have radically different meanings depending on their orientation, to investigate the perception of animacy while retaining control over low-level image features (so-called 'anagram' stimuli).

      The authors present seven elegantly designed experiments in a commendably compact format.

      Experiments 1 and 2 involved a working memory paradigm in which participants had to spot which of five objects in an array changed after a pause. Importantly, the changed object was an anagram stimulus that in one orientation matched the animacy/inanimacy of the changed object, and in the other orientation was the opposite (e.g., a rabbit is replaced by either a dog or a boot, where the dog and boot stimuli are actually identical, just rotated by 90 degrees). They found a difference in accuracy depending on whether the animacy of the objects matched.

      Experiments 3 and 4 used a visual search task in which the participants had to localize the target, and the distractors were anagrams that either matched the target in terms of animacy or did not. There was a significant cost in terms of response time when the animacy of the target was the same as that of the distractors. Experiments 5 and 6 also used a similar visual search design, except that the task was to determine if the target was present or absent from the display, and the distractors again either matched or differed from the target in terms of animacy. Again, the authors found slower responses when the distractor arrays matched the animacy of the target than when they differed.

      An obvious potential concern about the studies is addressed by Experiment 7. It is unclear if the observed effects are related to the specific orientations of the target and distractor stimuli selected in each condition. For example, it could be that all the animate versions of the anagrams involved tall and skinny shapes, while all the inanimate versions involved wide and short objects, due to the 90-degree rotational difference between the two versions of the stimuli. To control for this, the authors repeated the visual search experiment but with convex-hull silhouettes of each of the stimuli. In other words, all targets and distractors from each trial were replaced by a black splotch with approximately the same overall outline (envelope) as the corresponding stimulus. Importantly, in contrast to the anagram stimuli, the silhouettes had had no meaningful semantic interpretation, and their animacy did not change depending on their orientation.

      Strengths:

      The main strength is the elegant use of stimuli that control almost perfectly for low-level image features.

      Thank you for this kind feedback. This summary perfectly captures both our empirical contribution and the claims we are making.

      Weaknesses:

      My only real concern about the study is whether the findings truly provide evidence for a high-level visual representation of animacy independent of the low-level stimulus characteristics, or whether, instead, the effects are essentially semantic priming, which is independent of visual processing per se. For example, if all the stimuli in the experiments were replaced with the verbal names of the depicted objects instead of pictures, would we expect different results? Words can also access semantic representations of the animacy of objects, and also don't suffer from low-level visual confounds. It would be helpful to add a discussion of this possibility to the article.

      Wow, we love this question! And so we’ve now conducted exactly the experiment the Reviewer suggests here. In a new pre-registered study (Experiment 8), we presented participants with a present/absent search task (as in Experiments 5–7). One half of trials consisted of the anagram stimuli (such that we could, once again, replicate the mixed-animacy search advantage); but the other half of trials consisted of the words describing the anagrams (e.g., “dog”, “boot”, “sheep”, “car”, etc.). The experiment worked beautifully: We found no effect with the words, but replicated the search advantage with the pictures — and also found a significant difference between the effects elicited by the two stimulus types.

      We agree with the Reviewer that this now rules out the possibility that semantic representations alone explain these visual effects. Thank you! 

      Reviewer #4 (Public review):

      In this article, the authors investigate whether perceived animacy influences visual processing independently of lower-level visual features by using "visual anagrams." Across seven experiments, they test whether animacy, isolated from many lower-level visual properties, structures visual working memory and guides visual attention. The central claim is that the visual system may represent animacy itself, rather than animacy emerging solely from associations among low-level visual properties.

      I find this investigation compelling. The experiments described provide strong control over several lower-level visual features, including curvature, texture, and related image properties. However, the visual anagrams are not pixelwise-identical across orientations. Because the images are rotated, the retinal configuration of pixels and the spatial organization of some low- to mid-level shape features also change. As a result, the configural arrangement of mid-level visual features may still contribute to perceived animacy.

      We are glad to hear the Reviewer finds our investigation “compelling”.

      I encourage the authors to discuss how independent perceived animacy is in this context from the contribution of mid-level visual features, such as configural shape cues that are diagnostic of animacy. This distinction would help sharpen the interpretation of the results and more precisely define the level of visual representation isolated by the visual-anagram approach.

      This is a helpful point, and it also echoes a sentiment expressed by Reviewer #2. While configural shape is diagnostic of animacy writ large, it can’t account for our observed effects here because rotating an image does not vary its configural shape. We now mention this in our work, and we agree that it helps sharpen the interpretation of our studies.

      Additionally, previous studies have argued that low- and mid-level curvilinear features may contribute to animate/inanimate categorization, and may in some cases be sufficient to support such distinctions (e.g., PMID: 33798259; PMID: 28654965). I encourage the authors to clarify how these previous findings on curvilinearity and rectilinearity fit with the overarching claim of the current study, namely that the visual system may represent animacy itself rather than animacy emerging solely from associations among lower-level visual properties.

      Yes, many studies from exactly that corner of the field actually motivated the present work, which is why we cited them in our submission. In a way, we are approaching this issue from the other side of the equation. Whereas the papers the Reviewer points to (along with many others) ask whether mid-level features (such as curvilinearity and rectilinearity) are sufficient to support perceived animacy, we ask whether these and other features are necessary to support perceived animacy. Prior work is relatively split on this issue, leaving the question wide open. We take our work to show that differences in curvature are not necessary for differences in perceived animacy, because our anagrams have identical curvature yet differ in animacy — and the visual system capitalizes on that difference. Put the other way around, representation of animacy can and does go beyond representation of its low- and mid-level correlates. Thank you!

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This study approaches an important topic providing insight into the neuronal circuitry that interconnects memory consolidation and sleep. The data were collected and analysed using a solid methodology, contributing new findings for neurobiologists working on how memories are stored and the roles of sleep. However, the data is incomplete to support the proposed role of the PAM-DPM circuits as the link between sleep state and long-term memory consolidation.

      We sincerely appreciate the editor and reviewers’ thoughtful and constructive comments on our study. Your insightful feedback has not only affirmed the significance of our work on the interplay between memory consolidation and sleep, but also provided valuable inputs for improving the clarity, rigour, and impact of our study.

      We have carefully addressed all the comments raised by the reviewers and revised the manuscript accordingly. We have also streamlined the paper with the goal of making it more accessible to readers. We feel this revised version strengthens our conclusion that the PAM-DPM circuits as the link between sleep and memory consolidation.

      The main improvements in terms of data addition are three complementary sets of circuit-specific experiments:

      (1) To better characterize the dynamics of the PAM-DPM circuit following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. This experiment addresses the activity of the microcircuit in a much more relevant time frame than the CRTC data in the previous version of the paper which looked only at the first hour after training. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, both PAM-α1 and DPM neurons displayed similar neural activity over the 3-hour recording period. The first hour was characterized by a synchronous decrease in activity for both neuron types, with hours 2 and 3 achieving a stable baseline. Since the decrease in the first hour is also seen in the trained condition, we think it is likely a reflection of the animals becoming acclimated to the recording tubes.

      Notably, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit in the LTM consolidation time window. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 8C-F). These findings not only support the hypothesized role of the inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics.

      (2) To further support the functional connectivity of the PAM-DPM microcircuit, we conducted in vivo experiments to complement the dissected brain prep P2X2 data. Optogenetic activation of PAM neurons in intact flies via the red light-gated cation channel CsChrimson (Klapoetke NC et al., 2014) resulted in a significant reduction in GCaMP signals within DPM neurons (revised Figure 2B). These findings strongly confirm that PAM neurons exert direct inhibitory control over DPM neurons in the intact brain.

      (3) Further, we investigated how dopamine signaling to the DPM inhibits its activity, and issue which has not been investigated previously. We conducted a series of experiments:

      Firstly, we verified which dopamine receptors (Dop1R1, Dop1R2, DopEcR, and Dop2R) express on the DPM neurons via double-labeling with gene-embedded GAL4 lines. We found that DPM neurons have expression of both Dop1R1 and Dop1R2 (revised Figure 10A).

      Secondly, to clarify which receptors on DPM neurons respond to dopamine and how they signal, in addition to EPAC experiments in the first submission, we recorded neural activity changes when we knocked down Dop1R1 and Dop1R2 in DPM neurons. DPM neurons exhibited a significantly reduced GCaMP level with DA application, regardless of whether Dop1R1 or Dop1R2 was intact or knocked down knockdown in comparison to the no-DA control condition (revised Supplemental Figure 3C-E). These data suggest that either residual Dop1R1 and Dop1R2 remaining in the RNAi condition is sufficient or that the two receptors may coordinate to mediate the inhibition of neural activity.

      Finally, we investigated the behavioral contributions of Dop1R1 and Dop1R2 in DPM neurons to sleep and memory processes (revised Figure 10C-H). Dop1R1 knockdown resulted in a marked reduction in daytime sleep and a significant impairment of 24 h memory expression. In contrast, Dop1R2 knockdown selectively compromised 24 h memory without affecting sleep.

      When integrated with our EPAC assay findings from the initial submission, which demonstrated, that Dop1R1 is the primary receptor mediating dopamine-induced cAMP elevation, these new data collectively delineate a more complex mechanistic framework: dopamine signaling in DPM neurons coordinates the dual regulation of sleep and memory predominantly via Dop1R1. Meanwhile, Dop1R2 are engaged in the selective modulation of memory.

      All newly generated experimental datasets, comprehensive statistical analyses, and their corresponding figure panels (revised Figures 2B, 10, 11 and Supplemental Figure 3) have been fully incorporated into the revised manuscript.

      In addition to adding the experiments described above, we have reorganized and streamlined the paper. First, the CRTC data have been replaced by the Tric-luc data. The CRTC data were taken in the first hour after training and do not shed light on the bulk of the consolidation window. Since the behavioral and sleep effects we see with manipulation of the PAM/DPM microcircuit all occur with a time delay, examining later times in consolidation is more relevant. Additionally, the first hour post-training is quite complex since there are sensory changes and STM processes overlaid on the processes we want to study. Second, we have moved the data in Figure 8 to supplemental (revised Supplemental Figure 2) since they are basically a control for the experiments in Figure 7 validating known requirements for appetitive LTM.

      We have also substantially expanded the Discussion section to contextualize the PAM-DPM circuit within the broader framework of well-characterized memory-regulatory pathways, such as the intrinsic circuits of the mushroom body, and to explicitly delineate the hierarchical interplay between sleep-dependent synaptic plasticity and LTM consolidation.

      We contend that these complementary experimental assays and targeted revisions markedly strengthen the causal evidence underscoring the role of the PAM-DPM circuit as a pivotal regulatory node bridging sleep states and LTM consolidation. We are confident that these revisions essentially address the concerns raised by the reviewers.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to use state-of-the art behavior, imaging, and connectome techniques to identify the neural interaction between sleep and long-term memory consolidation in the PAM-DPM circuits, a well-known dopaminergic pathway within Drosophila Mushroom Body.

      Strengths:

      From a Drosophila sleep researcher's perspective, the investigation follows a clear and logical strategy to collect a huge dataset of sleep, appetitive memory, and live imaging. The authors clearly identified and showed that activation of a PAM subset: alpha-1 reduces sleep quality and memory consolidation in a starvation-dependent manner. The authors also convincingly demonstrated the corresponding neuronal responses of DPM neurons following PAM alpha-1 activation, and the positive role of DPM neural activity in sleep and memory consolidation. Moreover, the authors applied a new way of sleep statistics to demonstrate hour-by-hour changes between treatment and genotypes. Importantly, the authors demonstrated that memory loss derived from PAM alpha 1 activation can be partly restored by ectopic sleep enhancement via feeding THIP during the memory consolidation period after training.

      Weaknesses:

      Two investigatory gaps relate to the misalignment between circuital activity and behaviors, due to the nature of large circuital functional analysis like this. Firstly, the central observation of the study indicates that PAM alpha1 activation causes DPM inhibition which disrupts sleep and memory consolidation. Therefore one would expect a reduced PAMalpha1 and increased DPM activities after memory training, but the authors found that the endogenous CRTC::GFP reported neuronal activity for PAMalpha1 and DPM are both increased after memory training (Figure 9). This can be due to the difficult functional demarcation among the 14 PAMalpha1 projections. Secondly, the authors acknowledged the contradicting finding that memory defect is detected in PAMalpha1 inactivation (Figure 7C), yet suggested a tight link between sleep and memory consolidation; it is clear loss of PAM subset activity can disrupt memory consolidation without affecting sleep (cf Figure 7C and 7I).

      Thank you for your insightful analysis and the relevant possibilities you've raised. We agree that given that memory consolidation and sleep are time-dependent processes, the 1-hour window employed to capture neural activity changes via the CRTC::GFP reporter may not fully reflect the overall dynamics of neural activity in this microcircuit. To better characterize the dynamics of the PAM-DPM circuit in the consolidation window following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, PAM-α1 neurons displayed stable neural activity over the entire 3-hour recording period. However, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 8C-F). These findings not only support to the hypothesized role of inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics.

      Regarding the second question, the core finding underlying the link between sleep and memory elucidated in the present study lies in the whole PAM-α1-DPM microcircuit rather than the specific DANs alone. MB299B and MB043B, the two split-GAL4 drivers employed to target PAM-α1 neurons, were originally characterized previously (Aso et al., 2014). However, these drivers also exhibit non-specific labeling of additional cells, and we can not rule out the possibility that such off-target labeling may have masked the subtype-specific necessity in sleep or memory processes.

      Reviewer #2 (Public review):

      Summary:

      Sleep plays a critical role in memory consolidation, but the neural mechanisms underlying this relationship remain poorly understood. The authors present novel findings implicating two small neuronal groups with inhibitory connections, PAM-a1 to DPM, in sleep regulation and LTM consolidation. However, whether the PAM-a1 to DPM microcircuit promotes LTM consolidation through sleep regulation requires further investigation.

      Strengths:

      The authors report several novel findings. Brief activation or inhibition of PAM-a1 neurons, or brief inhibition of DPM neurons during the first few hours after training, impairs 24-hour LTM. Notably, these brief manipulations disrupt sleep for many hours afterward, particularly at night. Interestingly, disruption of PAM-a1 and DPM neurons impairs sleep and appetitive memory consolidation only under starvation conditions, and pharmacological induction of sleep during the night rescues the LTM defects. These findings suggest that PAM-a1 and DPM neurons are involved in sleep regulation and LTM consolidation under starvation. These are important findings that advance our understanding of the link between sleep and memory consolidation.

      Weaknesses

      Some claims lack sufficient evidence or clarity:

      (1) All sleep experiments are conducted under the "training" (temperature-change) condition. While genotypic controls are helpful, additional no-training controls are required to confirm that the observed differences are due to training rather than unknown genotype-related factors. The fact that experimental genotypes exhibit significantly altered sleep even before "training" (e.g., Figs. 7H, J, K, 8A, B, D) highlights the necessity of these controls.

      Thank you for raising this important question. We have re-examined the sleep profiles recorded over two acclimation days and one day of baseline sleep, which preceded the implementation of the “training” paradigm (temperature manipulation) and thus served as a valid no-training control. As shown in Author response images 1-4, subtle yet discernible genotype-dependent differences were indeed observed under baseline conditions. However, when animals were subjected to starvation, the experimental manipulations (activation or inactivation of the target cells) elicited marked, statistically significant alterations in sleep patterns that cannot be accounted for by the baseline genotype differences. Collectively, these data confirm that the observed sleep phenotypes are attributable to the “training” intervention, rather than to confounding, pre-existing genotype-related factors.

      Author response image 1.

      Baseline and manipulation day sleep profiles following PAM activation and PAM/DPM inactivation under starvation conditions.

      Author response image 2.

      Baseline and manipulation day sleep profiles following PAM activation and PAM/DPM inactivation under non-starvation conditions.

      Author response image 3.

      Baseline and manipulation day sleep profiles following PAM- α1 activation and inactivation under starvation conditions.

      Author response image 4.

      Baseline and manipulation day sleep profiles following PAM- α1 activation and inactivation under non-starvation conditions.

      (2) Previous studies on disrupted memory due to sleep reduction have primarily examined conditions with severe sleep deprivation. In contrast, this report claims that relatively small decreases in total sleep accompanied by sleep fragmentation are responsible for impaired memory consolidation. It remains unclear whether sleep fragmentation at this level is truly critical for memory consolidation. The authors should cause sleep loss and fragmentation of similar magnitude through other means and determine whether it can impair LTM.

      We appreciate the reviewer’s insightful suggestion. While alternative assays for inducing sleep loss or sleep fragmentation are indeed available, this line of investigation lies beyond the core scope of the present study. We will certainly take this valuable suggestion into consideration for the future studies.

      (3) The authors employed a neural activity reporter to show that starvation increases the basal activity of PAM-a1 but not DPM neurons in untrained flies (Figures 9C-E). They observed small increases in the activity of both neuron groups immediately after training but not one hour later. Given the inhibitory connection from PAM-a1 to DPM, it is unclear why both neuron groups show increased activity after training. Additionally, as the authors acknowledge, it is puzzling how the inactivation of PAM-a1 produces similar effects on sleep and memory as DPM inhibition and PAM-a1 activation. Further experiments are needed to clarify these findings, such as manipulating PAM-a1 activity during the one-hour post-training period and evaluating the effect on DPM activity. Including data from training under fed conditions would provide a more comprehensive understanding of state-dependent neural activity. Even if certain experiments are not feasible, these issues warrant further discussion. It is also important to clarify that the term "synchronized" does not imply single-spike-level synchrony.

      Thank you for raising these critical questions. To deepen our understanding of these issues, we have conducted additional experiments and have incorporated them into the revised manuscript. Below are our specific responses to each of your points:

      (1) Regarding the contradiction between "PAM-α1 inhibition of DPM" and a transient increase in the activity of both neurons immediately after training:

      PAM/PAM-α1 neurons are well-documented to respond to reward signals (Liu et al., 2012, Ichinose et al., 2015), while DPM neurons have been shown to respond to both olfactory stimuli and electric shocks, and to form delayed olfactory memory traces (Yu et al., 2005). Thus, the concurrent increase in the activity of PAM-α1 and DPM neurons immediately following training is likely a response to the olfactory and/or sucrose stimuli in the assay. Given that memory consolidation and sleep are time-dependent processes, the 1-hour window employed to capture neural activity changes via the CRTC::GFP reporter likely does not fully reflect the overall dynamics of neural activity in this microcircuit. Additionally, this time window overlaps with the period in which the animals are adapting to the new tubes and is likely contaminated with other sensory information.

      To better characterize the dynamics of the PAM-DPM circuit following associative memory training, we performed 3-hour continuous neural activity recording in freely behaving flies. Specifically, we expressed the Tric-LUC reporter gene, a calcium-responsive tool that harnesses the interaction between calmodulin and its cognate binding peptides to drive rapid luciferase transcription in a calcium-dependent manner (Gao et al., 2015; Guo et al., 2017), in PAM-α1 and DPM neurons, respectively. Flies were then subjected to either associative memory training or a no-training control condition, with real-time luciferase levels monitored throughout the recording window.

      In the absence of training, PAM-α1 neurons displayed stable neural activity over the entire 3-hour recording period. Notably, associative memory training profoundly reshaped the activity profile of the PAM-DPM circuit. Training induced a mild yet statistically significant elevation in PAM-α1 neural activity specifically during the third hour of recording, while concurrently eliciting a robust reduction in DPM neuron activity over the last two hours (revised Figure 9F-I). These findings not only support to the hypothesized role of inhibitory PAM-α1-DPM circuit in sleep and memory consolidation, but also advance our mechanistic understanding of underlying neural dynamics post-training. We have replaced the CRTC data with this more relevant data set.

      (2) Regarding the state-dependent neural activity:

      We agree that investigating state-dependent neural activity would be an interesting extension of our study. However, this falls beyond the scope of the current study and will be considered in future research. Our primary findings, including sleep disruptions and the associated memory impairments, were specifically observed under starvation conditions, which align with the appetitive memory paradigm employed here. Delving into neural activity changes under non-starvation state would not yield direct evidence to support the core conclusions of the present work, as the study’s focus is on the starvation-dependent interplay between sleep, neural circuitry, and appetitive memory consolidation.

      (3) Regarding the terminology of “synchronization”:

      We believe that the use of the term “synchronization” in our study is appropriate. In the context of neural circuitry, synchronization refers to the process by which distinct neurons or neural populations achieve temporal alignment of their activity, a phenomenon that supports neural communication and information integration. In the present work, this specifically describes how PAM-α1 and DPM neurons exhibit phase-related temporal coordination of their activity to regulate the interplay between sleep and memory consolidation.

      (4) The authors considered that PAM-a1 and DPM might function in parallel, independent pathways for sleep and LTM. They rejected this possibility based on the lack of additive effects when both neuronal groups were simultaneously inactivated. However, they found that MB299B-labelled neurons exert stronger memory effects than MB043B-labelled neurons, while MB043B neurons have stronger sleep effects. If sleep is a primary driver of memory consolidation, a stronger correlation between memory and sleep effects would be expected. This observation merits further discussion.

      We appreciate the reviewer’s constructive suggestions. We have performed additional experiments to explore a well-characterized memory-related PAM-α1 recurrent loop in sleep regulation. The new data, along with further discussion, have been incorporated into the revised manuscript.

      The two split-GAL4 drivers (MB299B and MB043B) used to target PAM-α1 neurons were originally characterized previously (Aso et al., 2014). However, these drivers exhibit non-specific labeling of additional neuronal populations, a technical limitation that may have masked the subtype-specific functional requirements of PAM-α1 in sleep and memory processes.

      In addition, we assessed sleep and LTM following the thermoactivation of DPM neurons (revised Supplemental Figure 1), and no significant changes were observed in either phenotype.

      PAM-α1 has previously been demonstrated to drive appetitive LTM formation and consolidation via a recurrent loop with MBON-α1 (Ichinose et al., 2015). To investigate whether MBON-α1 also participates in sleep regulation, we activated or inactivated MBON-α1 neurons under both starvation and non-starvation conditions. Our results revealed that inhibition of MBON-α1 under both starvation and non-starvation conditions resulted in a significant reduction in sleep and a reduced arousal threshold (revised Figure 11B, D), suggesting that MBON-α1 participates in regulating sleep in a state-independent manner. However, no significant changes were observed upon activation of MBON-α1 neurons (revised Figure 11A, C). Combined with our observation that inhibition of MBON-α1 during the memory consolidation phase also impaired 24 h LTM, these new data indicate that MBON-α1-mediated sleep is necessary for effective memory consolidation. Notably, while activation of MBON-α1 during consolidation phase similarly impaired LTM, this manipulation did not alter the sleep profile, suggesting a dissociation between MBON-α1’s mechanistic roles in sleep regulation and LTM processing.

      Taken together (see Author response table 1 and the new schematic diagram of revised Figure 12), these findings reveal a dedicated hierarchical, modular regulatory network that mediates sleep-LTM coupling via an activity-dependent mechanism. Within this network, activation of PAM-α1 acts as an upstream modulator to inhibit the activity of DPM, a downstream integrative hub that coordinates the execution of sleep and memory processes via recruiting different signaling cascades mediated by distinct dopamine receptors. MBON-α1, which is likely inhibited by PAM-α1, serves as parallel pathway to suppress sleep and impair LTM. Conversely, inactivation of PAM-α1 relieves its inhibitory control over MBON-α1, leading to MBON-α1 activation; MBON-α1 then functions as a signal amplifier that further exacerbates the reduced activity of PAM-α1, ultimately resulting in LTM impairment. Inactivation of PAM-α1, together with non-PAM-α1 neurons labeled by MB043B, contributes to the regulation of sleep. Sleep and memory are highly intertwined within this circuit, where distinct neuronal populations exhibit specialized yet interdependent functional roles, with overlapping and divergent regulatory contributions to sleep and LTM. The inherent complexity of this regulatory network thus merits further dedicated investigation in future studies.

      Author response table 1.

      (5) Given prior knowledge that PAM neurons are heterogeneous and that the R58E02 driver is broadly expressed, data in Figures 1-5 concerning PAM are outdated. The use of more restricted PAM-a1 drivers from the outset would make the manuscript easier to read and interpret.

      We sincerely appreciate the reviewer’s point of view regarding the selection of PAM drivers. While we acknowledge the well-characterized heterogeneity of PAM neurons and the broad expression profile of the R58E02 driver, and fully agree that employing subtype-restricted drivers enhances the precision of functional interpretation, this set of experiments serves as an essential foundational step and logical basis for subsequent subtype-specific investigations and thus merits retention in the manuscript. As detailed above, the more specific drivers also have some drawbacks in terms of additional expression, making the broad driver critical for setting the stage.

      (6) Some figures lack relevant data, certain experiments are missing necessary controls, and anomalies are present in some data sets.

      We sincerely appreciate the reviewer’s detailed suggestions, and we have revised the manuscript comprehensively in accordance with them.

      Reviewer #3 (Public review):

      Summary:

      Understanding the neural circuits that link sleep and memory remains a fundamental challenge in neuroscience. In this study, Lin Yan and colleagues investigate how dopamine signaling in Drosophila regulates long-term memory (LTM) formation in the context of sleep. They identify a specific microcircuit between protocerebral anterior medial dopamine neurons (PAM-DANs) and dorsal paired medial (GABAergic DPM) neurons that modulates memory consolidation. Their findings suggest that disrupting the basal activity of PAM-α1 neurons during early consolidation impairs LTM, with particularly pronounced effects under starvation conditions. Notably, sleep fragmentation caused by this disruption can be pharmacologically rescued, restoring LTM. These results provide compelling evidence that dopamine signaling plays a crucial role in linking sleep and memory, offering new insights into the underlying mechanisms.

      Strengths:

      This study presents a well-executed investigation into sleep-memory interactions, utilizing a combination of connectomics, behavioral assays, functional imaging, and pharmacological manipulations. The authors convincingly demonstrate that the PAM-α1 and DPM circuits interact, highlighting a potential mechanism by which sleep influences memory consolidation. The anatomical and functional dissection of this circuit is of high interest to the field, and the study's integration of sleep and memory processes contributes significantly to our understanding of dopamine's role in cognitive functions.

      Weaknesses:

      While the study is well-designed and presents compelling findings, some aspects require further clarification. The interpretation of dopamine receptor signaling remains incomplete, particularly regarding inhibitory pathways. The role of DPM in memory consolidation is not entirely conclusive, as different genetic approaches yield variable results. Additionally, some inconsistencies in neuronal activity patterns and experimental variability, especially regarding sleep patterns or pharmacological rescue, should be addressed to strengthen the mechanistic framework.

      Conclusion:

      Overall, this study provides valuable new insights into how sleep and dopamine circuits interact to regulate memory consolidation. While the findings are compelling, addressing the points above-particularly receptor signaling and the specific role of DPM and its activity patterns within the microcircuit would further solidify the study's conclusions.

      We sincerely appreciate the reviewer’s constructive feedback and useful suggestions, which have been instrumental in enhancing the rigour and completeness of our study.

      To address these points, we have performed a series of additional experiments that we believe strengthen the mechanistic framework of our work. The key new findings are summarized below:

      (1) Regarding the dopamine receptor signaling

      To define the dopamine receptor (DAR) signaling mechanisms underlying DPM neuron activity and its regulatory roles in sleep and memory, we first characterized DAR expression profile of the DPM neurons. Using double-labeling assay, we detected robust expression of Dop1R1 and Dop1R2 in DPM neurons, whereas no detectable colocalization was observed for DopEcR and Dop2R (revised Figure 10A). Accordingly, we refined our FRET-based EPAC data by removing the DopEcR knockdown group, and now present cAMP changes in DPM neurons following Dop1R1 and Dop1R2 knockdown, in direct comparison with the intact receptor control group (revised Figure 10B). These data conform that Gαs-coupled Dop1R1 is the primary receptor mediating DA-dependent cAMP elevation in DPM neurons.

      To further identify the DARs responsible for transducing DA-induced inhibitory effect on DPM neural activity, we quantified GCaMP levels in DPM neurons with targeted knockdown of individual DARs. Knockdown of either Dop1R1 or Dop1R2 failed to abolish DA-induced Ca<sup>2+</sup> decrease; only Dop1R2 knockdown exhibited a trend toward attenuating this Ca<sup>2+</sup> decrease (revised Supplemental Figure 3C-E), suggesting that the two receptors cooperate to modulate DPM neural activity.

      Finally, to dissect the specific contributions of DARs in DPM neurons to sleep and/or memory regulation, we performed sleep monitoring and memory assays in animals with DPM-specific knockdown of distinct DARs (revised Figure 10C-H). Knockdown Dop1R1 in DPM neurons resulted in statistically significant sleep reduction, decreased arousal threshold, and impaired 24 h LTM memory (revised Figure 10C-E). In contrast, knockdown Dop1R2 in DPM neuron selectively impaired 24 h LTM memory with no effect on sleep (revised Figure 10F-H). Collectively, these findings demonstrate that coupling sleep and LTM requires Dop1R1 in DPM neurons through the modulation of both cAMP signaling and neuronal activity, while Dop1R2 specifically mediates LTM regulation, likely through modulating DPM neural activity alone.

      (2) We have additionally characterized the role of MBON-α1 in sleep, which has been previously shown as a PAM-α1-related recurrent feedback loop in the regulation of memory formation and consolidation (Ichinose et al., 2015).

      To investigate whether MBON-α1 also participates in sleep regulation, we activated or inactivated MBON-α1 neurons under both starvation and non-starvation conditions (revised Figure 11A-D). Our results revealed that inhibition of MBON-α1 under both starvation and non-starvation conditions resulted in a significant reduction in sleep and a reduced arousal threshold (revised Figure 11B, D), suggesting that MBON-α1 participates in regulating sleep in a state-independent manner. However, no significant changes were observed upon activation of MBON-α1 neurons (revised Figure 11A, C). Moreover, inhibition of MBON-α1 during the memory consolidation phase significantly impaired 24 h LTM (revised Figure 11E-F). These results indicate that MBON-α1mediated sleep is necessary for effective memory consolidation. Notably, while activation of MBON-α1 during consolidation phase similarly impaired LTM, this manipulation did not alter the sleep profile, suggesting a dissociation between MBON-α1’s mechanistic roles in sleep regulation and LTM processing.

      Taken together (see Author response table 1 and the new schematic diagram of revised Figure 12), these findings reveal a dedicated hierarchical, modular regulatory network that mediates sleep-LTM coupling via an activity-dependent mechanism. Within this network, activation of PAM-α1 acts as an upstream modulator to inhibit the activity of DPM, a downstream integrative hub that coordinates the execution of sleep and memory processes via recruiting different signaling cascades mediated by distinct dopamine receptors. MBON-α1, which is likely inhibited by PAM-α1, serves as parallel pathway to suppress sleep and impair LTM. Conversely, inactivation of PAM-α1 relieves its inhibitory control over MBON-α1, leading to MBON-α1 activation; MBON-α1 then functions as a signal amplifier that further exacerbates the reduced activity of PAM-α1, ultimately resulting in LTM impairment. Inactivation of PAM-α1, together with non-PAM-α1 neurons labeled by MB043B, contributes to the regulation of sleep. Sleep and memory are highly intertwined within this circuit, where distinct neuronal populations exhibit specialized yet interdependent functional roles, with overlapping and divergent regulatory contributions to sleep and LTM. The inherent complexity of this regulatory network thus merits further dedicated investigation in future studies (See Author response table 1).

      We have modified the schematic diagram in the revised manuscript to illustrate the mechanistic framework (revised Figure 12).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here I listed details for potential clarification or further investigation related to the weaknesses:

      (1) Line 145-147: I suspected the authors used previously verified RNAi lines, but it would be informative to include a citation or their own validation for the effectiveness of these RNAi lines.

      We sincerely appreciate the reviewer’s suggestion. As our double-labeling assays confirmed that only Dop1R1 and Dop1R2 are colocalized with DPM neurons (see our responses to the public review from Reviewer #2 and #3), we refined the revised data to focus exclusively on these two DARs. Corresponding revisions have been made to the Materials and Methods, Results and Discussion sections. Additionally, we conducted qPCR analysis to verify the knockdown efficiency of these DARs, providing further support for our findings that Dop1R1 and Dop1R2 are functionally required in DPM neurons for the regulation of sleep and memory (revised Supplemental Figure 3).

      (2) Line 169: Moving from describing Figure 3A/B to Figure 3C, it is not immediately clear from 3C-H, the authors follow the training paradigm of 3B?

      To enhance clarity, we have added the referenced figure citations in the “Memory assay” section: “For all 24 h sucrose-odour memory, a single training session of sucrose paired with an odour for 2 min was employed (Figure 3A).”

      (3) Line 179: before the PAM inactivation data are shown in Figure 7, the authors seem to be getting ahead of themselves by stating "suggesting that activity of PAM neurons is necessary for the consolidation window or that heterogeneity in the subsets of PAM neurons masks any phenotype." when the Figure 3 data collectively indicate that "suppression" of PAM is necessary.

      This statement is based on our observation that inactivation of the majority of PAM neurons labeled by R58E02 results in sleep disruption but leaves memory intact, and we stand by this conclusion.

      (4) Line 186-188: The labelling of DP1 is not entirely aligned between the figures and the text for a reader to follow which time period is described, as DP1 is embedded within the dark phase in the figures.

      These experiments spanned two consecutive days. LP1 and DP1 denote the light and dark periods on the first day, respectively, whereas LP2 designates the light period on the second day. As only one full dark phase was monitored across the experimental interval, we characterized the relevant phenotype using the general terms dark phase or nighttime, rather than specifying DP1. We thank the reviewer for this thoughtful observation; nonetheless, we consider the original description correct and unambiguous, and thus appropriate for inclusion in the manuscript.

      (5) Line 321: The statistics for Figure 9 CRTC::GFP measurement is crucial for interpretation but the referring and labelling for this on Figure 9 is poor: it is not apparent which comparisons are indicated. There is inconsistency between Table 1 Figure 9D and Table 3 Figure 9D entries: no significant between train and untrain indicated in Table 1 but it is described as significant in the text and Table 3?

      Thank you for this observation. As described above, we have removed these data from the paper and replaced them with Tric-Luc data that capture the consolidation window more completely.

      (6) Line 423: The starvation-mediated sleep suppression is not clear in this manuscript, can the author comment on this? The response to this may also alter the summary concept cartoon.

      This is an important point. To directly address the reviewer’s question regarding starvation-mediated sleep suppression, we have generated a representative response figure comparing sleep duration under starvation versus non-starvation conditions (Author response image 5). This figure clearly demonstrates that sleep is suppressed under starvation, providing straightforward evidence to address this concern.

      Author response image 5.

      Examples of starvation-mediated sleep suppression.

      However, the key focus of our study is that changes in neuronal activity disrupt sleep under starvation conditions but not under non-starvation conditions. To emphasize this critical distinction, we have incorporated additional discussion focused specifically on this point.

      “It is well established that starvation induces sleep suppression (MacFadyen, 1973; Thimgan et al., 2010; Melnattur and Shaw, 2019; Keene et al., 2010; He et al., 2020; Yangkyun et al., 2022), and our results are consistent with these previous findings: all genotypes exhibited less sleep under starvation than under fed conditions (i.e. Figures 4A-B, 5A-B and 11). Under normal appetitive memory training, starvation-induced sleep loss does not necessarily impair memory processing (Thimgan et al., 2010; Chouhan et al., 2021). PAM-α1 neuronal activity is higher in starved, trained flies than in fed or untrained flies (data not shown), suggesting that these neurons act as a critical node for integrating internal motivational and arousal states, as well as conveying positive valence for the normal appetitive memory process, independently of starvation-induced sleep loss. While DPM neurons are less sensitive to starvation, they still exhibit training-induced elevated activity (data not shown), indicating coherent responsiveness to upstream signaling. In the present study, we found that under fed conditions, sleep remained intact even when excessive changes in neural activity occurred within the PAM(-α1)-DPM circuit; in contrast, under starvation conditions, significant sleep reduction and fragmentation were observed. These observations indicate that starvation may trigger a transition from a physiologically normal brain state to an unstable, abnormally active state, which consequently elicits negative behavioral outputs.”

      (7) Line 1121: The data points for Figure 9 D-E are surprisingly low considering there are 14 PAMalpha1 labelled, the data presented here indicated potentially only 1-2 neurons were counted per fly brain. Can this contribute to the large variation and the contradiction of PAM's memory-suppressing role?

      We sincerely appreciate the reviewer’s critical comments regarding the sample size of labeled PAM-α1 neurons in Figure 9D–E. We have revisited our raw data, incorporated additional brain samples, and reanalyzed the dataset. For this updated analysis, we included all clearly distinguished neurons, excluded overlapping ones, and calculated a single NLI per brain for statistical analysis. The key conclusions remain consistent with those in the original submission, confirming the robustness of the observed phenotype.

      Memory consolidation is a time-dependent process. To further elucidate the link between neural activity and behavioral outputs, we performed additional experiments with an extended recording period. A detailed response to this point is provided in the response to public review, and we therefore do not reiterate the details here.

      (8) Line 345: the effect size and data spread of THIP restored memory is different from the controls in Figure 10, perhaps warranting a more conservative interpretation of the role of sleep in memory consolidation.

      We appreciate this critical comment. We fully agree that the role of sleep in memory consolidation requires cautious interpretation, a point we have integrated into the revised manuscript.

      Drug treatment in Drosophila, particularly for group-based assays, can introduce substantial variability at both the individual and group levels. To account for this, we employed a statistically valid sample size for our analyses to ensure robust conclusions. While minor quantitative discrepancies exist in the data, this technical consideration does not significantly alter the core conclusions of the study.

      Reviewer #2 (Recommendations for the authors):

      (1) As mentioned in the public review, all data using the broad PAM-DAN driver should be removed. Concerns regarding the experiments involving the broad driver are not included here.

      A detailed response to this point is provided in the response to public review, and we therefore do not reiterate the details here.

      (2) In GCaMP experiments (Figure 9B), the ΔF/F traces for the AHL and AHL+ATP conditions start diverging before the addition of ATP. The quantification shows they are not significantly different in the first 30 seconds, but the fact that in two separate experiments (2A and 9B), they diverge in the same direction makes me wonder whether the AHL condition is different from the +ATP condition even before the ATP treatment. Also, the traces should include standard errors.

      We observed the same diverging trend in the first 30-second baseline as the reviewer. We reviewed the raw data for each sample and found that this divergence is likely attributable a small number of outliers. Given the absence of a statistically significant difference, this divergence does not affect our conclusions.

      We have also added standard errors to the revised figures.

      (3) Figure 9B. The authors need to show data for a control genotype. +>P2X2; VT064246-LexA > GCaMP6f that does not include MB299B-Gal4 is crucial to demonstrate that expression of P2X2 in PAM-α1 is responsible for the inhibitor effect, as LexA-P2X2 may be leaky.

      One of the UAS-P2X2 lines was found to exhibit leaky expression, so we instead used a non-leaky UAS-P2X2 line for all related experiments. To address the reviewer’s comments and further validate our findings, we have added complementary experiments with a control genotype. In addition, we also added a control to confirm the non-leaky expression of LexA-P2X2 under the driver of R58E02-LexA. As shown in revised Figures 8B, application of ATP in the absence of MB299B-GAL4 failed to induce a significant inhibitory effect. These data strongly and convincingly support our conclusion.

      (4) Figure 9B. Some of the individual data show values lower than -100% ΔF/F0. By definition, ΔF/F cannot be less than -100%, as this would require negative fluorescence, which is physically impossible. The calculation of fluorescence changes using ΔF/F should be carefully reconsidered.

      We thank the reviewer pointing out this potential confusion. We used a standard method of calculating the change in fluorescence over time using △F/F = (Fn-F<sub>0</sub>) / F<sub>0</sub>×100% as we previously described (Liu et al., 2019). Changes of greater than +100% of △F/F would not be unusual, since the reported value is a ratio to the initial level of fluorescence, not a subtraction of the baseline value from the signal (which obviously could not go below 100%). We have included a sentence in the results explaining this (page 7): “As previously described, we used the percent change in fluorescence over time as a ratio to the initial level, △F/F = (Fn-F0)/F0×100% for quantification (Liu et al., 2019).” And we have carefully reviewed our raw and processed data and confirmed that our analysis was correct.

      (5) Figure 2B. The number of UAS transgenes should be controlled, as Gal4 could be diluted with 3 UAS constructs in experimental conditions compared to only 1 UAS construct in controls. Are Dop1R2 and DopEcR significantly different from wt? Why do they present an average ΔF/F in 2A and a maximum in 2B?

      We appreciate the reviewer’s careful observations and valuable comments.

      As the reviewer noted, the EPAC imaging experiment utilizes three UAS transgenes, which enable Gal4 enhancement via Dicer, targeted manipulation of dopamine receptor expression levels, and neural activity monitoring in DPM neurons. All other imaging experiments in the study employ only one or two UAS transgenes. Given the robustness of the observed phenotypes, the potential dilution effect is not a major concern. Knockdown of Dop1R2 and DopEcR showed no significant differences relative to the WT control group; the maximum values presented in Fig. 2B are included solely to illustrate statistical significance. While the EPAC (CFP/YPF) signal reflects an obvious cAMP elevation, no differences were detected in the averaged signal across groups.

      Notably, in the revised manuscript, our double-labeling assays confirmed that only Dop1R1 and Dop1R2 are colocalized with DPM neurons (see our responses to the public review from Reviewer #2). Accordingly, we have refined our data analysis to focus exclusively on these two DARs.

      (6) Figures 7H, J. Why is almost every MB299B>TrpA1 fly sleeping at ZT0?

      To align the starvation protocol for sleep analysis with that used in the memory assay, MB299B>TrpA1 flies and their genetic controls were transferred to fresh sleep tubes containing starvation food during the ZT0–1 time window. This transfer resulted in no detectable locomotor activity during this period, a pattern indicative of sleep in all flies.

      (7) The number of episodes and P(wake) should be presented for all sleep data.

      We have added these two parameters as new panels to all relevant sleep figures. The corresponding statistical analyses have also been included in the supplemental tables.

      Reviewer #3 (Recommendations for the authors):

      The study's findings provide compelling insights into the neural circuits connecting sleep and memory and the role of dopamine in general. While the anatomic dissection of the microcircuit and its overall involvement in sleep and memory is convincing and of high interest to the field and beyond, some statements of the study need further clarification, particularly the interpretation of receptor signaling and the role of DPM.

      Major Points

      (1) Figure 2: cAMP Imaging and Dopamine Receptor Involvement

      The authors present calcium and cAMP imaging to support the inhibitory connection between PAM and DPM neurons. While using both sensors is a robust approach, I am not entirely convinced that cAMP imaging is the ideal approach for identifying the dopamine receptors involved. To my knowledge, only Dop1R1 is classically linked to Gs-mediated cAMP signaling. Dop1R2 is typically coupled to Gq (PLC and DAG), while DopEcR is non-canonical and can engage both pathways. Additionally, these receptors are classically excitatory, yet the authors did not analyze Dop2R, the primary inhibitory dopamine receptor - which would represent the most relevant candidate for an inhibitory PAM-DPM connection.

      We have addressed this point in our response to the public comments, so will not reiterate here.

      While dopamine receptor functions can vary by neuronal context, I would appreciate clarification on the following points:

      (a) Why was Dop2R not tested? Was it omitted or found to have no effect?

      We sincerely appreciate the reviewer’s critical questions. This point has been addressed in our response to the public comments. Briefly, Dop2R is not colocalized with DPM neurons; instead, only Dop1R1 and Dop1R2 are detected in DPM neurons, which is why we focused exclusively on these two receptors in the revised manuscript.

      (b) Why was cAMP imaging chosen for receptor identification? Was calcium imaging performed, and if so, what were the results?

      This is an excellent point, and we sincerely appreciate the reviewer’s valuable input, which has helped to strengthen the logical framework of our analysis on receptor-mediated neural activity. These dopamine receptors are well-characterized as members of the Gas-coupled protein receptor family, and cAMP signaling serves as a reliable readout of their functional activity. To strengthen the logic flow of our analysis on the target inhibitory circuit, we have made the following key revisions to the manuscript: 1) defined the expression profile of dopamine receptors in DPM neurons; 2) refined our cAMP imaging data analyses based on specific receptor subtypes; and 3) assessed DPM neural activity via calcium imaging under conditions of targeted receptor knockdown. For further details, please refer to our response to the public comments.

      (c) Since the data suggest multiple receptor involvements and complex interactions, I encourage a more detailed discussion of the working hypothesis, particularly regarding the unexpected finding that classically excitatory receptors contribute to an inhibitory connection.

      We appreciate the suggestion to elaborate on our working model. Accordingly, we have revised the schematic diagram and refined the manuscript to clearly illustrate the underlying mechanistic framework. For further details, please refer to our response to the public comments.

      (2) Figure 3: DPM Involvement in Memory Consolidation

      The authors show that PAM activation and DPM inhibition during consolidation impair appetitive LTM. However, the role of DPM is critical. While the c316-GAL4 driver yields strong effects, VT064246 inhibition shows only slight significance, requiring more than twice the sample size of other experiments. Given that c316-GAL4 is not DPM-specific and also labels MB Kenyon cells, I suggest using MB-GAL80 to restrict expression - or commenting on the possibility that other neurons like MB-KCs could directly participate in the phenotype. This is particularly relevant since VT064246 efficiently modulates sleep, indicating that it is generally effective in altering behavior. These issues weaken the claim that DPM plays a crucial role in linking sleep and memory, and should be addressed. Minor comment on this Figure: In the Figure legend, the driver and "n" are not mentioned for 3C, while this is the case for all other panels. Moreover, the DPM schematic only depicts the MB, making it somewhat confusing. DPM innervates the entire MB, still, it would be helpful to shade the DPM projections more distinctly within the MB for clarity.

      We thank the reviewer for the suggestion to improve the precision of our figures.

      Regarding the expression specificity concern, in all experiments using c316-GAL4, we had eyeless-GAL80 and MB-GAL80 co-expressed to restrict GAL4-driven expression to DPMs. While complete suppression of expression of MB-KCs was not achievable, we largely eliminated the potential confounding effects from majority of these cells. VT064246-GAL4 is known to exhibit weak expression (Jenett et al., 2011; Haynes et al., 2015), but high relative specificity. Importantly, the overall conclusion derived from experiments using c316-GAL4 with GAL80s and VT064246-GAL4 are consistent, which strongly supports the role of DPM neurons in mediating the link between sleep and memory.

      As suggested, we have added sample sizes for all panels and refined the depiction of DPM projections in revised Figure 3C.

      Minor Comments

      (1) Introduction:

      The authors introduce dopamine's role in forgetting but focus on aversive rather than appetitive memories. To avoid confusion, this distinction should be mentioned explicitly (likewise in the discussion). Regarding references: Zhang et al. (line 95) do not discuss DPM or APL. Donlea et al. (line 97) do not cover dopamine - I think Pimentel et al. (2016) would be a more appropriate citation.

      This is a good point. We have removed Zhang et al. (2013) and replaced Donlea et al with Pimentel et al. 2016 as suggested.

      (2) Figure 9: DPM Activation During Consolidation:

      The authors show that PAM neurons are activated by starvation and further enhanced by appetitive training. Surprisingly, DPM neurons also increase activity post-training, despite the proposed inhibitory connection between PAM and DPM. The authors state that "PAM-α1-DPM microcircuit exhibits synchronized neural activity changes during the consolidation window" (line 326), yet they do not address this apparent contradiction. If I have not overlooked key information, this should be clarified/addressed e.g. in the discussion.

      This is an excellent point. We have addressed this in our response to point (3) from Reviewer #2 in the public comments, so we will not reiterate it here.

      (3) Figure 10D/E: THIP Rescue of LTM Deficits:

      Some inconsistencies in the THIP rescue experiments need clarification:

      (a) In Figure 10D, MB299B activation with THIP appears not to significantly restore memory relative to zero, nor to differ from untreated conditions in Figures 7A or 10E.

      (b) In Figure 10E, MB299B activation +/- THIP shows a much clearer effect.

      (c) Are Figures 7A, 10E, and 10D independent experiments, or were they conducted together?

      (d) Should the left bar in 10D and the right bar in 10E be identical? If not, I do not fully understand the discrepancy and suggest discussing the variation.

      Upon revisiting the raw datasets and conducting a one-sample t-test to analyze the group differences, the experimental group in Figure 7A showed no significant difference from the theoretical mean (set at zero). This group also did not differ from the two genetic controls, indicating that the restored memory was comparable to control levels. In Figure 10E, the group with MB299B activation plus THIP treatment exhibited a significant difference from the theoretical mean (one-sample t-test) and from the non-THIP control group, confirming a significant restoration of memory function. Owing to our laboratory relocation, the starvation duration at the new facility was adjusted based on a recalibrated starvation curve; the higher overall 24 h memory index in Figure 10E is likely attributable to a relatively longer starvation period. However, this experimental parameter variation does not alter the study’s overall conclusions.

      (4) Sleep Phenotypes and Starvation Effects:

      Sleep scores are shown under starvation/fed conditions but not under baseline conditions (without inhibition/activation). Could the authors indicate whether they observe basal starvation-induced sleep changes? The authors frequently state that PAM-DPM effects on sleep are context-dependent, yet mild but significant changes occur under fed conditions. I suggest rewording to clarify that the effect is enhanced in a context-dependent manner rather than strictly context-dependent.

      Starvation-induced sleep reduction is a well-characterised phenotype. Our study focused on the key question of whether altered neuronal activity modulates sleep under innate starvation conditions. Accordingly, all comparisons were made between the experimental and control groups under both starvation and fed conditions. We appreciate the reviewer’s suggestion to improve clarity and have revised the text as suggested.

      (5) Starvation Duration in Methods:

      The authors use 30-46h of starvation, which is longer than the ~20h typically used in appetitive memory studies. Could the authors explain why such extended starvation times were necessary?

      Determining starvation levels via survival curves is a well-established and relatively objective method, one that has been widely adopted in prior studies. For the memory test, we standardized the total starvation duration for each genotype to the time point at which mortality reached 20%. Owing to inherent differences in to starvation resistance across distinct genotypes, the final starvation durations ranged from 20 hours to 46 hours.

      (6) Variability in PAM-α1 Sleep Effects:

      (a) The extent and timing of sleep effects differ across PAM-α1 drivers (e.g. night vs. light-period effects). Could MBON co-targeting by these drivers contribute to the variability?

      We have supplemented additional experiments to investigate the effects of MBON-α1 neurons on 24h memory and sleep. For detailed findings, please refer to our response to your public comments.

      (b) Even within the same driver, results differ (e.g., Figure 7H vs. 10A). A general comment on these differences would be important, e.g. regarding the relevance of day and night sleep for memory consolidation.

      We sincerely appreciate the reviewer’s incisive observation regarding these details. The discrepancy stems from the timing of neuronal activity inhibition, during which a laboratory relocation led to adjustments in starvation duration for memory experiments, which in turn indirectly altered sleep patterns.

      (c) Technical note: Similar y-axis scales for sleep plots (Figures 10A and B) would make comparison easier.

      We have unified the y-axis scales to the same range.

      (7) Discussion, Line 376:

      The phrase "sleep deprivation is important for memory consolidation" is misleading, as it could imply that deprivation aids memory formation. Please clarify.

      We appreciate the reviewer’s suggestion. We have revised the text to: “These results demonstrate that preserving unperturbed sleep during the critical memory consolidation window is essential for stabilizing appetitive long-term memory.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study builds upon a major theoretical account of value-based choice, the 'attentional drift diffusion model' (aDDM), and examines whether and how this might be implemented in the human brain using functional magnetic resonance imaging (fMRI). The aDDM states that the process of internal evidence accumulation across time should be weighted by the decision maker's gaze, with more weight being assigned to the currently fixated item. The present study aims to test whether there are (a) regions of the brain where signals related to the currently presented value are affected by the participant's gaze; (b) regions of the brain where previously accumulated information is weighted by gaze.

      To examine this, the authors developed a novel paradigm that allowed them to dissociate currently and previously presented evidence, at a timescale amenable to measuring neural responses with fMRI. They asked participants to choose between bundles or 'lotteries' of food times, which they revealed sequentially and slowly to the participant across time. This allowed modelling of the haemodynamic response to each new observation in the lottery, separately for previously accumulated and currently presented evidence.

      Using this approach, they find that regions of the brain supporting valuation (vmPFC and ventral striatum) have responses reflecting gaze-weighted valuation of the currently presented item, where as regions previously associated with evidence accumulation (preSMA and IPS) have responses reflected gaze-weighted modulation of previously accumulated evidence.

      A major strength of the current paper is the design of the task, nicely allowing the researchers to examine evidence accumulation across time despite using a technique with poor temporal resolution. The dissociation between currently presented and previously accumulated evidence in different brain regions in GLM1 (before gazeweighting), as presented in Figure 5, is already compelling. The result that regions such as preSMA response positively to |AV| (absolute difference in accumulated value) is particularly interesting, as it would seem that the 'decision conflict' account of this region's activity might predict the exact opposite result. Additionally, the behaviour has been well modelled at the end of the paper when examining temporal weighting functions across the multiple samples.

      In response to reviewer comments, the authors have explicitly tested for the effects of gaze-weighting over and above any main effect of value, and convincingly shown that these effects are both present in the main regions of interest - namely |SV| and gazeweighted |SV| in the vmPFC, alongside |AV| and |AV_gaze| in the pre-SMA. This provides clear evidence in support of the notion of gaze-weighting of value signals in these regions.

      We thank the reviewer for their comments.

      Reviewer #2 (Public review):

      Summary:

      In this paper the authors seek to disentangle brain areas that encode the subjective value of individual stimuli/items (input regions) from those that accumulate those values into decision variables (integrators) for value-based choice. The authors used a novel task in which stimulus presentation was slowed down to ensure that such a dissociation was possible using fMRI despite its relatively low temporal resolution. In addition, the authors leveraged the fact that gaze increases item value, providing a means of distinguishing brain regions that encode decision variables from those that encode other quantities such as conflict or time-on-task. The authors adopt a region-of-interest approach based on an extensive previous literature and found that the ventral striatum and vmPFC correlated with the item values and not their accumulation whereas the preSMA, IPS and dlPFC correlated more strongly with their accumulation. Further analysis revealed that the pre-SMA was the only one of the three integrator regions to also exhibit gaze modulation.

      The study uses a highly innovative design and addresses an important and timely topic. The manuscript is well-written and engaging, while the data analysis appears highly rigorous.

      Weaknesses:

      With 23 subjects the study has relatively low statistical power for fMRI although the within-subjects design and relatively high trial count reduces these concerns.

      We thank the reviewer for their comments.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Something seems to have gone slightly wrong in (I think) the labelling of new figure 7 (the correlation matrix between the different regressors). There are five variables on both the x- and y-axes of the figure, and they are the same five variables - meaning the diagonal of the matrix would be all equal to 1 (being the correlation of a regressor with itself - e.g. |AVgaze| with |AVgaze|). In the figure legend, six variables are mentioned, including lagged |deltaAVgaze| - but this doesn't appear on the x or y-axes. I suspect that the authors may need to check that this matrix has been calculated correctly, and isn't being mislabelled?

      We thank the reviewer for noticing this issue. We have now corrected Figure 7.

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank the Reviewers for the favourable feedback. There is no additional comments from Reviewers 1, 3, and 4, and we address the minor concerns from Reviewer 2 as follows.

      (1) I appreciate the authors acknowledge that testing the physical properties of the degradasome puncta is necessary to explore whether they indeed represent condensates. The term "condensates" implies liquid-liquid phase separation (rightly or wrongly). However, this question has not yet been resolved in the case of degradasomes. I therefore suggest the term "condensates" to be avoided. A simple morphological description as "puncta" may suffice.

      We have revised our manuscript to state: “The DC has been proposed to exist as biomolecular condensates.” Additionally, we have included a time-lapse image showing that AXIN1-GFP puncta exhibit dynamic fusion behaviour in cells (Fig. S8), suggesting that the DC may be liquid-like, at least with AXIN1 overexpression.

      (2) I thank the authors for including the additional data comparing tankyrase binding by IWR and IWRPOMA. I agree that using the BRET signal of IWR-POMA is informative. Adding the IC<sub>50</sub> values directly to the figure panels (S3E, S3G) would help the reader to quickly assess binding. The comparison between these two panels is insightful.

      Added.

      (3) Regarding the use of the terms TNKS, TNKS1 and TNKS2, if the authors would like to use the name "TNKS" to refer to both paralogues collectively, can this please be specified early in the manuscript to limit confusion with the official gene name "TNKS", which of course only refers to one paralogue? Regarding the use of the terms TNKS, TNKS1 and TNKS2, if the authors would like to use the name "TNKS" to refer to both paralogues collectively, can this please be specified early in the manuscript to limit confusion with the official gene name "TNKS", which of course only refers to one paralogue?

      We now specify in the Introduction that TNKS1/2 are encoded by TNKS/TNKS2, and are collectively referred to as TNKS in this manuscript.

    1. Author response:

      The following is the authors’ response to the current reviews.

      We thank the editor and the reviewers for their comments on the revised manuscript. Based on the comments, we decided to go for a minor revision which will address all the comments of reviewer 1.

      Towards the comments of the reviewer 2, we would like to state that we already provided the results from the new experiments and the reasons why we did not perform some of the suggested ones. Interestingly, we have not deviated from standard practices in the field in our approaches. Yet, the reviewer is not convinced and raised concern about the robustness of our observations. We therefore decided to carry out a few more control experiments which in our opinion are redundant as they already were carried out multiple times by us as well as by the other field experts under identical conditions and using the identical cell lines.

      Reviewer #1 (Public review):

      Summary:

      This study identifies a mechanism responsible for the accumulation of the MET receptor in invadopodia, following stimulation of Triple-negative breast cancer (TNBC) cells with HGF. HGF-driven accumulation and activation of MET in invadopodia causes the degradation of the extracellular matrix promoting cancer cell invasion, a process here investigated using gelatine-degradation and spheroid invasion assays.

      Mechanistically, HGF stimulates the recycling of MET from RAB14-positive endodomes to invadopodia, increasing their formation. At invadopodia, MET induces matrix degradation via direct binding with the metallo protease MT1-MMP.

      The delivery of MET from the recycling compartment to invadopodia is mediated by RCP which facilitates the colocalization of MET to RAB14 endosomes. On this compartment, HGF induces the recruitment of the motor protein KIF16B promoting the tubulation of the RAB14-MET recycling endosomes to the cell surface.

      This pathway is critical for the HGF-driven invasive properties of TNBC cells as it is impaired upon silencing of RAB14.

      Strengths:

      The study is well organized and executed using state of the art technology. The effects of MET recycling in the formation of functional invadopodia are carefully studied taking advantage of mutant forms of the receptor that are degradation-resistant or endocytosisdefective.

      Data analyses are rigorous and appropriate controls are used in most of the assays to assess the specificity of the scored effects. Overall, the quality of the research is high.

      The conclusions are well supported by the results and the data and methodology are of interest for a wide audience of cell biologists.

      Previous Weaknesses:

      The role of the MET receptor in invadopodia formation and cancer cell dissemination has been intensively studied in many settings including Triple Negative breast cancer cells. The novelty of the present study mostly consists in the detailed molecular description of the underlying mechanism based on HGF-driven MET recycling. The question of whether the identified pathway is specific for TNBC cells or represents a general mechanism of HGFmediated invasion detectable in other cancer cells is not addressed or at least discussed.

      Comments on revised version:

      The authors have partially replied to my previous concerns.

      We sincerely thank the reviewer for careful evaluation of our manuscript and recognizing the strength of our study. We are grateful for the positive assessment about the well-executed methods, rigorous data analysis, usage of appropriate controls, and for acknowledging that we have addressed, at least in part, the concerns raised in the previous round of review. We appreciate the reviewer’s constructive comments, which have helped us further clarify the scope and significance of our findings.

      Reviewer #1 (Recommendations for the authors):

      The authors have partially replied to my previous concerns. The following points still have to be addressed.

      (1) Despite many TNBC tumours present high expression of the EGFR, trials with EGFR inhibitors have been very disappointing (please read PMID: 41651315). EGFR inhibition has been extensively attempted with negative outcome, and this is well known.

      In the clinical practise, Triple Negative breast cancer patients are commonly treated with chemotherapy, not with EGFR or MET inhibitors.

      Line 31-38 are misleading at best and should be removed and the incipit of the study modified. The biology of this study is sound there is no need to push the clinical relevance with instances that are notoriously not applicable.

      We are thankful to the reviewer for bringing up this point. We will modify the manuscript as per the reviewer’s suggestion.

      (2) In reply to point 4, the authors claim that they checked MMP2 and the experiment is shown in FigS5I, which is not......

      We sincerely apologise for uploading an incorrect file. The correct file will be uploaded.

      (3) I noticed that, in the previous version of the manuscript in Fig. S4A the panel showing the mutation frequency was erroneously indicated in the legend as referring to MET.

      The authors replied that they fixed this but actually the legend is still wrong...

      We thank the reviewer for bringing this error into our attention. We will attentively correct the legend in the revised version of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Khamari and colleagues investigate how HGF-MET signaling and the intracellular trafficking of the MET receptor tyrosine kinase influence invadopodia formation and invasion in triple-negative breast cancer (TNBC) cells. They show that HGF stimulation enhances both the number of invadopodia and their proteolytic activity. Mechanistically, the authors demonstrate that HGF-induced, RAB4- and RCP-RAB14KIF16B-dependent recycling routes deliver MET to the cell surface specifically at sites where invadopodia form. Moreover, they report that MET physically interacts with MT1MMP - a key transmembrane metalloproteinase required for invadopodia function- and that these two proteins co-traffic to invadopodia upon HGF stimulation.

      Although the HGF-MET axis has previously been implicated in invadopodia regulation (e.g., by Rajadurai et al., Journal of Cell Science 2012), studies directly linking ligandinduced MET trafficking with the spatial regulation of MT1-MMP localization and activity have been lacking.

      Overall, the manuscript addresses a relevant and timely topic and provides several novel insights.

      Comments on revised version:

      I appreciate the authors' efforts to revise the manuscript and address the reviewers' comments. While the revised version includes additional experiments and several improvements in data presentation, the major methodological and conceptual concerns raised in the initial review remain largely unresolved. In my opinion, these issues critically undermine the central mechanistic conclusions of the study.

      We thank the reviewer for critically re-evaluating our manuscript and acknowledging the novel insights and relevance of our study.

      We thank the reviewer for pointing out the study by Rajadurai et al., Journal of Cell Science, 2012, which provided crucial evidence about the role of MET signaling in invadopodia formation [1]. However, the experimental system largely used by Rajadurai et al. is fundamentally different from the receptor trafficking mechanism investigated in the present study. The group have mostly used overexpression of Tpr-MET, which is a cytosolic MET mutant, that does not undergo the canonical ligand-induced RTK endocytosis and subsequent degradation or recycling. Our study did not only establish another link between MET signaling and invadopodia formation; rather, we identified a trafficking-dependent mechanism whereby HGF stimulation regulates the spatial redistribution and recycling of full-length MET to invadopodia, thus providing more physiologically relevant insights.

      We also respectfully disagree that the methodological concerns raised critically undermine our mechanistic conclusions. Although other approaches as suggested by the reviewer could provide complementary information, we believe that the methods used in our study are appropriate for assessing invasive behaviour of the TNBC cells. These methods have been used by us and other research groups in the filed as reflected from the existing literature [2–6].

      In this context, we would also like to add that the study referred by the reviewer above, has used a more off target prone approach (SiGenome Smartpool) compared to ONTARGETplus (chemically modified SiRNA pool for minimizing off target effect) in addition to the same small molecule inhibitor used in our study. Additionally, we also used a shRNA-based silencing to verify the phenotype. Since the silencing was not as pronounced as the siRNA-mediated knockdown, the reviewer has expressed concern.

      We would like to point out that the antibody we used in IF for MT1-MMP (MMP14) have been published in multiple peer-reviewed journals by various research groups using the identical cell lines (MDA-MB-231, ATCC- HTB-26) [6–8]. MET antibodies used in the study (CST, D1C2 XP & L6E7) has been validated in MDA-MB-231 and MET-depleted cells [9-12]. Since these antibodies have been utilized for IF since a long time across various research groups under identical laboratory/experimental conditions, the exercise of validation suggested by the reviewer is surprising.

      (1) Inappropriate experimental design for studying MET trafficking

      A major concern remains the use of prolonged HGF stimulation times (2-6 hours) to study MET endocytosis and recycling. This is not an appropriate experimental design for investigating receptor tyrosine kinase trafficking dynamics. Ligand-induced internalization of MET occurs within minutes, with maximal endosomal accumulation typically observed within 5-15 minutes, whereas recycling occurs over approximately 15-60 minutes.

      Importantly, the authors have not included short stimulation time points or any kinetic analysis that would allow a proper assessment of MET internalization or recycling. The additional surface biotinylation experiment does not address this issue, as it still does not provide temporal information regarding receptor trafficking.

      Therefore, the current data do not support the conclusions regarding MET endocytosis or recycling, and this major methodological concern has not been adequately addressed in the revised manuscript.

      We want to clarify that, our prime objective is to determine how MET trafficking is regulated at the time points at which we observe the functional effects of HGF on invadopodia formation and matrix degradation. Since our functional assays were performed following 2-3 h of HGF stimulation, we specifically examined MET localization and trafficking at these same time points.

      Though shorter time points could provide information on the kinetics of MET trafficking, but their absence does not invalidate our conclusions regarding the role of MET trafficking in HGF-induced invasive function. In other words, our conclusions are made for the time points for which we have conducted the experiments. It is needless to add that different cargo molecule will show different kinetics.

      In summary, we wanted to study the MET trafficking at the late hours in accordance with our functional assays and accordingly designed our experiments. Also, we have supported our results through biochemical methods which is considered to be one of the gold standards in the field.

      (2) Insufficient validation of antibody specificity in immunofluorescence

      The validation of antibody specificity for MET, phospho-MET, and MT1-MMP in immunofluorescence experiments remains insufficient. While the authors demonstrate knockdown efficiency by immunoblotting and show some reduction in fluorescence signal, they do not provide rigorous evidence that the immunofluorescence signal is specifically abolished upon gene silencing under identical imaging conditions. Such validation is essential, particularly because the manuscript relies heavily on imaging-based localization and colocalization analyses. Without these controls, it cannot be excluded that the observed signal represents non-specific staining.

      Importantly, the authors attempt to justify antibody specificity primarily by citing previous publications that used the same antibodies. However, this is not an adequate substitute for experimental validation within the current study. Previous reports do not guarantee specificity under the present experimental conditions, particularly in immunofluorescence, where staining patterns can be strongly influenced by fixation procedures, antibody concentrations, imaging settings, and cell type. Moreover, those studies may themselves lack sufficiently rigorous validation of antibody specificity. Therefore, antibody specificity should be demonstrated directly in the experimental system used in this manuscript, especially given that the principal conclusions rely extensively on the subcellular localization of MET, phospho-MET, and MT1-MMP.

      We understand the reviewer’s concern regarding antibody specificity and agree that appropriate validation is important for imaging-based analyses. However, we strongly disagree with the statement that the MT1-MMP, MET antibody were not adequately validated. The specificity of the MT1-MMP antibody was independently validated by both siRNA- and sgRNA-mediated gene silencing, where we observed a substantially diminished MT1-MMP signal by immunoblotting (Fig S5I, M’). Moreover, the antibody has been used for IF in the same cell line by multiple research groups [6–8]. So, in our opinion, this validation is completely redundant.

      We have validated the MET staining/ signal using the antibody in gene-silenced cells by immunoblotting and immunofluorescence (Fig S1I, K, L), as also acknowledged by the reviewer in comment-4. In addition, we would also like to clarify that the references cited in support of antibody specificity were not selected simply because they used the same antibodies. They include studies that provide experimental validation of the antibodies by gene silencing.

      To further confirm the antibody specificity, we will add immunofluorescence images of MET or MT1-MMP silenced cells stained with respective antibodies. However, we may not want to add these data to the manuscript as they do not carry any additional values to the manuscript.

      (3) Questionable MET localization in TIRF microscopy

      The presence of punctate MET signal in TIRF microscopy under unstimulated conditions raises additional concerns. Under basal conditions, MET is generally expected to exhibit a predominantly diffuse distribution at the plasma membrane, whereas prominent punctate structures are typically associated with ligand-induced clustering, endocytosis, or trafficking events.

      The observation of numerous MET-positive puncta in unstimulated cells, together with the insufficient validation of antibody specificity, raises the possibility that at least part of the observed signal represents non-specific staining or imaging artefacts rather than bona fide MET localization. This concern is further compounded by the lack of rigorous immunofluorescence antibody validation discussed above and significantly undermines the interpretation of all TIRF-based trafficking analyses presented in the manuscript.

      We would like to highlight the apparent similarities between Fig 2A, B and the unstimulated condition in Fig. 2H. In figure 2A, B, MET is detected using an anti-MET antibody, whereas in Figure 2H, GFP-MET is imaged under live cell condition. We believe the reviewer would agree that imaging GFP-MET in live cells avoids fixation- and antibody-related artifacts. The comparable localization observed using these two independent approaches therefore provides additional support that the MET distribution shown in Fig. 2A, B reflects genuine receptor localization rather than an imaging or staining artefact.

      (4) The evidence supporting a MET-specific role in invadopodia remains unconvincing

      The authors argue that the role of MET in invadopodia formation is validated using three independent approaches: shRNA-mediated knockdown, SMARTpool siRNA-mediated knockdown, and pharmacological inhibition with PHA665752. However, I do not agree that these constitute three independent orthogonal validations of MET function.

      First, the shRNA-mediated knockdown presented in this study achieves only modest depletion of MET protein. The authors themselves acknowledge this limitation and therefore selected cells with visibly reduced MET staining for imaging. Consequently, the shRNA experiments cannot be considered a robust or independent validation of MET function.

      Second, although pooled SMARTpool siRNAs are widely used to improve knockdown efficiency, they cannot exclude off-target effects, as each individual guide RNA contributes its own potential off-target profile. Therefore, pooled siRNAs cannot by themselves establish that an observed phenotype is specifically attributable to depletion of the intended target and do not replace validation using independent individual siRNAs or rescue experiments.

      Third, the pharmacological data should also be interpreted with caution. Throughout the manuscript, PHA665752 is presented as a MET inhibitor supporting the specificity of the observed phenotype. However, there is essentially no such thing as a truly selective receptor tyrosine kinase inhibitor. PHA665752 inhibits multiple kinases in addition to MET, particularly at concentrations commonly used in cell-based assays. Consequently, the inhibitor cannot be considered an independent validation of MET-specific function.

      Importantly, the newly added siRNA experiments do not resolve my original concern regarding the role of MET in invadopodia formation. Although siRNA-mediated MET depletion is substantially more efficient than the shRNA-mediated knockdown presented in the original manuscript, this marked difference in MET depletion is not accompanied by a correspondingly stronger inhibition of invadopodia formation or ECM degradation. If MET were indeed the principal driver of the observed phenotype, one would expect the magnitude of the biological effect to correlate with the efficiency of MET depletion. This inconsistency raises the possibility that the observed phenotype is not solely attributable to MET depletion and calls into question the specificity of the proposed mechanism.

      Taken together, the three perturbation approaches used by the authors cannot be regarded as independent orthogonal validation of MET function. One approach provides only modest target depletion, another relies on pooled RNAi reagents that cannot exclude off-target effects, and the third employs a multi-kinase inhibitor rather than a MET-specific compound. Collectively, these limitations substantially weaken the conclusion that the reduction in invadopodia formation is specifically attributable to loss of MET. A convincing demonstration of MET-specific function would require rescue experiments or another truly orthogonal validation strategy.

      We had adopted three independent approaches to validate the phenotype. All three approaches are well practiced in the field. The small molecule inhibitor has been used in the study by Rajadurai et al, J Cell Science, 2012 and it is very much accepted in studying cellular kinases [1].

      We agree that even though the smart pool has always chance of off-target effects, the OnTargetPlus Smart pool has the minimum chance of off-target effects because of the patented chemical modifications, compared to the individual oligos and SiGenome SMARTpool, which was used by Rajadurai, et. al. in their study [1].

      The shRNA mediated silencing resulted in less reduction in the MET level (~50-60%) but is it scientifically not acceptable, particularly when it showed similar phenotype over n=3 sets of experiments?

      The arguments made by the reviewer in this context seems to be harsh. However, we decided to carry out MET silencing using two independent oligos from the SMART pool.

      (5) Weak evidence for MET-MT1-MMP interaction

      The evidence supporting a physical interaction between MET and MT1-MMP remains unconvincing. The newly added co-immunoprecipitation experiment does not reveal a convincing MET-MT1-MMP interaction, and I am unable to appreciate a specific coimmunoprecipitated MT1-MMP signal in the presented blot. As presented, these data do not convincingly demonstrate a specific or functionally relevant interaction. Given that this interaction constitutes a central component of the proposed mechanistic model, this remains a major weakness of the study.

      We have detected the interaction in both GFP pulldown assay in 4 different cell lines and the corresponding reverse His-Ni-NTA pulldown assay, providing complementary evidence for their physical association (Fig 6E, S5F, G). We have also clearly stated in the manuscript that this interaction is weak in nature and have not claimed it to be a strong interaction. While we acknowledge that the signal is modest, disregarding reproducible positive results would not be an appropriate interpretation of the pulldown assays. We believe the reproducibility of these findings supports a genuine MET and MT1-MMP association, and we have reported it accordingly.

      (6) Overinterpretation of the data

      Taken together, the study proposes a mechanistic model linking MET trafficking to MT1MMP localization and invadopodia function. However, the experimental evidence largely supports correlative observations rather than demonstrating a direct mechanistic relationship.

      Specifically, MET endocytosis and recycling are not properly demonstrated because of the inappropriate temporal resolution of the trafficking experiments; the localization data remain uncertain owing to insufficient validation of the immunofluorescence reagents; and the proposed interaction between MET and MT1-MMP is not convincingly demonstrated. Consequently, the manuscript establishes correlation rather than causality, and the central mechanistic conclusions appear to be substantially overstated relative to the presented data.

      We agree that there is scope for further investigation of MET and MT1-MMP cotrafficking, however, we have provided preliminary evidence demonstrating the cotrafficking of MET and MT1-MMP at the cell surface (Fig 6F). Importantly, MET and MT1-MMP co-trafficking represents only one component of the manuscript and not a central mechanistic conclusion. As the title of the study indicates, the major component of the study is focused on MET trafficking and its implication in invadopodia-associated TNBC invasion, for which we have provided direct experimental evidence. Therefore, we believe that describing the overall study as primarily overstated and correlative underestimates the extent of the experimental evidence supporting our mechanistic conclusions.

      Conclusion:

      While the manuscript addresses an interesting and biologically relevant question, the current experimental evidence does not adequately support the proposed mechanistic model. The combination of inappropriate experimental design for trafficking studies, insufficient validation of key imaging reagents, questionable interpretation of the localization data, lack of convincing evidence for the proposed MET-MT1-MMP interaction, and the absence of a clear relationship between the degree of MET depletion and the biological phenotype substantially limits the reliability of the conclusions.

      In my opinion, these issues cannot be addressed by further revision of the current manuscript, as they require substantial additional experimentation, including appropriately designed trafficking assays with short kinetic time points, rigorous validation of antibody specificity for immunofluorescence, and stronger mechanistic evidence linking MET trafficking to MT1-MMP-dependent invadopodia function.

      We thank the reviewer for outlining the remaining concerns. We will address the points raised by performing additional antibody validation and independent oligo-mediated MET silencing experiment, providing further support for the specificity and robustness of our findings.

      However, we respectfully disagree, that the conclusions require the extensive additional experimentation suggested by the reviewer. As clarified above, our trafficking experiments were designed around the time points at which the functional invasive phenotype is observed, rather than to define the kinetics of MET internalization.

      In conclusion, we would expect that the views of the reviewer 2 towards the manuscript should change and the reliability of our manuscript to the public should improve.

      References:

      (1) Rajadurai CV, Havrylov S, Zaoui K, Vaillancourt R, Stuible M, Naujokas M, Zuo D, Tremblay ML, Park M. Met receptor tyrosine kinase signals through a cortactinGab1 scaffold complex, to mediate invadopodia. J Cell Sci. 2012 Jun 15;125(Pt 12):2940-53. doi: 10.1242/jcs.100834. Epub 2012 Feb 24. PMID: 22366451; PMCID: PMC3434810.

      (2) Sharma P, Parveen S, Vinod Shah L, Mukherjee M, Kalaidzidis Y, Joseph Kozielski A, et al. Title: SNX27-retromer assembly directs MT1-MMP trafficking to invadopodia and promotes breast cancer metastasis.

      (3) Mader CC, Oser M, Magalhaes MAO, Bravo-Cordero JJ, Condeelis J, Koleske AJ, et al. Molecular and Cellular Pathobiology An EGFR-Src-Arg-Cortactin Pathway Mediates Functional Maturation of Invadopodia and Breast Cancer Cell Invasion [Internet]. doi:10.1158/0008-5472.CAN-10-1432

      (4) Parveen S, Khamari A, Raju J, Coppolino MG, Datta S. Syntaxin 7 contributes to breast cancer cell invasion by promoting invadopodia formation. J Cell Sci. 2022 Jun 15;135(12). doi:10.1242/jcs.259576 PubMed PMID: 35762511.

      (5) Joffre C, Barrow R, Ménard L, Calleja V, Hart IR, Kermorgant S. A direct role for Met endocytosis in tumorigenesis. Nat Cell Biol. 2011 Jun 5;13(7):827–37. doi:10.1038/ncb2257 PubMed PMID: 21642981.

      (6) Monteiro P, Rossé C, Castro-Castro A, Irondelle M, Lagoutte E, Paul-Gilloteaux P, et al. Endosomal WASH and exocyst complexes control exocytosis of MT1-MMP at invadopodia. J Cell Biol. 2013 Dec 23;203(6):1063–79. doi:10.1083/jcb.201306162 PubMed PMID: 24344185.

      (7) Wenzel EM, Pedersen NM, Elfmark LA, Wang L, Kjos I, Stang E, et al. Intercellular transfer of cancer cell invasiveness via endosome-mediated protease shedding. Nat Commun. 2024 Feb 10;15(1):1277. doi:10.1038/s41467-024-45558-8

      (8) Pedersen NM, Wenzel EM, Wang L, Antoine S, Chavrier P, Stenmark H, Raiborg C. Protrudin-mediated ER-endosome contact sites promote MT1-MMP exocytosis and cell invasion. J Cell Biol. 2020 Aug 3;219(8):e202003063. doi: 10.1083/jcb.202003063. PMID: 32479595; PMCID: PMC7401796.

      (9) Duan Q, Jia HR, Chen W, Qin C, Zhang K, Jia F, Fu T, Wei Y, Fan M, Wu Q, Tan W. Multivalent Aptamer-Based Lysosome-Targeting Chimeras (LYTACs) Platform for Mono- or Dual-Targeted Proteins Degradation on Cell Surface. Adv Sci (Weinh). 2024 May;11(17):e2308924. doi: 10.1002/advs.202308924. Epub 2024 Feb 29. PMID: 38425146; PMCID: PMC11077639.

      (10) Wei J, Wang J, Guan W, Li J, Pu T, Corey E, Lin TP, Gao AC, Wu BJ. PlexinD1 is a driver and a therapeutic target in advanced prostate cancer. EMBO Mol Med. 2025 Feb;17(2):336-364. doi: 10.1038/s44321-024-00186-z. Epub 2025 Jan 2. PMID: 39748059; PMCID: PMC11822115.

      (11) Yamasaki A, Miyake R, Hara Y, Okuno H, Imaida T, Okita K, Okazaki S, Akiyama Y, Hirotani K, Endo Y, Masuko K, Masuko T, Tomioka Y. Dual-targeting therapy against HER3/MET in human colorectal cancers. Cancer Med. 2023 Apr;12(8):9684-9696. doi: 10.1002/cam4.5673. Epub 2023 Feb 7. PMID: 36751113; PMCID: PMC10166911.

      (12) Soonnarong R, Putra ID, Sriratanasak N, Sritularak B, Chanvorachote P. Artonin F Induces the Ubiquitin-Proteasomal Degradation of c-Met and Decreases AktmTOR Signaling. Pharmaceuticals (Basel). 2022 May 21;15(5):633. doi: 10.3390/ph15050633. PMID: 35631459; PMCID: PMC9145792.


      The following is the authors’ response to the original reviews.

      We sincerely thank the editor and reviewers for thoroughly evaluating the manuscript. Following the comments from the reviewers we caried out four major sets of experiments and added the results and the conclusions derived from them in the revised manuscript. We also modified the abstract and the introduction. As suggested by the reviewers, we have rewritten the discussion. All the mislabelling and typing errors have been corrected, and representative graphs has been replaced as suggested. The list of the newly carried out experiments are -

      (i) In the original submission, we carried out a microscopy-based study to investigate the recycling of MET. We now added surface biotinylation approach to show the RCP or KIF16B-mediated surface delivery of MET (Fig. 5J).

      (ii) To demonstrate the functional effect of KIF16B silencing on TNBC invasion we have performed ECM degradation assay with depleted KIF16B cells (Fig. S4K-L). Further, the MET degradation in KIF16B-silenced cells has also been investigated using immunoblotting (Fig. S4H).

      (iii) To rule out the off-target effect of the siRNA used in this study, we have validated the invadopodia-associated function using 2 individual siRNAs for RAB14 and RCP (Fig. S4K-L).

      (iv) We have now introduced MMP2 as a positive control as a substrate of MT1-MMP to show the effect of silencing of the protease on its cleavage (Fig. S5I).

      We also incorporated following changes, largely additions of new plots, data in the revised manuscript.

      (i) RAB4 and RAB14 colocalization with MET in BT-549 cell line has also been added in Fig. 4 (A, B, C).

      (ii) Data showing MET silencing using siRNA and its effect on invadopodia has been added to Fig 1D.

      (iii) Graph showing percentage of cells forming invadopodia has added to Fig. S1G.

      (iv) Line intensity plots of MET-containing invadopodia has been added in Fig. 2G’.

      (v) We have added quantification of all the blots to the figures.

      (vi) A graph representing MET degradation kinetics with HGF over 3 experiments has been added to Fig S2F.

      (vii) The full field of view of Fig 1F has been added in the Fig S1L. A quantification of the gelatin degradation has been added to Fig 1F.

      (viii) The blot for loading control of Fig S1K has been changed from Actin to Vinculin.

      (ix) The blot showing expression of MT1-MMP in the SCR and knockout cells has been added to Fig S5M’.

      (x) The survival plot for patients with altered or unaltered MET and RCP has been removed. The graph showing frequency alteration of MET has also been removed.

      (xi) Additional immunoblots associated with all the figures are now provided in a newly added supplementary figure (Fig S6).

      Reviewer #1 (Public review):

      Summary:

      This study identifies a mechanism responsible for the accumulation of the MET receptor in invadopodia, following stimulation of Triple-negative breast cancer (TNBC) cells with HGF. HGF-driven accumulation and activation of MET in invadopodia causes the degradation of the extracellular matrix, promoting cancer cell invasion, a process here investigated using gelatin-degradation and spheroid invasion assays.

      Mechanistically, HGF stimulates the recycling of MET from RAB14-positive endosomes to invadopodia, increasing their formation. At invadopodia, MET induces matrix degradation via direct binding with the metalloprotease MT1-MMP. The delivery of MET from the recycling compartment to invadopodia is mediated by RCP, which facilitates the colocalization of MET to RAB14 endosomes. In this compartment, HGF induces the recruitment of the motor protein KIF16B, promoting the tubulation of the RAB14-MET recycling endosomes to the cell surface. This pathway is critical for the HGF-driven invasive properties of TNBC cells, as it is impaired upon silencing of RAB14.

      Strengths:

      The study is well-organized and executed using state-of-the-art technology. The effects of MET recycling in the formation of functional invadopodia are carefully studied, taking advantage of mutant forms of the receptor that are degradation-resistant or endocytosis-defective.

      Data analyses are rigorous, and appropriate controls are used in most of the assays to assess the specificity of the scored effects. Overall, the quality of the research is high.

      The conclusions are well-supported by the results, and the data and methodology are of interest for a wide audience of cell biologists.

      We sincerely thank the reviewer for the positive feedback and for considering our study to be well executed and rigorous. The valuable suggestions and comments certainly improved the understanding of the role of the RAB14-RCP-KIF16B axis in MET trafficking and breast cancer invasion.

      Weakness

      The role of the MET receptor in invadopodia formation and cancer cell dissemination has been intensively studied in many settings, including triple-negative breast cancer cells. The novelty of the present study mostly consists of the detailed molecular description of the underlying mechanism based on HGF-driven MET recycling. The question of whether the identified pathway is specific for TNBC cells or represents a general mechanism of HGF-mediated invasion detectable in other cancer cells is not addressed or at least discussed

      We thank the reviewer for raising this point. We would like to clarify that in TNBCs, the overexpression of EGFR and MET in the null background of the hormonal receptors; progesterone receptor, estrogen receptor, and HER2 is considered to be very crucial in terms of prognosis and treatment (PMID: 27655711, 25368674). Hence study of MET signalling and trafficking is more relevant for TNBCs compared to other cancer cells. In the current study, we therefore focused on two TNBC cell lines. We have added this in the first paragraph of introduction. Line no: 31-38.

      Reviewer #1 (Recommendations for the authors):

      Major points

      (1) My major concern refers to the clinical data presented in this study. Different from the mechanistic findings, the quality of these analyses is too low, the description in the figure legends is scant, and is absent in the Method section. The results concerning the prognostic value of RCP are the most problematic. What is shown in the Kaplan Meier? Is it the correlation between the RCP mRNA levels and the patient's survival? More importantly, to study the correlation between genetic alteration and prognostic outcome, multivariable analyses should be performed comparing the genetic alteration with known prognostic factors (sex, age, tumor size, node status, ER/PrR, HER2, Ki67 if available, tumor grade). This is because breast cancer prognosis depends on multiple interrelated factors, and only multivariate models can adjust for confounding and identify which variables independently predict outcome-providing far more accurate and clinically useful prognostic information than univariate analysis. Furthermore, overexpression of RTKs has been extensively reported and studied in TNBCs. Similarly, the relevance of MET in cancer cell invasion has been firmly established. Therefore, the data presented here, whose quality does not match the mechanistic part of the study, can be considered unnecessary. I would, therefore, recommend removing the "clinical" data. The introduction should be modified accordingly.

      We thank the reviewer for this insightful comment. To generate the graphs, we selected studies available in the publicly accessible database cBioportal (cbioportal.org/). The graphs generated by the database, representing the genetic alterations of genes has added in the Fig. S4A. However, as suggested by the reviewer, we have removed the survival plots for MET and RCP, the gene alteration frequency of MET and modified the text accordingly.

      (2) Overexpression of KIF16B has been shown to inhibit the degradative pathway, stimulating the recycling one. In agreement, silencing of KIF16B accelerates EGFR degradation (PMID: 15882625). Does this also apply to the MET receptor? In the present setting, does the overexpression of KIF16B result in prolonged MET expression?

      We thank the reviewer for raising this question. Although we did not analyze the expression level of MET in the KIF16B overexpressed cells, we have analyzed the total MET levels in the control and the KIF16B silenced cells using immunoblotting and did not observe any significant changes in the MET protein levels (Fig S4H). Line no: 390-391. This led us to believe that the depletion or overexpression of KIF16B may not have any direct effect on MET expression.

      (3) The contribution of MMP2 and MMP9 to the degradative properties of HGFstimulated TNBC cells should be investigated and compared to MT1-MMP.

      We believe that this is a relevant note from the reviewer. However, MT1-MMP is one of the most well-established metalloproteases, known till date for its role in invadopodia-associated activities in breast cancer (PMID: 35008569, 27501444). Moreover, it is the best-known candidate protease, which is a transmembrane metalloprotease could be the model in studying membrane recycling to invadopodia (PMID: 20605060, 19692588). However, as pointed out by the reviewer, we completely agree that MMP2 and MMP9 also contribute significantly to invadopodia-associated functions in TNBCs (PMID: 23902685, 25699257). HGF is also known to promote the expression and activity of MMP2 and MMP9 (PMID: 23320110, 26259977). Interestingly, the cellular machineries involved in their enhanced activity due to HGF stimulation may be distinct from what was observed in the current study and may require a distinct, elaborated study, which is beyond the scope of the current one.

      (4) The authors appropriately tested the possible shedding effect of MT1-MMP on MET. They should repeat the experiments, adding a positive control of shedding. Furthermore, the legends referring to these experiments, shown in Supplementary Figure S6, seem to be wrong (or mislabeled).

      We thank the reviewer for the suggestion. MT1-MMP is known to proteolytically cleave and initiate the activation of MMP2 (PMID: 11161720, 15095267). We now carried out the experiment with MMP2 as a control (Fig S5I). We observe an increase in the unprocessed MMP2 level in the MT1-MMP silenced cells, whereas the MET levels are unaltered. Line no: 467-468.

      We sincerely apologize for the oversight in the mislabelling. We have now corrected it in the revised manuscript.

      (5) Does altered expression of RAB14 and/or KIF16B affect MT1-MMP delivery to invadopodia in the TNBC cell lines? Does KIF16B silencing affect invasion?

      The role of RAB14 and KIF16B in MT1-MMP delivery to podosomes, a structure similar to invadopodia in macrophages has been studied by Hey S. et al. (PMID: 37696580). The study suggests that KIF16B silencing reduces the invasion of macrophages. However, the effect is not known for TNBCs. Thus, we have conducted the ECM degradation assay in TNBC cell lines to show the effect of KIF16B gene silencing on breast cancer invasion to the revised manuscript (Fig S4K-L). In both MDA-MB-231 and BT-549 cells we observed reduced ECM degradation activity upon KIF16B depletion, corroborating the observation from Hey S. et al. Line no: 392-400.

      Minor points

      (1) I recommend authenticating cell lines and stable populations by STR profiling.

      All the cell lines used in the study have been purchased from ATCC, and STR is a standard practice followed by ATCC. Further, to avoid any alteration, cells were discontinued after 20 passages. We have added this statement to the methods section in the revised manuscript. Line no: 643-644.

      (2) In the legend to Supplementary Figure S4A, the panel is described as the frequency of alterations of MET, while, if I correctly interpret it, the bar graph refers to RCP. As mentioned above, these data could be removed.

      We thank the reviewer for pointing out the mistake. We have rectified this in the revised version.

      (3) Check for typos. Sometimes invadopodia is written with the capital: "Invadopodia", in other instances it is not. The authors should be consistent throughout the manuscript. English language editing would help.

      We sincerely apologize for the inconsistency in the writing. We have removed the unnecessary capitalization of invadopodia in the revised manuscript. Line no: 142, 159, 280, 436.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Khamari and colleagues investigate how HGF-MET signaling and the intracellular trafficking of the MET receptor tyrosine kinase influence invadopodia formation and invasion in triple-negative breast cancer (TNBC) cells. They show that HGF stimulation enhances both the number of invadopodia and their proteolytic activity. Mechanistically, the authors demonstrate that HGF-induced, RAB4- and RCP-RAB14-KIF16B-dependent recycling routes deliver MET to the cell surface specifically at sites where invadopodia form. Moreover, they report that MET physically interacts with MT1-MMP - a key transmembrane metalloproteinase required for invadopodia function- and that these two proteins co-traffic to invadopodia upon HGF stimulation.

      Although the HGF-MET axis has previously been implicated in invadopodia regulation (e.g., by Rajadurai et al., Journal of Cell Science 2012), studies directly linking ligand-induced MET trafficking with the spatial regulation of MT1-MMP localization and activity have been lacking.

      Overall, the manuscript addresses a relevant and timely topic and provides several novel insights. However, some sections require clearer and more concise writing (details below). In addition, the quality, reliability, and robustness of several data sets need to be improved.

      Strengths:

      A key strength of the study is the novel demonstration that HGF-mediated, RAB4- and RAB14-dependent recycling of MET delivers this receptor, together with MT1MMP, to invadopodia -highlighting a previously unrecognized mechanism, regulating the formation and proteolytic function of these invasive structures. Another strong point is the breadth of experimental approaches used and the substantial amount of supporting data. The authors also include an appropriate number of biological replicates and analyze a sufficiently large number of cells in their imaging experiments, as clearly described in the figure legends.

      We greatly appreciate the positive assessment from the reviewer, who also acknowledged the novelty and relevance of our study. Below, we have carefully addressed the comments/concerns raised regarding this study and that have strengthened the reliability and robustness by revisiting the data, providing additional analyses where required, and clarifying methodological details.

      Weakness

      (1) Inappropriate stimulation times for endocytosis and recycling assays. The experiments examining MET endocytosis and recycling following HGF stimulation appear to use inappropriate incubation times. After ligand binding, RTKs typically undergo endocytosis within minutes and reach maximal endosomal accumulation within 5-15 minutes. Although continuous stimulation allows repeated rounds of internalization, the temporal dynamics of MET trafficking should be examined across shorter time points, ideally up to 1 hour (e.g., 15, 30, and 60 minutes). The authors used 2-, 3-, or 6-hour HGF stimulation, which, in my opinion, is far too long to study ligandinduced RTK trafficking.

      We understand the reviewer’s concern regarding the HGF stimulation time point for endocytosis and recycling. We want to highlight that to study the recycling/surface delivery of MET in response to HGF, we performed TIRF microscopy-based imaging, where images were taken within 1h of HGF addition (Fig. 2I). Additionally, we have incorporated surface biotinylation to show the recycling of MET as suggested in comment-7 (Fig. 5J). For this experiment we have used 30 min of HGF stimulation. Line no: 382-391.

      Moreover, we have observed the functional effect of HGF on ECM (gelatin) degradation and invadopodia formation after 3 h of HGF stimulation. We were curious to know where does the MET localises with prolonged ligand stimulation. Hence, to study the localization of MET to invadopodia or the endocytic markers, the cells were stimulated with HGF for 2-3 hours.

      (2) Low efficiency of MET silencing in Figure S1I. The very low MET knockdown efficiency shown in Figure S1I raises concerns. Given the potential off-target effects of a single shRNA and the insufficient silencing level, it is difficult to conclude whether the reduction in invadopodia number in Figure 1F is genuinely MET-dependent. The authors later used siRNA-mediated silencing (Figure S5C), which was more effective. Why was this siRNA not used to generate the data in Figure 1F? Why did the authors rely on the inefficient shRNA C#3?

      We understand the concern raised by the reviewer. We want to emphasize that we have employed three different approaches to investigate the effect of MET silencing/inhibition on invadopodia formation. (i) A MET kinase inhibitor, PHA665752, which shows reduced invadopodia formation (Fig. 1E, E’). (PMID: 21973114, 41009793) (ii) Silencing with shRNA: Since the level of silencing of MET with the shRNA was not sufficient, cells were stained with MET as a readout for MET silencing, and images of the cells with reduced MET expression were captured. ECM degradation activity and invadopodia numbers were found to be reduced in the MET-depleted cells (Fig. 1F). (iii) We have now added the data showing the effect of siRNA-mediated MET depletion on invadopodia formation to the revised figure 1D. Line no: 123-125. To draw a robust conclusion regarding the role of MET on invadopodia-associated TNBC invasion, we have integrated all three complementary approaches.

      (3) Missing information on incubation times and inconsistencies in MET protein levels. The figure legends do not indicate how long the cells were incubated with HGF or the MET inhibitor PHA665752 before immunoblotting. This information is crucial, particularly because both HGF and PHA665752 cause a substantial decrease in the total MET protein level. Notably, such a decrease is absent in MDA-MB-231 cells treated with HGF in the presence of cycloheximide (Figure S2F). The authors should comment on these inconsistencies. Additionally, the MET bands in Figure S1J appear different from those in Figure S1C, and MET phosphorylation seems already high under basal conditions, with no further increase upon stimulation (Figure S1J). The authors should address these issues.

      We apologise for the unintentional omission of experimental detailing about HGF or drug incubation time, which we have incorporated into the figure legend appropriately. Regarding the decreased MET level in the drug-treated condition: literature suggests that the MET inhibitor PHA665752 also promotes MET degradation, corroborating our result shown in Fig. S1J (PMID: 15788682, 18327775). Further in Fig. S1J, the relative phosphorylation of MET when compared to the total MET level in the HGF-treated condition is higher (~2-fold). Quantification of the blot has been added now.

      Further, addition of HGF for 3 h leads to 40±15% reduction in the MET protein levels as seen in the Author response image 1 representing quantification of different immunoblot. The degradation of MET in the Fig S1J is 65% which nearly fall in the range for HGF-mediated MET degradation.

      Author response image 1.

      Quantification of immunoblots showing MET signal intensity in the presence or absence of HGF normalized with the loading control. N=6.

      Next, in the fig. S1A, K the rabbit anti-MET (CST, D1C2 XP) antibody has been used, which binds to a C-terminal motif of MET and identifies both the 170kDa as well as 140kDa protein representing the uncleaved and cleaved form of MET. In Fig. S1J, the mouse antiMET (CST, L6E7) antibody has been used, which binds to an N-terminal motif of MET and recognizes only the 140kDa protein.

      (4) Insufficient representation and randomization of microscopic data. For microscopy, only single representative cells are shown, rather than full fields containing multiple cells. This is particularly problematic for invadopodia analysis, as only a subset of cells forms these structures. The authors should explain how they ensured that image acquisition and quantification were randomized and unbiased. The graphs should also include the percentage of cells forming invadopodia, a standard metric in the field. Furthermore, some images include altered cells - for example, multinucleated cells - which do not accurately represent the general cell population.

      We thank the reviewer for raising this point. The single-cell images are shown for clarity and to visualize the subcellular features; however, the conclusions are made based on the quantitative analysis of multiple cells collected from multiple fields of view (Frames). At least 30 such frames per condition having 4-7 cells/ frame has been acquired and analysed for the quantification throughout this manuscript. We would like to highlight that the image acquisition has been done over random fields on a coverslip. In the revised manuscript, for a better representation of the population of cell-forming invadopodia, a graph showing the percentage of cells forming invadopodia have been added (Fig S1G). Line no: 118-119. The percentage of cells forming invadopodia increased upon HGF stimulation in MDAMB-231.

      (5) Use of a single siRNA/shRNA per target. As noted earlier, using only one siRNA or shRNA carries the risk of off-target effects. For every experiment involving gene silencing (MET, RAB4, RAB14, RCP, MT1-MMP), at least two independent siRNAs/shRNAs should be used to validate the phenotype.

      We would like to clarify that we are using SMARTPool siRNA, which contains 4 individual siRNAs for the target gene. Literature suggests that using a pool of siRNA has reduced off-target effects compared to using single oligos for gene silencing (PMID: 14681580, 33584737, 24875475).

      While SMARTpool siRNA minimizes the off-target effect, it does not eliminate the possibility of it. To confirm that the observed phenotypes are specifically attributable to the genes investigated in this study, we now performed functional experiments using two independent siRNAs targeting RCP and RAB14. The results have been added to figure S4K-L. Silencing of RCP or RAB14 using single oligos resulted in decrease in the degradation index comparable to SMARTpool siRNA, thus phenocopied the SMARTpool siRNA. Line no: 392-400.

      Further, RAB4 is well established to be associated with MET trafficking and it served as a positive control in our study (PMID: 21664574, 30537020). Additionally, a recent study by Hey et al. have used individual oligos for KIF16B to demonstrate the effect of KIF16B silencing on gelatin degradation, which corroborate with our observation from the KIF16B silencing using the SMARTpool siRNA (PMID: 37696580).

      For MET, we used siRNA, shRNA and an inhibitor to show the effect of MET inhibition/perturbation in the invadopodia-associated activity, which validates the observations of siRNA-mediated gene silencing (detailed in point 2).

      We did not perform any experiments using single oligos targeting MT1-MMP, since in our MT1-MMP siRNA-based study, now we have taken an appropriate positive control to validate the efficacy of MT1-MMP silencing (Fig. S1I). In addition, we have shown the effect of MT1-MMP depletion on invadopodia formation using a CRISPR-based gene knock-out study, and another study from our group has shown a similar effect using siRNA (PMID: 31820782), which supports our MT1-MMP KO cell observation.

      (6) Insufficient controls for antibody specificity. The specificity of MET, p-MET, and MT1-MMP staining should be demonstrated in cells with effective gene silencing. This is an essential control for immunofluorescence assays.

      The anti-MET antibody (CST, D1C2 XP) has been used in several studied (PMID: 41166312, 41152910, 39748059). The CST L6E7 anti-MET antibody has been used in studied by Radke et al, Wang et el., Kong et al. (PMID: 36435874, 38262412, 32214092). In our study, immunoblots demonstrating depletion of MET in the siRNA or shRNA-treated cells has been provided in Fig. S1I, K respectively. Further, we have demonstrated MET silencing using immunofluorescence. We also have now added the entire field of view in Fig S1L of showing cells treated with control or shRNA against MET. In the shRNA-treated condition, the cell at the centre shows low MET fluorescence intensity indicating depletion of the RTK, while the surrounding cells have MET staining similar to control.

      Tyr 1234/35 are present in the active site of MET kinase domain and upon binding of the ligand promotes their autophosphorylation (PMID: 17667909). Earlier studies have established the specificity of the phosphor-MET antibody by immunoblotting and immunofluorescence using MET inhibitors (PMID: 21973114, 41009793). In our study we have shown that the inhibition of MET kinase activity using PHA665752 abolished the MET phosphorylation at the Tyr 1234/35, as shown in Fig S1J which revalidates the specificity of the antibody.

      Additionally, in a previous study Joffre et al. have shown that an oncogenic mutant form of MET, M1250T is highly phosphorylated at the Tyr 1234/1235 (PMID: 21642981). Using the phospho-MET antibody, we have shown in Fig 3C, S2I the increased Tyr phosphorylation of M1250T MET mutant as reported by Joffre et al.

      The anti-MT1-MMP antibody is also a very well-established antibody reported in multiple studies (PMID:32479595, 31820782, 35762511). In our study, we have shown the specificity of the antibody using immunoblot analysis. Immunoblots showing significant depletion of MT1-MMP protein level following the SMARTpool siRNA and sgRNA-mediated gene silencing has been provided in Fig. S5I, M’, respectively. Further MT1MMP silencing has been also validated by immunofluorescence in the following studies. PMID: 22291036, 21571860, 20505159.

      (7) Inadequate demonstration of MET recycling. MET recycling should be directly demonstrated using the same approaches applied to study MT1-MMP recycling. The current analysis - based solely on vesicles near the plasma membrane - is insufficient to conclude that MET is recycled back to the cell surface.

      We appreciate the reviewer’s suggestion for an alternative approach to show MET trafficking. We have demonstrated MET trafficking using surface biotinylation, where we have shown that the RCP and KIF16B depletion affect the surface delivery of MET (Fig 5J). Line no: 382-391.

      In addition, to study the surface delivery of cargo, TIRF is a widely used reliable approach and it is highly sensitive technique for detection of surface delivery events (PMID: 24344185, 20971701). We have also tried to investigate the trafficking of MET using antibody uptake approach; however, it could not be established as the binding of the antibody hindered the ligand binding and vice versa.

      (8) Insufficient evidence for MET-MT1-MMP interaction. The interaction between MET and MT1-MMP should be validated by immunoprecipitation of endogenous proteins, particularly since both are endogenously expressed in the studied cell lines.

      We thank the reviewer for pointing out the insufficient evidence for MET-MT1-MMP interaction at the endogenous level. We now carried out the immunoprecipitation of endogenous MET to validate the interaction with MT1-MMP (Fig S5H). A light (low intensity) band corresponding to MT1-MMP was detected in the anti-MT1-MMP immunoblot for the immunoprecipitated sample. We believe that the interaction between MT1-MMP and MET may be weak in nature, resulting in limited co-immunoprecipitation of the endogenous MT1-MMP by MET. The immunoblot is now added to the revised manuscript. Line no: 460-461.

      (9) Inconsistent use of cell lines and lack of justification. The authors use two TNBC cell lines: MDA-MB-231 and BT-549, without providing a rationale for this choice. Some assays are performed in MDA-MB-231 and shown in the main figures, whereas others use BT-549, creating unnecessary inconsistency. A clearer, more coherent strategy is needed (e.g., present all main findings in MDA-MB-231 and confirm key results in BT549 in supplementary figures).

      MDA-MB-231 and BT-549 are two well-characterized TNBC cell lines that readily form invadopodia. These cell lines have been extensively used to study invadopodia-associated breast cancer cell invasion (PMID: 32697977, 35915226, 31533971). These two cell lines also show overexpression of MET, making them suitable model cell lines for our study (PMID: 36139568, 20687930, 27502396).

      Overall, most of the conclusions reported in this manuscript are derived from multiple experimental approaches using two TNBC cell lines for generalization.

      We agree with the reviewer that showing the results from one type of cell line in the main figure would have been better, and wherever possible, we now provided the observations from a single cell line in the main figures and the data from the other cell line in the supplementary figures. However, some of the overexpression studies were performed in BT-549 cells to derive robust statistically meaningful conclusions. Therefore, we could not avoid adding results from both the cell lines in some of the figures. We would like to add that the legends for these figures have been edited to clearly mention the cell lines associated with each of the figure panels to avoid any confusion or inconsistency.

      (10) Inconsistency in invadopodia numbers under identical conditions. The number of invadopodia formed in Figure 1E is markedly lower than in Figure 1C, despite identical conditions. The authors should explain this discrepancy.

      We sincerely thank the reviewer for pointing out the inconsistency in invadopodia numbers across 2 experiments. Fig. 1C has 2 conditions: UT and the HGF-treated condition. The Untreated condition has the serum-free media without any stimulation. Whereas we have added vehicle (DMSO) in Fig. 1E, E’, since the drug is resuspended in DMSO. This difference in the treatment is likely to be responsible for the decreased numbers of invadopodia in Fig. 1E. In different studies it has been shown that DMSO is not biologically inert and can affect invasive properties of cells by perturbing actin dynamics and metalloprotease activity (PMID: 33552397, 22529897, 7188610).

      (11) Questionable colocalization in some images. In some figures - for example, Figure 2G - the dots indicated by arrows do not convincingly show colocalization. The authors should clarify or reanalyze these data.

      As suggested by the reviewer, we have now re-analyzed the data for figure 2G. The apparent visual lack of colocalization is likely due to the relatively lower fluorescence intensity of MET at these structures. We have now added the line intensity plots for the indicated puncta to show the intensity of both channels at the ‘dots’ in the figure 2G’ and they show correlation in their intensity distribution.

      We would also like to elaborate that to quantify the colocalization of two channels, we have used the automated image analysis software Motiontracking (motiontracking.mpi-cbg.de) (PMID: 16143105), which has been detailed in the method section. Briefly, the algorithm works on object-based co-localization. If the fluorescence intensity distributions at a given object corresponding to any two different channels (fluorophores) show 35% or more overlap (Author response image 2), the object is considered as a multi-colour object and the overlapped area value is used to calculate the degree of co-localization. The calculation is carried out over all the objects in a given field of view (frame) and over all the field of views (frames) acquired for a given condition. Also, the apparent colocalization is corrected for random colocalization, which is the random permutation of object colocalization. This makes object-based colocalization more reliable than intensity-based colocalization.

      Author response image 2.

      Image showing the object identification and contour of the multicolour object identified by Motiontracking. The plot shows the intensity distribution of these two objects as analyzed by Motiontracking.

      (12) Abstract, Introduction, and Discussion require substantial rewriting.

      (a) The abstract should be accessible to a broader audience and should avoid using abbreviations and protein names without context.

      (b) The introduction should better describe the cellular processes and proteins investigated in this study.

      (c) The discussion currently reads more like an extended summary of results. It lacks deeper interpretation, comparison with existing literature, and consideration of the broader implications of the findings.

      We thank the reviewer for this suggestion. We have substantially modified the abstract, and the introduction following the reviewer’s suggestion. The introduction has been edited to describe the cellular processes investigated in this study and some of the key associated molecular machineries. In the discussion section, we have avoided redundant descriptions of the results but retained some of them wherever necessary for interpretation and relevant discussion in the light of existing literature.

      Reviewer #2 (Recommendations for the authors):

      (1) Quality of charts. Several charts (e.g., Figure 1B, 1C, 1F) are of poor visual quality. The authors should provide higher-resolution graphs with clearer axis labels, consistent formatting, and properly scaled data.

      We thank the reviewer for pointing out the insufficient visual quality of some of the charts. We believe the resolution of the charts/graphs have changed while converting to PDF, due to image compression. We will provide the uncompressed charts with much improved visual quality, provided they are not restricted by file size limitation.

      (2) Full protein names on first mention. Whenever a protein appears for the first time in the manuscript, its full name should be provided, if possible, before using the abbreviation.

      We have incorporated the full name of the protein while reporting for the first time in the manuscript.

      (3) Correct use of "invadopodium" vs. "invadopodia." Invadopodia is the plural form; the singular is invadopodium. The sentence "Invadopodia, an actin-rich membrane protrusion decorated with proteases, is a tool for ECM and basement membrane degradation during cancer cell invasion" should be corrected accordingly.

      We are thankful to the reviewer for pointing out the grammatical error. We corrected the error in the revised version. Line no: 44.

      (4) Unclear sentence about resistance and invasion.

      The sentence "However, often patients develop resistance to EGFR-targeted therapies due to overexpression of MET; yet, the mechanistic understanding of MET-dependent cancer invasion is unclear" is confusing because the shift from drug resistance to invasion is abrupt. The authors should revise this sentence for clarity and logical flow.

      We are thankful to the reviewer for highlighting this sentence. We have rewritten the sentence as follows “Since one of the receptor tyrosine kinases (RTK), EGFR is often amplified in TNBC patients, they are usually targeted for its treatment. However, often patients develop resistance to EGFR-targeted therapies due to overexpression of another RTK MET”. Line no: 35-38.

      (5) Incorrect figure reference. In the paragraph describing the results related to RAB proteins, there is an incorrect reference to Figure 3 instead of Figure 4. This should be corrected.

      We sincerely apologize to the reviewer for the incorrect figure reference. We have corrected the reference to figures in the revised manuscript. Line no: 251, 261.

      (6) Ambiguous sentence regarding MET activation. The sentence "MET, upon activation by HGF, triggers the activation of the RTK that induces cancer cell invasion" is unclear and should be rewritten for precision and clarity.

      We are thankful to the reviewer for highlighting the unintentional mistake. We have now added a clearer sentence “MET-HGF signalling axis are reported to promotes invasion in gastric cancer cells and melanoma cells”. Line no: 96-97.

      (7) Questionable wording of figure legend. The phrase "Immunoblotting of indicated cell lines with MET and Tubulin" is an informal shortcut. The authors should rephrase it.

      We are thankful to the reviewer for highlighting this sentence. We have modified the figure legend with appropriate text. The modified text is as follows: Lysates of MDA-MB-231, BT-549 and MCF10A DCIS were separated by SDS-PAGE and analyzed by Western blot. Membranes were probed with anti-MET and anti-Vinculin antibody.

      (8) Unnecessary capitalization. Terms such as invadopodia and cortactin should not be capitalized. The authors should correct capitalization throughout the manuscript.

      We are thankful to the reviewer for raising this point. We have modified this accordingly in the revised version. Line no: 142, 159, 280, 436.

    1. Author response:

      We thank the editors for sending our work for review and the reviewers for their thorough and constructive evaluations. In response to their feedback, we will submit a revised version of the manuscript soon. Below, we provide clarification on several of the concerns raised and outline the changes planned for the revised manuscript.

      Reviewer 1, weaknesses

      (1) The two-neuron circuit model is a good choice for the analytics, but it may have hidden a covariance effect of the "rate-dominated" symmetric spike-based kernel that would appear when several inhibitory neurons, each sharing a different spike correlation with the postsynaptic neuron, converge onto it. The rate homeostasis achieved by the rate-dominated model arises from adjusting inhibitory weights according to their initial correlation with the output neuron, so that after learning, the weights are distributed such that these correlations are cancelled out (Vogels et al., 2011). In other words, even the rate-dominated rule is covariance-driven under the hood: with a single inhibitory input, the two-neuron circuit cannot expose this, but with several differently correlated inputs, the covariance dependence should reappear.

      Reviewer 1  highlights that the rate-homeostatic rule by Vogels et al. (2011) also includes a covariance-dependent component. Thus, in a network with multiple excitatory units, differences in pre-postsynaptic correlations can drive a redistribution of weights, resulting in stronger inhibitory weights for higher correlations. Our two-neuron circuit, which contains only one plastic inhibitory input, cannot reveal this competitive effect. We note, however, that although in the Vogels rule the covariance-dependent term is non-zero, it is typically much smaller than the rate-dependent term (see Methods, Section 4.3). Consequently, covariance-dependent organization may emerge on a slower timescale (a similar effect was also shown by Lagzi and Fairhall 2024; https://doi.org/10.1126/sciadv.adi4350). In the revised manuscript, we will clarify this point and analyze small motifs with multiple excitatory units and heterogeneous correlations, focusing on both the timescale of weight redistribution and the resulting steady-state weights. We will also revise our interpretation of Supplementary Figure S9: within the simulated time window, the rate-dominated rule produces uniform inhibitory connectivity, but this does not exclude slower covariance-dependent reorganization.

      (2) It is unclear whether the distribution of inhibitory weights has stabilised after 25 minutes of simulation time (Figure 3C), given that a considerable proportion of (mutual) weights reach the maximum allowed weight while (unidirectional) weights appear to vanish. Without a maximum-weight bound, and given sufficiently long simulations, the weights might diverge to infinity or decay to zero, so the apparent stationarity may be imposed by the bound rather than reflecting a true steady state. This could also be a finite-size effect, given the small number of excitatory connections per neuron.

      The reviewer raises the possibility that the apparent stationarity in Figure 3C is influenced by the imposed weight bounds. In the revised manuscript, we will discuss this point and present extended versions of the simulations in Figure 3 where we remove the weight bounds. We will also test whether the observed behavior depends on network size or excitatory connection density. We note that, under different external input regimes, the weights stabilize without reaching the hard bound (Figure S11), suggesting that saturation is not a necessary outcome.

      (3) The connections from excitatory neurons to the two inhibitory populations are different in the ring model (exc to PV is wider than exc to SST according to Table 3), and it is not clear whether this width difference, rather than the plasticity rules themselves, is responsible for the emergence of the Mexican hat.

      We recognize that we did not fully justify the parametrization used in the ring model. The pre-existing ring architecture determines the spatial correlations available to iSTDP and therefore contributes to the learned connectivity. In the revised manuscript, we will add control simulations in which the excitatory inputs to the PV and SST populations have either identical or markedly different widths, while the plasticity rules are kept fixed. These controls will allow us to assess the relative contributions of these factors to the formation of a Mexican-hat effective-connectivity profile.

      Reviewer 2, weaknesses

      (1) The main limitations concern the extent to which the learned motifs are fully self-organized and how broadly the results generalize. In particular, the ring-network results rely on a pre-specified ring-like excitatory architecture and on two inhibitory populations with distinct plasticity rules, making it important to clarify which aspects of the Mexican-hat effective connectivity emerge from iSTDP itself. The conclusions would also be strengthened by intermediate plasticity rules. Finally, the ring-network simulations provide an interpretable proof of principle, but the authors should clarify whether the PV/SST effects depend on this specific architecture or would also arise in a more generic recurrent or cortex-like connectivity motif.

      This comment raises an important distinction between the components that are specified and those that emerge through plasticity. In our simulations, the excitatory architecture is fixed, whereas the initially weak and unstructured inhibitory-to-excitatory connections are learned through iSTDP. Thus, in the ring network, the spatial organization of excitation is prescribed, but the inhibitory connectivity and resulting Mexican hat-like effective connectivity emerge from the interaction of this architecture with the two iSTDP rules. Importantly, the central result that symmetric and antisymmetric rules promote distinct reciprocal and lateral E/I motifs is not restricted to the ring network, but is also observed in sparse randomly connected spiking networks and in networks with intrinsically generated irregular activity. The ring network is therefore used to demonstrate how these general motif-forming mechanisms can support specific circuit computations.

      Our framework parametrizes a broader family of pairwise iSTDP rules rather than relying exclusively on isolated, preselected rules, allowing the contribution of rule shape and rate-dependent terms to be understood analytically. Intermediate or “mixed” plasticity rules, including those considered by Yang and Doiron (2026) https://doi.org/10.1103/9nv2-y63v, as well as additional network architectures, are valuable directions for extending the framework. We will revise the Discussion to clarify which components are prescribed, which emerge through plasticity, and which conclusions apply beyond the ring-network implementation.

      (2) The authors largely achieve their aim [...] by demonstrating that different temporal forms of iSTDP lead to distinct learned E/I motifs and can shape effective connectivity and cortical-like response patterns. However, the broader biological interpretation remains more suggestive because some results depend on specific assumptions for the network architecture and plasticity rules.

      This comment points to an important distinction between biological generality and mechanistic insight. Our models are deliberately simplified to isolate how the temporal structure of iSTDP interacts with internally generated correlations to select distinct E/I connectivity motifs, and to make this relationship analytically tractable. The resulting predictions are then reproduced in large conductance-based spiking networks with different connectivity structures and sources of irregular activity. Thus, although we do not claim that the specific biological implementations considered here capture the full diversity of cortical circuits, the conclusions are not restricted to a single minimal model or network architecture. Adding further biological detail would introduce additional parameters and architecture-specific assumptions, but would not by itself establish greater generality or provide the same mechanistic understanding. In the revised Discussion, we will clarify the distinction between the general mechanistic principles established by our framework and the more specific biological interpretations that remain to be tested.

      (3) [...] At present, the work identifies rules that are sufficient to generate these motifs in model networks, while the mapping of these rules onto specific interneuron types remains for future experimental testing.

      Our use of the PV and SST labels is intended as a biologically motivated implementation rather than as a universal assignment of plasticity rules to these interneuron classes. The symmetric and antisymmetric kernels were motivated by in vitro measurements from PV and SST interneurons in mouse orbitofrontal cortex, respectively (Lagzi et al., 2021; https://doi.org/10.1101/2021.09.06.459211). Because inhibitory plasticity can vary across brain regions, developmental stages, and experimental conditions (Feldman, 2012; https://doi.org/10.1016/j.neuron.2012.08.001), the general conclusion of our work concerns the mapping from the temporal structure of an iSTDP rule to the E/I motif it promotes, rather than a fixed correspondence between a particular rule and an interneuron identity.

      Our models therefore establish more than the sufficiency of two isolated rules: the analytical framework explains how features of the plasticity kernel and rate-dependent terms determine whether reciprocal, lateral, or blanket inhibitory connectivity emerges. The specific association of these mechanisms with PV and SST interneurons in other circuits remains an experimentally testable prediction. We will clarify this distinction in the revised manuscript and emphasize that, once cell-type-specific plasticity rules are measured in a given circuit, the framework can predict the E/I connectivity motifs that those rules are expected to promote.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Laura Korobkova and Brian Dias describes an interesting study of the role of GABAergic neurons in the zona incerta (ZI) in incentive motivation for reward.

      The authors report that DREADD inhibition of ZI neurons reduced the effort breakpoint in a progressive ratio task, which measures the intensity of incentive motivation to obtain food rewards. In other tests, chemogenetic inhibition did not alter food consumption or memory.

      Conversely, DREADD excitation of ZI neurons increased incentive motivation in the progressive ratio task, expressed as a higher breakpoint for food rewards.

      Korobkova and Dias report that prior stress exposure to a series of stressors (e.g., forced swim & water submersion, restraint, mild footshock) by itself reduced the breakpoint for food reward under vehicle, though it did not impair the ability to learn an instrumental response. However, DREADD excitation of ZI neurons in previously stressed mice increased the breakpoint to normal levels equivalent to the never-stressed group. This important finding indicates the ability of ZI stimulation to rescue the incentive motivational deficit induced by prior stress.

      In fiber photometry studies using vGAT-CRE mice to specifically identify GABA neurons, Korobkova and Dias report that ZI GABA neurons are excited by sensory signals, including neutral cues. However, after reward conditioning, ZI GABA neurons increase their activation to the CS+ cue that predicts reward, but not to the CS- cue that doesn't. ZI neurons also respond in an instrumental reward task during both lever press and reward delivery. The authors conclude that ZI neurons respond to sensory stimuli, but specifically code the motivational significance of reward-related stimuli.

      In optogenetic studies, the authors find that ZI GABA neuron stimulation during a reward CS+ enhances motivated responding to obtain reward, particularly in females, but not stimulation outside the CS+. This suggests the ZI stimulation in females may specifically enhance the incentive salience of the CS+, namely the cue's ability to trigger an increase in 'wanting' for the reward. However, that effect was not found here in males.

      Altogether, this is a fine contribution to the literature, and the authors deserve congratulations on their study and manuscript.

      Strengths:

      This is a powerful and creative set of studies that clarifies the roles of ZI neurons in sensory processing and especially in incentive motivation for rewards. The use of multiple methods and test situations to triangulate on reward motivation functions gives a well-rounded perspective on ZI function. The discovery of incentive motivation roles for ZI neurons is intriguing and improves understanding of ZI, which traditionally has been a relatively understudied brain structure. The finding that ZI stimulation may rescue stress-induced deficits in motivation is especially notable and may have therapeutic implications.

      Weaknesses:

      Minor: This version of the manuscript focuses the introduction and discussion specifically on ZI GABA neurons. The ZI may be primarily GABAergic, but also contains other neurons, and DREADD studies may have used the hSyn promoter, which would impact all types of ZI neurons. Other studies here did more specifically target GABA neurons using vGAT Cre mice and specific targeting. The manuscript might be slightly improved by distinguishing in the discussion a bit more clearly which effects implicate GABA neurons specifically, and which effects might include other neurons too, to more clearly parse out the relative roles of GABA vs broader neuronal populations in ZI.

      The current version of our manuscript notes “Our chemogenetic manipulations targeted all GABAergic ZI neurons, and emerging evidence suggests molecular heterogeneity within this population that may map onto distinct functional roles (Arena et al., 2024; Wilt et al., 2025). Future work to first profile the molecular heterogeneity of GABAergic cells in the ZI and then using intersectional strategies to target defined ZI sub-populations will be essential to providing a more nuanced view of ZI GABAergic influences on motivation.” Our revision will make sure to discuss newer literature demonstrating cellular heterogeneity of ZI and emphasize that this heterogeneity will provide more nuanced contributions of ZI cells beyond the studied GABAergic population on motivation.

      Reviewer #2 (Public review):

      Summary:

      This paper describes a study that uses a combination of observational and experimental techniques to investigate the hypothesis that the zona incerta is a neural loci where sensory information is integrated to interpret the motivational value of reward-associated cues. They show that manipulation of GABAergic neurons in this region bidirectionally modulates responding during a progressive ratio test, that activating these neurons recovers motivational deficits incurred by chronic stress, and that they fire in response to reward-associated visual or auditory cues. They also showed that activity in these neurons is not necessary for incentive salience of reward-associated cues, because inactivating them did not prevent Pavlovian-instrumental transfer. However, activating them did enhance responding during the presentation of reward-associated cues in females but not in males.

      Strengths:

      The study has a very systematic and elegant approach to assess how this region responds first to intrinsic motivation and then to motivation-enhancing effects of reward-associated cues.

      Weaknesses:

      Males and females are used throughout, but sample sizes are generally too small to make a meaningful interpretation of sex differences (which is not the focus of the study, but is worth bearing in mind). In the last experiment, the lack of discrimination between CS+ and CS- conditions across training for males confounds any interpretation of sex-differences in the outcomes.

      The goal of our study was not to determine sex differences in motivation mediated by the ZI but rather demonstrate a general role of GABAergic cells in the ZI on motivation. As such, while all our experiments used male and female mice and our statistical analyses did not uncover any sex differences in most experiments, our sample sizes of each sex are too small to definitively make any statements about sex differences. As noted already in our Discussion, the sex-specific cue-utilization behavioral strategy in our optogenetic experiment warrants further investigation as a contributing factor to motivation that may or may not be influenced by GABAergic cells in the ZI. It bears mentioning that, to our knowledge, none of the recently published literature on the role of the zona incerta in learning, memory and appetitive behavior that is cited in this manuscript (including our own prior work) has uncovered sex differences in the contributions of the zona incerta to these behaviors.

      The ZI is known to be a region where there is notable convergence of neural inputs from a diverse and heterogenous range of sensory and other cortical inputs. To my knowledge, this is the first study that has directly tested whether it may serve to encode motivational/incentive properties of reward-associated cues. The outcomes are not definitive - it appears that they are sufficient but not necessary. However, this study represents an important first step - the ZI also has notable heterogeneity in the genetic identity of neurons, and properly dissecting the function of ZI microcircuits will likely require characterising function based on more than one molecular marker. This is addressed by the authors in the discussion.

      In summary, this study will have a significant impact on our understanding of how motivation is calculated based on complex environmental signals.

      Reviewer #3 (Public review):

      Summary:

      The authors investigated the role of the zona incerta in motivation and cue-reward associations. Using chemogenetic and optogenetic manipulations of the ZI, they altered motivation in cued and uncued variants of the progressive ratio task and rescued deficits in motivation induced by chronic stress. They further use fiber photometry to demonstrate that the ZI tracks the formation of cue-reward associations.

      Strengths:

      (1) The authors fill an important gap in the literature linking sensory input to motivation via the zona incerta.

      (2) The authors demonstrate that ZI tracks cue value rather than just tracking sensory input.

      (3) The authors demonstrate that the ZI excitation rescues stress-induced suppression of motivation.

      (4) The authors perform several important control tasks, demonstrating that their findings are not a result of alterations in locomotor activity, food consumption, or memory.

      Weaknesses:

      In Figure 1D and E (inhibitory vs excitatory DREADDS), the control groups in the Gi group appear to have more elevated breakpoints than the control groups in the Gq group, although a statistical comparison between the two is not reported. It is not clear if this is because the two groups were given a different reinforcement schedule, this should be made clearer.

      In our revision, we will be sure to insert language re-emphasizing that the Gi and Gq experiments were performed using different reinforcement schedules, FR3 and FR1, respectively.

      In Figure 1E, it is important to note that although the authors found a significant planned comparison between Gq VEH and Gq CNO, the interaction was not significant, nor were comparisons to mice injected with control virus. Thus, activation of ZI GABA neurons appears to be a relatively weak effect.

      In our revision, we will insert language to acknowledge that reducing motivation after inhibiting GABAergic ZI cell activity is stronger than increasing motivation seen after stimulating the activity of these cells.

      In Figure 5, the authors see what is likely a significant difference in lever presses during acclimation between the Gi and GFP groups, which they state is an expected difference. However, it is difficult to see why this would be expected. While Gi:CNO manipulation yielded lower breakpoints in Figure 1D, it did not yield lower FR1 responding for food in Fig S3 (although this was FR1 for food dispenser visits rather than lever press). One reason I ask is that the authors highlight the differences in CS+/CS- between groups, but the biggest difference between groups appears to be in acclimation, which may be driving the group x block interaction.

      In a revision, we will revise the language to state that the acclimation difference that is more pronounced in the GFP control group vs the Gi group is to be expected because we had already shown that inhibition of GABAergic cells in the ZI would reduce lever pressing. We will also report a planned comparison of CS+ versus CS− responding that omits the acclimation block, which shows discrimination between sound (CS+) and light (CS−) in the Gi group but not in the GFP group, confirming that the cue effect is not driven by the acclimation difference. We will also note that acclimation responding is non-reinforced and effortful, whereas FR1 dispenser visits (Fig. S3) are reinforced and low-effort, which is why the two measures dissociate.

      In Figure 6, the authors demonstrate that optogenetic stimulation during cue light increases the breakpoint in females, but not in males. They suggest that this may be because the males did not sufficiently discriminate the cue light before optogenetic manipulation began. If this were the case, then the authors would need to use "cue discrimination" as a factor to determine if it is a better predictor than sex.

      In a revision, we will add an analysis that includes cue discrimination as a factor, to test whether it is a better predictor of the optogenetic effect on breakpoint than sex.

      The authors' work demonstrates that chemogenetic inhibition of GABAergic ZI cells reduces uncued motivation for reward but enhances cued responses under extinction. The authors state that this is a paradoxical finding that suggests that the ZI operates within a redundant motivation network. However, a critical difference between the two tasks is that one measures motivation for food while the other measures persistent responding under food extinction, which are not the same process. Thus, a simpler explanation is that ZI inhibition reduces motivation and impairs extinction.

      Our revision will include discussion of this important point and tie it into our previous work (Venkataraman et al. 2019 and 2021 – cited in this version) that included extinction-like protocols, albeit in classical (not operant) conditioning protocols.

  2. Jul 2026
    1. Author response:

      General Statements

      Please find below a revision plan for our manuscript entitled ‘Engulfment by brain macrophages in a short-lived vertebrate’ that was reviewed by Review Commons and that we wish to be considered for publication in eLife.

      We were very happy to see that all three Reviewers were interested in our study and we are sincerely grateful to all of them for their constructive suggestions, which we believe can largely be addressed and will improve our manuscript.

      We would like to inquire whether you would consider publishing our manuscript at eLife in its current form, along with the reviews. We will then provide an updated version of the manuscript, based on our revision plan below when we are able.

      Thank you so much for your consideration.

      Description of the planned revisions

      Reviewer #1:

      Major comments

      Comment 1: Although the authors describe and analyze their data from the viewpoint of engulfing macrophages, the paper would benefit from a broader perspective and a comparison to other studies on microglia in different species. Along this line, the title does not really seem to cover the data presented here very well, and the introduction lacks a proper explanation of terminology on microglia/brain macrophages and their known roles, cell types versus cell states and the current state of the art in fish versus other model species in the context of aging.

      We thank the Reviewer for these suggestions. We will expand the introduction to add more information and references on microglia/brain macrophages across fish and other species. We will also edit the title to more closely reflect the data in the manuscript.

      Comment 2: The result that nearly all myeloid cells in the killifish brain are of the engulfing macrophage type is somewhat surprising. This appears to differ from other studies in for instance zebrafish (e.g. ref 80, that describes the heterogeneity of the myeloid cells in detail). There are two questions we like to raise: (Q1) What is the evidence towards this homogeneity? and (Q2) Could there be a technical bias?

      We agree with the Reviewer, and we were also surprised to find a fairly homogenous myeloid population even in wildtype (non-transgenic) killifish brains, especially given that all these 3 different populations of myeloid cells (microglia, non-microglial macrophages, and dendritic-like cells), could be found in the adult zebrafish brain by Rovira et al. using single-cell RNA sequencing of cd45+ cells (ref. 80).

      We will better describe our evidence towards this homogeneity among killifish brain myeloid cells in the manuscript by adding a paragraph in the text at the end of the section referring to Figure 2. We will also provide additional experiments and analyses (see also below, our detailed responses to Q1 and Q2):

      i) In our and Ayana et al.’s single-cell RNA sequencing datasets, we could not find distinct populations of myeloid cells corresponding to microglia, non-microglial macrophages, and dendritic-like cells, even in wildtype (non-transgenic) killifish young adult and old brain, and even when subclustering only the myeloid cells. However, cd45+ cells were not specifically enriched in these experiments.

      ii) When using PCA, we also could not separate killifish brain macrophages into microglia, border-associated macrophages, and monocyte-derived macrophages (whereas we could separate mouse brain macrophages into these three groups as a positive control for this approach).

      iii) As noted by the Reviewer below, our analysis of the Ayana et al. wildtype dataset was limited to the young adult timepoint from this dataset. In the revised version of our manuscript, we will also include analysis of the old timepoint from this dataset.

      iv) As described in our responses to Q1 and Q2 below, we will provide more detailed subclustering that highlights the heterogeneity we do observe among killifish brain myeloid cells. We find this heterogeneity to be subtle and likely corresponding to different cell states rather than functionally distinct cell types. We will also use hybridization chain reaction (HCR) to test whether cd74, a key marker we use to propose a monocyte-derived macrophage-like identity for almost all detected killifish brain macrophages, labels apoeb+ cells in young wildtype brains in situ (as would be predicted based on our single-cell RNA sequencing analysis). We will additionally use HCR to test whether batf3, a key marker used by Rovira et al. to identify dendritic-like cells, labels any cells in the killifish brain (based on our single-cell RNA sequencing analysis, we do not expect to find batf3+ cells in the killifish brain).

      We agree that there could be technical confounds that could lead to the observed homogeneity of the myeloid population and lack of canonical microglia and dendritic-like cells even in wildtype killifish in our analysis. These could include our/Ayana et al.’s protocol for brain dissociation, our FACS-sorting step, or the step of loading cells into the 10x Genomics chip for single-cell RNA sequencing – each of these could be a step where canonical microglia and/or dendritic-like cells could be lost. In addition, we did not enrich for cd45+ cells, so it is possible that rare populations of immune cells in the killifish brain may not be represented in our dataset. As described below in our responses to Q1 and Q2, we will perform additional analyses that we hope will help address some of these key points, and we will more extensively discuss all possible technical confounds leading to potential cell type bias in the Discussion section.

      In addition to discussing potential technical confounds, we will also better highlight biological possibilities that could explain the differences between our and Rovira et al.’s findings. For example, the observed lack of microglia and dendritic-like cells in killifish brains could reflect a species (killifish vs. zebrafish) difference. Alternatively, even though both studies analyzed young adult timepoints, there could also be differences in biological age between the young adult killifish and zebrafish analyzed, which may be additionally influenced by husbandry conditions (e.g., feeding, pathogen exposure).

      Regarding (Q1): What is the evidence towards this homogeneity? The markers used are overlapping with markers for microglia. It would be helpful to clarify how canonical microglia populations are represented in the dataset. What is the heterogeneity of the oScarletHIGH cells? On several plots (Fig1f, Fig2d, Fig4a) this population of cells seems more heterogeneous than described. Are there different cell states or types? What is the percentage of myeloid cells that is oScarletLOW? To what extent do these cells compare transcriptionally to the oScarletHIGH cells?

      The Reviewer makes a series of excellent points. To address them:

      i) We will more prominently highlight markers not expressed by canonical microglia (e.g., mrc1, cd74) that are expressed by nearly all detectable killifish brain myeloid cells, even in wildtype datasets. We could not find canonical microglia (e.g., tmem119+, sall1+) in any of our killifish datasets. We will highlight these results in the text and will move some of the relevant graphs in the main figure.

      ii) We agree that oScarlet<sup>HIGH</sup> cells are somewhat heterogeneous. To better understand the source of this heterogeneity, we will present UMAPs/Seurat FeaturePlots focusing on cell state markers (e.g., cycling [mki67+] and activated [iba1+] cells) as well as specific cell type markers from mammalian/zebrafish literature (canonical microglia, non-microglia macrophages, dendritic-like cells, etc.). We expect this analysis to show distinct cell states, but fairly uniform expression of cell type markers across all oScarlet<sup>HIGH</sup> cells in our datasets.

      iii) The percentage of detected myeloid cells that are oScarlet<sup>LOW</sup> is very low (<1%). However, as noted in the next comment from the Reviewer, our sorting protocol de-enriches for oScarlet<sup>LOW</sup> myeloid cells, so the actual percentage of all myeloid cells that is oScarlet<sup>LOW</sup> may be higher. We will include additional panels to illustrate this point.

      iv) oScarlet<sup>LOW</sup> myeloid cells are largely transcriptionally similar to oScarlet<sup>HIGH</sup> cells (which are almost all myeloid). The main differentially expressed genes in oScarlet<sup>LOW</sup> vs oScarlet<sup>HIGH</sup> myeloid cells are markers of non-myeloid genes (e.g., elavl3, mpz). These could reflect phagocytosis of non-myeloid cells by oScarlet<sup>LOW</sup> myeloid cells but could also reflect ambient RNA contamination, since oScarlet<sup>LOW</sup> myeloid cells (but not oScarlet<sup>HIGH</sup> cells) were sequenced alongside non-myeloid cells. We will include the differential expression analysis and add commentary in the text to discuss these possibilities.

      Fig1f-i depict an enriched oScarletHIGH group alongside oScarletLOW cells. This representation is a bit misleading since it seems to indicate that really all myeloid cells are of the engulfing macrophage type whereas it is the majority, but not all.

      We agree with the Reviewer. We will also include different representations that provide more accurate estimates of the proportion of all myeloid cells that are oScarlet<sup>LOW</sup>/not engulfing macrophages.

      Line 52: The authors describe that the oScarletHIGH cell group is "enriched for signatures characteristic of macrophage functions". This finding is logical, as the isolation procedure of this population of cells was based on the endocytic and phagocytic properties of the cells. This result appears more consistent with a validation of the isolation strategy than with definitive evidence for myeloid cell identity.

      We agree and we will reframe this result as a validation of the isolation strategy.

      Regarding (Q2): Could there be a technical bias? An alternative explanation that may warrant discussion is whether aspects of the experimental pipeline (cell dissociation, FACS, scRNA-seq) could influence myeloid cell states. For instance, it is conceivable that dissociation induces a reactive program that enhances uptake of fluorescent protein, potentially enriching for oScarletHIGH cells. As the authors use a similar experimental setup to prove uptake of dextran and ovalbumin, such a technical artefact may merit consideration. As this would influence the major conclusions of the paper, the authors might want to address this comment with additional experimental controls, such as single-nuclei RNA-seq on control young and aged brains to profile the natural myeloid population when not submitted to a cell dissociation and FACS procedure.

      We agree with the Reviewer that brain dissociation, FACS, and single-cell RNA sequencing could all represent technical biases that could all influence the representation of different myeloid cell types in our dataset. We will clearly note this potential confound in the Results and the Discussion.

      We agree that it would be valuable to test some of these issues experimentally. We will perform non-dissociative HCR (in situ) to test (1) whether in young wildtype brains, apoeb+ cells (macrophages) express cd74 (a key marker for our argument re: MDM-like cell type identity) and (2) whether we can find cells expressing batf3 (a key marker used by Rovira et al. to identify dendritic-like cells, which we cannot detect in the killifish brain in our single-cell RNA sequencing data). We will also acknowledge that these are only individual markers and would not provide as comprehensive a picture of cell type identity as single-nuclei RNA sequencing.

      Comment 3: The authors compare the oScarletHIGH cell transcriptomes to mouse and killifish datasets. Both the mouse (Barr et al) and killifish (Nagvekar, this paper) dataset are from enriched immune cells (mouse= CD45+ cells, and the 3 cell types selected from that). Why did the authors not compare to the whole mouse CD45+ dataset? Including zebrafish (Rovira et al, 2025) here would strengthen the evolutionary comparison. I also feel that the additional comparison with young killifish (Ayana et al) might not be that solid since this dataset was initially not enriched and has a significantly lower number of myeloid cells, and thus much less power. The old age time point in that study contained more myeloid cells, and might be interesting to include for cell type comparison. There are other, perhaps more unbiased ways of comparing cell types across species, for instance SAMap, developed by co-author Bo Wang. Did the authors consider using this or other methods?

      We thank the Reviewer for these excellent suggestions. We will do the following in our revised manuscript:

      i) We will include comparisons to all immune (CD45+) cells from mice (Barr et al.).

      ii) We will include comparisons to all immune (cd45+) cells in zebrafish (Rovira et al.).

      iii) We will also analyze the old age time point from the Ayana et al. wildtype killifish dataset.

      The percentage of oScarletHIGH cells in the aged condition is 8% (Fig1-suppl1) compared to 4% at young age. On the other hand, a lower number of cells was isolated at old age compared to young age (Figure4a). Can the authors elaborate on this difference? Later on, it is stated that the engulfing capacity declines with aging, but could this be linked to the lower or potentially biased recovery of cells?

      We appreciate the Reviewer’s point. We will elaborate on the percentage differences in oScarlet<sup>HIGH</sup> cells recovered in the different experiments, and more clearly indicate what they could originate from and whether they could influence the engulfment results:

      i) We will indicate more clearly that in each single-cell RNA sequencing experiment, we loaded the entire set of Live oScarlet<sup>HIGH</sup> cells collected from each condition onto the 10x Genomics chip. In the young vs. old comparison, more Live old oScarlet<sup>HIGH</sup> cells were loaded than Live young oScarlet<sup>HIGH</sup> cells (66K vs. 42K, based on the FACS sorter counts, though these numbers are likely overestimates). One possibility is that the lower recovery of successfully sequenced old oScarlet<sup>HIGH</sup> cells might be explained by differences with age in oScarlet<sup>HIGH</sup> cells’ ability to remain intact during loading onto the 10x Genomics chip.

      ii) In the ex vivo experiments where we find that engulfment capacity declines with age (current Fig. 5), the proportion of oScarlet<sup>HIGH</sup> cells among all Live cells slightly increased with age. We will include the data illustrating this, and indicate more clearly in the text that the differences in engulfment capacity observed in these experiments are unlikely to be explained solely by lower recovery of oScarlet<sup>HIGH</sup> cells from old brains.

      iii) We will more clearly acknowledge in the Discussion the possibility of potentially biased cell recovery with age.

      Figure4a: Transcriptional differences are stated between young and old (line 226), can a relevant selection be shown in e.g. a dotplot or heatmap?

      This is another great point from the Reviewer. We will show these genes in a dotplot/heatmap in a Supplemental Figure. We will also clearly indicate in the main text that many of the genes most strongly enriched in old oScarlet<sup>HIGH</sup> cells in our young vs. old comparison experiment were not strongly expressed in an independent old oScarlet<sup>HIGH</sup> cells dataset (which was compared to old wildtype cells, without a young counterpart). 

      The UMAP clustering does seem to indicate batch effects on panels a and g. Can the authors provide subclustering and show that young and old/ FACS sorted high and low cover similar cell types/states? The PCA plot (panel f) and marker analysis is not fully convincing, as PC1 and 2 alone do not suffice to explain all the variance in these cells, and the markers are common ones for many microglia/macrophage cell types (and thus likely to be expressed similarly).

      The Reviewer has another excellent suggestion. We will provide subclustering, which indeed shows that young and old oScarlet<sup>HIGH</sup> cells generally represent similar cell types and states. This analysis also shows a subcluster specific to old oScarlet<sup>HIGH</sup> cells, but this subcluster is far less pronounced in a second old oScarlet<sup>HIGH</sup> dataset (see our previous comment), so we will clearly indicate in the main text that this subcluster may not be robust. We will also remove the PCA plot.

      Figure 5: It would be informative to include the corresponding aged condition for panels c and e.

      We agree and we will include this.

      Minor comments

      The authors use the oScarlet fish in a heterozygous state. Is the homozygous line not viable? Or what is the reason to use the heterozygous state? Too high expression of OScarlett? Is apoptosis of neurons checked for this line? Compared to wild type?

      The homozygous SP-oScarlet line has not been characterized, and we will include this information in the Results. We did not check apoptosis of neurons in the SP-oScarlet line or in wildtype killifish, and we will indicate this in the Methods and Discussion.

      Is the secretion of OScarlett completely proven? Maybe the signal is engulfed by the clearance of cell debris from apoptotic neurons?

      This is a great point. We have not directly tested the secretion of oScarlet in the SP-oScarlet line, and it remains possible that engulfed oScarlet also comes from apoptotic neurons that are being engulfed by myeloid cells. We will acknowledge this possibility in the Discussion.

      We will also more clearly highlight references that adding a signal peptide to proteins (as we did for the SP-oScarlet line) is sufficient to induce their secretion in several contexts in different species (although this does not prove secretion in this particular case).

      We also do have another line (built for a separate project) in which oScarlet is expressed cytoplasmically in elavl3+ cells. In this line, in pilot data, we did not observe oScarlet<sup>HIGH</sup> cells by flow cytometry. While these experiments are not of publishable quality, these observations also suggest that the presence of a signal peptide on oScarlet is necessary for the engulfment phenotype.

      Line 329. What exact difference between teleosts and mice are you pointing at?

      We will edit the text to clarify that we are referring to our inability to find a cell population with a transcriptional profile like that of canonical adult mouse microglia in the adult killifish brain.

      Suggestions for Figures

      General remark IHC/HCR figures: Please add overview figures to guide the reader to ROI shown.

      Fig1.

      1.b Please add overview figures to show the overall distribution of the signal in the brain. Are there any hotspots or low-abundant regions?

      1.d/1.e Please add numbers of cells on figure panels.

      1.d: overview figure is unclear. Telencephalon seems to be missing? Can annotation be added to regions of the structure?

      Additional in vivo evidence of the homogeny/heterogeneity of oScarlet protein+/RNA- cells would be informative, e.g. double labeling (HCR) with some of the markers of panel Fig. 2a). This also would exclude location bias. One could expect myeloid cells that are not in close proximity to secreting neurons, for instance in dense neurogenic niches

      Fig. 2

      2.c Please add number of cells on graph

      2.d Please add brain regions to graphs of published datasets like was done for the own dataset.

      Subtext: Why is there a different number of cells for Barr in c and d? Please add number of cells of own killifish dataset in legend.

      Supplement fig. 2 Please add additional microglia-specific markers such as for instance HEXB, GPR34, SELL1, C1Q.

      Fig.3

      Please add overview pictures.

      Fig.4

      4.g grey color not very visible

      It would be informative to have the markers of 4.h plot on top of the UMAP to show the distribution/differences, maybe in supplementary information.

      We will implement all of these changes. We will add cd74 HCR labeling of oScarlet<sup>HIGH</sup> cells that do not express oScarlet RNA transcripts in situ. The different numbers of cells from the Barr et al. dataset in the current panels 2c and 2d do not reflect differences in brain regions from which macrophages were isolated (in both cases, the same whole-brain dataset was used). Instead, in this dataset, there is a small population of interferon-responsive macrophages that are not microglia, BAMs, or MDMs; these cells were included in the analysis in 2d but not 2c. We will clarify this and update the analysis in current panel 2d to include all immune cells from the Barr et all. dataset, as suggested by the Reviewer above.

      Reviewer #2:

      Red fluorescent proteins are notorious for being prone to aggregation. Are oScarlet proteins being internalized by macrophages aggregates or soluble proteins? This distinction is important as the clearance of extracellular molecules could be mediated by most cells yet aggregates could be removed specifically by macrophages. Can experiments be conducted to distinguish between these two possibilities? We realize this may be challenging. If not feasible, the discussion should be tempered to reflect this possibility.

      The Reviewer makes a great point, and we were also interested in this question. mScarlet and its derivatives (including oScarlet) would generally be expected to remain in the monomeric state to a greater extent than other red fluorescent proteins like mCherry (Bindels et al., Albakri et al.). We will indicate this with references in the Results section.

      But this does not preclude the possibility of oScarlet aggregation in SP-oScarlet brains. We did collect soluble and insoluble fractions from SP-oScarlet brains using a gentle lysis method that would be expected to preserve aggregates in the insoluble fraction (Avar et al.). However, we found that the levels of oScarlet in both fractions were below the limit of detection by western blot and by a plate reader-based fluorescence assay, thereby precluding us from directly testing how much oScarlet was aggregated. We will address the possibility of oScarlet aggregation, and its potential impact on engulfing macrophages, in the Results and Discussion sections.

      Brain dissociation tends to generate a lot of debris, especially from sheared neurons. Therefore, the high level of oScarlet inside macrophages could be an artifact of dissociation rather than a reflection of in vivo clearance. Authors should use internalization inhibitors during dissociation (CytoD, Dynasore, and pitstop) to exclude this possibility. Alternatively, if they have a transgenic killifish that expresses another fluorescent reporter in neurons (and preferably at a similar level to that of oScarlet), authors should dissociate brains together and quantify how many oScarlet+ cells are now also positive for that other fluorescent reporter. This could give an idea of how much engulfment is occurring due to the dissociation processes. It is not ideal, as macrophage eating could be happening during dissociation but before cells are in single cell suspension. However, given that RNAseq is needed to identify macrophages, this reviewer would be satisfied by this alternative approach if the aforementioned pitfall is also presented in the discussion.

      This is another excellent suggestion, and we have already done the following experiments:

      i) We have performed bulk RNA sequencing from FACS-sorted oScarlet<sup>HIGH</sup> cells from middle-aged SP-oScarlet brains dissociated with five inhibitors (cytochalasin D, dynasore, Pitstop 2, bafilomycin A, and wortmannin) in addition to transcription/translation inhibitors. Deconvolution analysis of these bulk RNA sequencing data showed that oScarlet<sup>HIGH</sup> cells in the presence of cytochalasin D, dynasore, Pitstop 2, bafilomycin A, and wortmannin are also almost all macrophages. While these data are not at single-cell resolution, they indicate that engulfment by macrophages is unlikely to solely occur during the dissociation process. We will include a revised Supplementary Figure with these experiments.

      ii) We have also already performed a pilot co-dissociation experiment (in a different genetic background) as suggested by the Reviewer using a ubb:GFP reporter as the second transgenic and observed very few GFP+ oScarlet<sup>HIGH</sup> cells (see Author response image 1). We will include co-dissociation data in the revised version of the manuscript.

      Author response image 1.

      We believe that the results of these experiments are consistent with the notion that oScarlet<sup>HIGH</sup> cells have engulfed oScarlet in vivo, prior to dissociation step. Nevertheless, we will also make note in the Discussion of the pitfall mentioned by the Reviewer.

      Related to the above, it appears based on the scRNAseq that dissociation heavily enriched for brain macrophages. Therefore, the claim that clearance is mostly macrophage mediated could be due to an enrichment of this population during dissociation rather than this cell type being responsible for most of the extracellular waste disposal. Authors should quantify the % of total oScarlet that is specifically in macrophages in the brain sections they already have that are stained against oScarlet and CSF1R/ApoEB transcript.

      The Reviewer’s point is well taken and we agree that experiments in intact brain sections are important to orthogonally test whether macrophages are responsible for engulfment. As suggested by the Reviewer, we will quantify the percentage of total oScarlet that is in apoeb+ macrophages in the brain sections we already have (current Fig. 3a). We note that it may not accurately reflect the actual in vivo percentage, as we have found that it can be challenging to draw cell boundaries around macrophages due to their irregular shapes.

      It is concerning that dextran and oScarlet are almost perfectly colocalized in the image presented (Figure 3a). It raises the possibility, among others, that dextran is sticking to potential oScarlet aggregates and then being internalized by macrophages. Therefore, it could be an artifact of the transgenic line. Authors should repeat the experiment in wildtype fish and use HCR against CSF1R/ApoEB to address this issue.

      We agree with the Reviewer. We have already injected dextran into non-transgenic killifish brains and observed dextran engulfment by apoeb+ cells. We will include this experiment in the revised version of the manuscript.

      Reviewer #3:

      Major comments

      The SP-oScarlet model enriches cells based on phagocytic capacity - by design, any phagocytic cell, including microglia, can be labeled. Only 0.5% of oScarlet<sup>LOW</sup> cells were myeloid cells, confirming that this method captures virtually the entire myeloid population. The transcriptional resemblance to BAMs/MDMs is therefore a post hoc characterization of brain phagocytes broadly, rather than evidence for a selectively labeled subset. The authors show examples of apoeb<sup>+</sup> cells near vasculature (Fig. 3b), but do not provide a comprehensive quantification of the full spatial distribution of oScarlet<sup>HIGH</sup> cells. Importantly, neither the SP-oScarlet macrophages nor previously published wild-type killifish brain macrophages could be transcriptionally separated into three subgroups analogous to mammalian microglia, BAMs, and MDMs by PCA. This suggests that fish brain macrophages may not exist as subpopulations that correspond with their mammalian counterparts. The authors should therefore describe these cells as brain myeloid cells that exhibit BAM/MDM-like transcriptional characteristics, rather than implying they are a population equivalent to mammalian BAMs/MDMs.

      We agree with the Reviewer and will describe killifish brain macrophages as “brain myeloid cells that exhibit BAM/MDM-like transcriptional characteristics.”

      As described in our responses to Reviewer 1’s comments, we will also provide/more prominently highlight analysis of killifish brain myeloid cell datasets showing:

      i) Markers (e.g., mrc1, cd74) that are shared between mammalian BAM/MDMs and killifish brain myeloid cells.

      ii) The heterogeneity we observe among killifish brain myeloid cells, which we believe corresponds to different cell states, but is not likely pronounced enough to correspond to functionally distinct cell types.

      iii) The presence of oScarlet<sup>LOW</sup> killifish brain myeloid cells, which are de-enriched by our sorting strategy.

      Minor comments

      The authors acknowledge that oScarlet is relatively resistant to degradation and could affect the state of cells that engulf it. The independent single-cell experiment comparing old SP-oScarlet and old wild-type macrophages addresses this concern at the transcriptional level. However, it remains possible that chronic accumulation of degradation-resistant protein in lysosomes could itself impair phagocytic capacity. If so, the age-related decline in ex vivo engulfment might be partially attributable to oScarlet burden accumulated over a lifetime, rather than to aging per se. While a direct functional comparison with old wild-type macrophages is experimentally challenging, the authors should at least acknowledge this interpretive caveat in the Discussion.

      The Reviewer makes a great point, and we will acknowledge this caveat in the Discussion.

      The transcriptional data show only modest changes in engulfment- and lysosome-related pathways with age. The authors should offer one or more specific, testable hypotheses (e.g., which phagocytic receptors or lysosomal components might be affected) to guide future investigation.

      We will provide a list of top engulfment-relevant genes that change the most with age, which also show only modest changes. We will better acknowledge that the age-related transcriptional changes in engulfment-relevant genes observed in our single-cell RNA sequencing experiment are modest and are unlikely to directly serve as a resource for testable hypotheses. We will also move our single-cell RNA sequencing data to the Supplementary Figures.

      In the text, gene names such as APOEB are capitalized, whereas the convention for fish gene nomenclature is lowercase italics (e.g., apoeb). The authors should ensure gene formatting follows species-appropriate conventions throughout the manuscript.

      We thank the Reviewer for this suggestion, and we will change killifish gene names to be lower-case italicized throughout the manuscript.

      Description of analyses that authors prefer not to carry out

      Reviewer #1:

      Regarding (Q2): Could there be a technical bias? An alternative explanation that may warrant discussion is whether aspects of the experimental pipeline (cell dissociation, FACS, scRNA-seq) could influence myeloid cell states. For instance, it is conceivable that dissociation induces a reactive program that enhances uptake of fluorescent protein, potentially enriching for oScarletHIGH cells. As the authors use a similar experimental setup to prove uptake of dextran and ovalbumin, such a technical artefact may merit consideration. As this would influence the major conclusions of the paper, the authors might want to address this comment with additional experimental controls, such as single-nuclei RNA-seq on control young and aged brains to profile the natural myeloid population when not submitted to a cell dissociation and FACS procedure.

      We believe that generating our own single-nuclei RNA sequencing dataset would be beyond the scope of this study. However, a preprint by Williams et al. from the Benayoun Lab (https://doi.org/10.64898/2026.04.09.717549) includes single-nuclei RNA sequencing data from wildtype killifish brains and could provide an excellent test of our findings. We will point readers to this preprint in the Discussion.

      Comment 3: The authors compare the oScarletHIGH cell transcriptomes to mouse and killifish datasets. Both the mouse (Barr et al) and killifish (Nagvekar, this paper) dataset are from enriched immune cells (mouse= CD45+ cells, and the 3 cell types selected from that). Why did the authors not compare to the whole mouse CD45+ dataset? Including zebrafish (Rovira et al, 2025) here would strengthen the evolutionary comparison. I also feel that the additional comparison with young killifish (Ayana et al) might not be that solid since this dataset was initially not enriched and has a significantly lower number of myeloid cells, and thus much less power. The old age time point in that study contained more myeloid cells, and might be interesting to include for cell type comparison. There are other, perhaps more unbiased ways of comparing cell types across species, for instance SAMap, developed by co-author Bo Wang. Did the authors consider using this or other methods?

      We thank the Reviewer for the SAMap suggestion. We did perform SAMap with a mouse reference from the Allen Brain Atlas. Our SAMap analysis identifed distinct microglia-, BAM-, and dendritic-like myeloid populations. However, we were not confident in these SAMap results because the differences between the populations were subtle and we thought they would be unlikely to represent differences between functionally distinct cell types. Of note, the dendritic-like cells we identified by SAMap in the killifish brain did not express canonical markers of dendritic/dendritic-like cells from mammalian and zebrafish literature (e.g., batf3). Additionally, other SAMap studies from the Wang Lab analyzing non-neural cell types in other vertebrates do not separate microglia from non-microglial macrophages (Kalakuntla et al., in preparation).

      Comment 4: Regarding the comparison with the aged brain:

      Figure 4: It would be nice to include the same comparisons as for young fish (cfr Fig.1 panels F-I).

      We agree with the Reviewer. Unfortunately, we did not sequence oScarlet<sup>LOW</sup> cells from old brains, for cost reasons, and this precludes the comparison requested by the Reviewer. We will more clearly indicate that we do not have the oScarlet<sup>LOW</sup> cells from old brain in the Results section.

      Reviewer #2:

      The flow cytometry strategy used does not distinguish between oScarlet protein that has been internalized versus that which is sticking to the surface of macrophages. Authors should stain non-premeabilized and permeabilized cell suspensions with a flow antibody against mCherry/RFP to get a sense of how much oScarlet is inside versus outside of the macrophage. For most antibodies this can be done on the same sample sequentially if the antibodies have a different fluorophore.

      The Reviewer makes an important point, and we will address this caveat in the Methods. We have performed a pilot of the flow cytometry experiment suggested by the Reviewer and, encouragingly, observed a higher proportion of cells labeled by the antibody (the same anti-mCherry antibody we used to label oScarlet in situ) in the permeabilized condition. However, we do not wish to publish these results as even in the permeabilized condition, the antibody labeled <0.2% of cells – meaning it likely did not label most oScarlet<sup>HIGH</sup> cells, making it difficult to draw strong conclusions about internalized vs. surface oScarlet in these cells.

      Reviewer #3:

      The age-related decline in oScarlet fluorescence in oScarlet<sup>HIGH</sup> cells in vivo could reflect either reduced phagocytic capacity of macrophages, or reduced oScarlet secretion by neurons, as the authors have discussed (Fig. 5a). The ex vivo assay addresses this by standardizing substrate concentration, which is a strength, but an in vivo functional assessment would provide a more physiologically relevant complement. The authors have already established the methodology for in vivo substrate injection (Fig. 3a, dextran). A similar experiment comparing substrate uptake in young and old fish would circumvent potential artifacts of the ex vivo approach, such as enzymatic dissociation altering surface receptor availability, and would directly test whether engulfment declines in the native brain environment.

      We agree with the Reviewer and performed a pilot of the suggested experiment, which did not show age-related differences in in vivo engulfment of injected dextran. However, in developing the brain injection procedure, we concluded that this procedure works well for comparisons within a sample, but not for comparisons across samples and conditions (e.g., young vs. old). This is because the injection site is determined by sight (aiming for the most medial point on the telencephalon/optic tectum border, with no standardized way to control injection depth), and not stereotactically with coordinates. With our current procedure, it is challenging to perform reproducible injections across ages due to known age-related differences in fish/brain size and skull thickness.

      We are interested in developing stereotactic injection approaches that would be compatible with comparisons across ages and other conditions, but we believe that this is beyond the scope of this manuscript. We will discuss this possibility as a future direction in the manuscript.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      The evidence described for the claim that this technique improves the alignment of the reconstruction of small complexes compared to standard techniques is incomplete. The authors could better evaluate the effects of model bias on the reconstructed densities, as suggested by reviewer #1.

      We thank the editors for highlighting this remaining concern. To better evaluate the effects of model bias, we have performed the FSC-based analyses suggested by Reviewer 1, including a map–model FSC of the omit map and an FSC between the half-maps (though the latter is unreliable for the composite), and added the results to the revised manuscript (new Figure 5—figure supplements 1–3; detailed under Requested FSC Analysis below). We have additionally revised the text to avoid overstating the reconstruction as “unbiased” and to soften comparisons with other reconstruction methods (detailed below).

      Public Reviews:

      Reviewer #1 (Public review):

      In the revised version, the refinement of atomic occupancies in the 2DTM-generated maps has been insightful: densities only come back at values ranging from 0.55–0.80, whereas residues included in the template remain at 1, suggesting that the 2DTM-reconstruction does suffer from model bias. Their newly added Omega calculations, which are helpful, also suggest that model bias is present in the 2DTM-based reconstructions. These observations therefore contradict the first subsection heading of the Results, which claims “unbiased reconstruction of omitted residues”.

      We agree that the previous subsection heading was too absolute. We have changed the heading from “Unbiased reconstruction of omitted densities in a 43 kDa protein kinase” to “Recovery of densities omitted from the template in a 43 kDa protein kinase”. The opening sentence now states that we evaluated the ability of 2DTM to “recover omitted ligand densities” (L123–125).

      We also revised nearby statements to describe the observations without claiming that the entire reconstruction is free of template bias. The manuscript now states that, because the corresponding features were omitted from the search template, the recovered densities cannot result from direct inclusion of those features in the template (L150–153).

      For the omitted alpha-helical turn, we now state simply that its density was recovered despite its absence from the search template (L191–192). We also revised the interpretation of the occupancy-refinement results from “confirming partial, unbiased recovery” to “supporting partial recovery of density in the omitted regions” (L198–199). The same wording has been applied to the Supplementary file 1 caption on page 27.

      Finally, the summary of the ligand-deletion experiments now states that a ligand and nearby residues can be deleted to reduce template bias while retaining sufficient signal for their density to be recovered (L248–251). In the Methods, we retain the description that the composite omit map was constructed to avoid template bias at the omitted locations (L953–954).

      We have also toned down comparative statements about reconstruction accuracy, including the relevant subsection heading (“Comparison of 2DTM and RELION reconstructions from the same particle stack” at L252–253) and the surrounding discussion (L274–279).

      Requested FSC Analysis

      The measurement of how much model bias is present in this OMIT map by FSC calculations is still pending. This could be done in two ways. My original suggestion was to calculate a mapto-model FSC for the OMIT map and the full reference. This should be compared with a similar map-to-model FSC on the map where only the ligand was omitted. Alternatively, they can use the cisTEM FSC uncorr procedure on the OMIT half-reconstructions and compare the resulting curve with the one presented in Figure 1b.

      We have now completed both analyses and added them to the revised manuscript (new Figure 5—figure supplements 1–3, with accompanying Results and Methods text).

      (1) Map–model FSC. We computed the map–model FSC between the composite OMIT map and a density simulated from the full 1ATP model, and compared it with the equivalent FSC for the Figure 1 reconstruction, in which the ligand and residues 222–227 were omitted from the template (Figure 5—figure supplement 1). The composite OMIT map crossed FSC = 0.5 at 3.5 Å and FSC = 0.143 at 3.0 Å, compared with 3.0 Å and 2.4 Å, respectively, for the Figure 1 reconstruction. Because each local region of the composite map was taken from a reconstruction in which the corresponding residues were absent from the template, this agreement reflects genuine recovery rather than direct inclusion of those local features in the template. The lower FSC values relative to the Figure 1 reconstruction are expected. In the Figure 1 reconstruction, most of the protein remained in the template, whereas the composite is assembled from disjoint local omit regions. This stitched construction introduces holes and mask boundaries that affect Fourier-space agreement across the curve, in addition to the weaker, partial recovery of locally omitted density. Thus, the map–model FSC provides a conservative Fourier-space assessment of recovered omit-region density. The composite map–model FSC also shows a negative dip at the lowest spatial frequencies, which does not indicate failed recovery. Radial binning around the omitted atoms (Supplementary file 3) shows that the atom-centred shells (≤2 Å) recover positive but weakened density upon omission, whereas the peripheral shells (2–3 Å) are negative and nearly identical whether the residue is present or omitted, leading to net-negative density within the molecular envelope. We therefore interpret the low-frequency dip as a consequence of the composite construction rather than as evidence for failed recovery of omitted density.

      (2) Half-map FSC<sub>uncor</sub>. We also computed the half-map FSC of the individual OMIT reconstructions (Figure 5—figure supplement 2): the 36 individual omit reconstructions cross FSC = 0.143 at a median resolution of 2.8 Å, comparable to the Figure 1b reconstruction (∼3.0 Å). The composite map’s own half-map FSC appears higher, but this value is not a reliable resolution estimate because the two composite half-maps are assembled using the same voxel-assignment masks. This shared support introduces artificial correlations, which we demonstrate with a phase-randomization control (Figure 5—figure supplement 3). We therefore use the individual omit-reconstruction FSCs and the composite map–model FSC, rather than the composite half-map FSC, to assess Fourier-space agreement.

      Together, these analyses show that the OMIT reconstructions contain high-resolution signal in regions absent from the corresponding search templates, while also identifying the low-frequency dip and the inflated composite half-map FSC as consequences of the conservative stitched composite construction.

      Reviewer #3 (Public review):

      Nor was it compared to more recent strategies for processing SPA data from small molecules, such as Blush regularization or HR-HAIR. [...] This places this method as a complementary technique, and whether it outperforms those methods for a wide variety of molecules is yet to be determined.

      We agree that a systematic comparison with recent small-particle SPA methods such as Blush regularization and HR-HAIR will be important. We have added this point to the Discussion (L831–846). We also note that such as comparison should consider not only particle stack and molecular mass, but also the fidelity of the 2DTM template forward model. In the ideal limit of an accurate template and forward model, 2DTM should provide a strong prior for particle detection and pose determination. In practice, however, current templates remain imperfect approximations to the experimental signal because of inaccurately modelled solvent-boundary effects and atomic scattering factors, bonding and charge redistribution, and conformational mismatch. Improving template generation is therefore an important direction for extending the range of molecular targets and imaging conditions where 2DTM can be applied.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study addresses the mechanism of action of benzoylurea insecticides and explores the metabolic consequences of inhibiting glycogen breakdown in insects. Both reviewers identify major flaws with the premise of the work. The strength of the provided evidence is inadequate as the data do not, or poorly, support several central claims. The significance of the findings is considered marginal.

      The Assessment stated that “both reviewers identify major flaws with the premise of the work” and that “the strength of the provided evidence is inadequate.” We have addressed both dimensions:

      (1) Premise: The Introduction has been substantially restructured to explicitly acknowledge the compelling CRISPR/Cas9 evidence establishing CHS as the primary site of BPU resistance (Reference 1). The study is now reframed as a systematic evaluation of GP as an independent insecticidal target and an investigation of metabolic compensation mechanisms — questions with scientific value independent of the BPU mechanism debate (see details in lines 47-54 of the revised manuscript).

      (2) Evidence: Four new sets of experiments directly address the specific evidence gaps identified by the reviewers: (i) GP enzyme activity measurements in RNAi-treated larvae; (ii) expression analysis of alternative glycogen catabolic enzymes; (iii) molecular docking and MM/GBSA binding free energy analysis; (iv) comprehensive fitness cost assessment including feeding rate, larval weight, pupal weight, adult wing area, and female fecundity.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors investigate whether glycogen phosphorylase is a potential molecular target of benzoylphenylurea insecticides and examine the physiological consequences of inhibiting glycogen breakdown in the diamondback moth Plutella xylostella. The authors express and characterize recombinant glycogen phosphorylase, test its inhibition by a mammalian glycogen phosphorylase inhibitor and by the insecticide diflubenzuron, and assess the physiological effects of glycogen phosphorylase inhibition through chemical exposure and RNA interference. Based on these experiments, the authors conclude that benzoylphenylurea insecticides do not target glycogen phosphorylase and propose that insects compensate for glycogen phosphorylase inhibition through activation of gluconeogenesis, allowing them to maintain glucose homeostasis and complete development despite strong suppression of the enzyme.

      Strengths:

      The study addresses an interesting and long-standing question in insect toxicology regarding the mechanism of action of benzoylphenylurea insecticides. The authors combine several complementary approaches, including recombinant enzyme characterization, inhibitor assays, RNA interference, gene expression analyses, and metabolite measurements. The biochemical characterization of the recombinant glycogen phosphorylase and the demonstration that the tested glycogen phosphorylase inhibitor can strongly inhibit enzyme activity represent important technical strengths. In addition, the study integrates biochemical and physiological observations to explore how insects might compensate for disruptions in central carbohydrate metabolism.

      We are grateful that Reviewer 1 recognized the study's strengths, including the complementary multi-approach strategy, the biochemical characterization of recombinant PxGP, and the integration of biochemical and physiological observations.

      Weaknesses:

      (1) The proposed compensatory mechanism relies on indirect evidence; direct measurements of gluconeogenic flux are lacking.

      We agree that isotopic tracer experiments would provide the most direct evidence for gluconeogenic flux. Such experiments are beyond the scope of the current revision, and we now explicitly acknowledge this as a key limitation and an important direction for future research (revised Discussion: Study limitations and future directions).

      However, we note that the convergent evidence from multiple independent lines collectively supports gluconeogenic activation: (i) transcriptional upregulation of PEPCK and G-6-Pase; (ii) declining protein levels now independently confirmed by our new enzyme activity data showing a 30.78% decrease in total protein concentration at 24 h post-RNAi (new Figure 10A); (iii) altered amino acid profiles; and (iv) maintained trehalose levels. The revised manuscript presents this evidence more cautiously, framing it as “consistent with gluconeogenic compensation” rather than establishing metabolic flux.

      Additionally, we now provide new data on GP enzyme activity (new Figure 10A, B; see response to Recommendation 4 below) and alternative glycogen catabolic enzyme expression (new Figure 10C, D; see response to Recommendation 2 below) that further strengthen the evidence chain.

      (2) Alternative glycogen degradation pathways are proposed but not experimentally examined.

      We have now directly addressed this concern. RT-qPCR analysis of glycogen branching enzyme (GBE) and α-amylase following PxGP knockdown reveals a striking and informative differential response (new Figure 10C, D):

      GBE was significantly upregulated at 24 h (+29.24%, P < 0.05), 48 h (+16.78%, P < 0.05), and 96 h (+44.46%, P < 0.001), indicating transcriptional activation of an alternative glycogen-remodeling enzyme in response to GP suppression.

      α-Amylase showed no significant change at any time point, demonstrating that the compensatory response is pathway-specific rather than a generalized upregulation of all glycogen-degrading enzymes.

      This differential pattern — GBE up, α-Amylase unchanged — provides the first evidence that P. xylostella selectively activates specific glycogen remodeling pathways when GP function is compromised. Upregulation of GBE, which increases glycogen branching and solubility, may facilitate glycogen mobilization through alternative routes even when GP-mediated phosphorolysis is impaired. These data are incorporated as new Figure 10C, D and discussed in the revised Results and Discussion.

      (3) Physiological consequences (fitness costs) are not explored.

      We have now conducted a comprehensive fitness cost assessment (new Figure 11). The results reveal a transient but significant fitness cost confined to the larval stage:

      Feeding rate: no significant difference between dsGP and dsGFP groups at any time point (24–120 h; Figure 11B), confirming that the observed metabolic changes are not attributable to reduced food intake.

      Larval weight: significantly reduced at 24 h (−29.10%, P < 0.05) and 48 h (−25.38%, P < 0.05; Figure 11C), demonstrating a measurable short-term cost of metabolic compensation.

      Pupal weight: no significant difference (Figure 11D), indicating full recovery before the pupal transition.

      Adult wing area: no significant difference (Figure 11E, Figures S5–S6), suggesting no impairment of flight capacity.

      Female fecundity (3-day egg production): no significant difference (Figure 11F), demonstrating no reduction in reproductive output.

      This pattern — transient larval weight loss with complete recovery of pupal weight, wing morphology, and reproductive performance — is consistent with our proposed model: GP suppression triggers protein catabolism to fuel gluconeogenesis (explaining the short-term weight loss), but the compensatory mechanism is sufficiently effective to restore metabolic homeostasis before pupation. These data strengthen the conclusion that GP is functionally non-essential for completing development and reproduction.

      (4) Broader conclusions regarding BPU class may require testing additional compounds.

      We agree. The revised manuscript now explicitly limits the biochemical conclusion to diflubenzuron: “DFB does not inhibit PxGP” rather than making broader claims about the BPU class as a whole. We discuss this as a limitation and note that testing additional BPU compounds would be needed before generalizing.

      (5) Some biochemical and cell-based observations would benefit from confirmation in whole insects.

      We have now provided whole-insect confirmation through: (i) GP enzyme activity measurements in RNAi-treated larvae (new Figure 10A, B); (ii) in vivo fitness assessment showing measurable physiological consequences of GP suppression (new Figure 11); and (iii) expression analysis of compensatory enzymes in intact larvae (new Figure 10C, D). These data bridge the gap between our cell-free biochemical observations and whole-organism biology.

      Reviewer #2 (Public review):

      (1) The central premise — that structural similarity among acylurea compounds implies shared targets — is not supported.

      We agree that the original manuscript overstated the significance of the shared acylurea core as a predictor of common biological activity. The Introduction has been substantially restructured to:

      Explicitly acknowledge the compelling genetic evidence from CRISPR/Cas9 experiments (Reference 5) establishing CHS as the primary site conferring BPU resistance.

      Reframe the study's objective: rather than proposing to “resolve” the BPU target controversy, the revised manuscript focuses on the systematic evaluation of GP as an independent insecticidal target and the discovery of a gluconeogenic compensation mechanism — questions with scientific value independent of the BPU mechanism debate.

      Remove the claim that the study “resolves the primary hypothesis.” The conclusion now states that our biochemical data demonstrate DFB does not inhibit PxGP, adding enzyme-level evidence to the existing genetic framework.

      (2) Target selectivity is determined by side-chain composition, not the shared acylurea core.

      We fully agree, and our new structural data now provide a molecular explanation for this principle at the atomic level. Molecular docking and MM/GBSA analysis (new Figure 12, new Table 1) reveal that both GPI and DFB anchor to PxGP through their common acylurea carbonyl groups (Arg193), but diverge dramatically in side-chain engagement:

      GPI's methoxyphenyl-methylurea moiety establishes extensive contacts with seven residues across both subunits (Asn44 and Val45 from chain A; Trp67, Gln71, Tyr75, Arg193, and Asp227 from chain B), binding at the allosteric site at the dimer interface — consistent with the experimentally determined binding mode of acylurea inhibitors in mammalian GP (PDB: 2ATI).

      DFB contacts six residues primarily from subunit B, and its difluorobenzoyl moiety remains entirely solvent-exposed without productive protein contacts.

      MM/GBSA analysis confirms GPI binds with substantially higher affinity (ΔG = −34.63 vs. −29.29 kcal/mol; ΔΔG = −5.34 kcal/mol), with van der Waals interactions as the dominant driver (Δ<sub>VDW</sub> = −11.49 kcal/mol), reflecting superior shape complementarity of GPI.

      These structural data directly support Reviewer 2's important point and are now presented as new Figure 12 and Table 1.

      (3) References 6–9 characterization.

      We have replaced the original citations (former References 6–9) in the Introduction with references that directly demonstrate the absence of CHS inhibition by BPUs in cell-free preparations: Cohen & Casida (1980) showed DFB did not inhibit Tribolium gut chitin synthetase; Mayer et al. (1981) and Cohen (1985) systematically confirmed that BPU-type insect growth regulators do not inhibit chitin synthase in cell-free assays; and Zhang & Zhu (2013) reported only slight in vitro CHS inhibition by DFB in Anopheles gambiae with no in vivo effect. We also cite authoritative reviews by Matsumura (2010) and Merzendorfer (2013) that contextualize this evidence gap. We thank Reviewer 2 for identifying this important citation issue, which has led to a substantially more accurate and well-supported presentation of the literature.

      (4) The term “dataology” is non-standard.

      This term has been removed and replaced with “data.” In accordance with eLife's policy on AI tools and technology, we have added a statement in the Materials and Methods section declaring that AI-based language editing tools were used for English grammar and style refinement. All scientific content was generated entirely by the authors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Direct assessment of gluconeogenic flux (e.g., metabolic tracer experiments).

      As discussed above, isotopic tracer experiments are beyond the current scope. We acknowledge this as a key limitation in the revised Discussion (see details in lines 560-565 of the revised manuscript). However, we now provide additional supporting evidence: the 30.78% decline in total protein at 24 h post-RNAi (from our new enzyme activity data, Figure S3) provides independent biochemical confirmation of protein catabolism, consistent with amino acid mobilization for gluconeogenesis (see details in lines 357-359 of the revised manuscript). We have also added discussion of potential future approaches, including measurements of key enzymes in amino acid catabolism (e.g., aspartate aminotransferase, glutamate aminotransferase) and lipid content dynamics (see details in lines 565-569 of the revised manuscript).

      (2) Expression or activity of alternative glycogen degradation enzymes (α-amylase, glycogen debranching enzymes).

      We have measured the expression of GBE (glycogen branching enzyme) and α-amylase by RT-qPCR in RNAi-treated insects. We also attempted to measure glycogen debranching enzyme (GDE), but multiple primer pairs failed to yield amplification products, likely due to sequence annotation issues; this is noted as a limitation.

      Results (new Figure 10C, D): GBE was significantly upregulated at 24 h (+29.24%), 48 h (+16.78%), and 96 h (+44.46%). α-Amylase was unchanged at all time points (see details in lines 377-381 of the revised manuscript). The selective upregulation of GBE but not α-amylase suggests a targeted compensatory response within the glycogen remodeling pathway.

      The absence of glycogen accumulation following GP knockdown may reflect reduced flux into glycogen synthesis (potentially through feedback inhibition of glycogen synthase) rather than activation of alternative degradative routes. This possibility is discussed in the revised manuscript, with glycogen synthase expression identified as a key target for future investigation (see details in lines 501-505 of the revised manuscript).

      (3) Fitness cost assessment (body size, flight capacity, reproductive performance).

      Complete data are now provided (new Figure 11A–F, Figures S5–S6):

      Author response table 1.

      The transient larval weight reduction (24–48 h) with complete pupal and adult recovery demonstrates that metabolic compensation carries a short-term physiological cost but is ultimately effective in maintaining developmental trajectory and reproductive fitness (see details in lines 388-408 of the revised manuscript).

      (4) Enzyme activity measurements in RNAi-treated insects.

      GP enzyme activity (GP-a) was measured in crude extracts from RNAi-treated larvae using a coupled-enzyme spectrophotometric assay kit (Solarbio BC3345) at 24, 48, 72, and 96 h post-injection (new Figure 10A, B).

      Two normalization approaches were used: - Per-protein activity: significant reduction only at 48 h (−10.35%, P < 0.05; Figure 10A). The modest per-protein reduction reflects a concurrent 30.78% decline in total protein at 24 h, which inflates per-protein specific activity when the protein pool shrinks. - Per-larva activity: significant reduction at 24 h (−27.57%, P < 0.05) and 48 h (−29.28%, P < 0.01; Figure 10B), confirming that RNAi-mediated transcript suppression translates to reduced enzyme function in vivo.

      The 30.78% decline in total protein at 24 h provides independent biochemical confirmation of protein catabolism — consistent with amino acid mobilization for gluconeogenesis (see details in lines 349-370 of the revised manuscript).

      (5) Scope of BPU conclusion — clarify whether additional compounds should be tested.

      The revised manuscript explicitly states that the biochemical conclusion applies to diflubenzuron specifically. We have added a Discussion paragraph noting that extending this conclusion to additional BPU compounds would require systematic testing, and that our structural analysis (Table 1) provides a framework for predicting which acyl urea side-chain architectures are compatible with GP binding (see details in lines 455-466 of the revised manuscript).

      (6) RNAi transcript recovery at 96 h and implications for non-essentiality.

      Our new enzyme activity data directly address this concern. GP activity (per-larva) showed partial recovery at 72 h and 96 h (Figure 10B), mirroring the transcript recovery pattern. However, the critical observation is that even during the period of maximum suppression (24–48 h), when per-larva GP activity was reduced by ~27–30%, larvae maintained glucose homeostasis and completed development. This confirms that GP is non-essential even during the period of strongest suppression. The revised Discussion addresses this point explicitly (see details in lines 479-488 of the revised manuscript).

      (7) GPI concentrations in larval exposure experiments and pharmacokinetic considerations.

      We have added a dedicated Discussion paragraph addressing this concern in detail. The GPI concentrations used (250–500 mg/L in diet) encompass a wide dose range; even at the highest concentration (500 mg/L, approximately 409,000-fold above the in vitro IC<sub>50</sub> of 2.96 nM), no toxicity was observed. We discuss several pharmacokinetic factors that may contribute to this apparent discrepancy, including limited oral bioavailability, metabolic inactivation by detoxification enzymes, and sequestration by hemolymph binding proteins. However, we note that the metabolic phenotype observed (elevated trehalose, reduced protein, upregulated gluconeogenic enzymes) provides indirect evidence that GPI does reach its target. The most parsimonious interpretation is that GPI achieves sufficient target engagement to partially suppress GP activity in vivo, but metabolic compensation renders this suppression non-lethal — an interpretation reinforced by our RNAi data, in which direct genetic suppression of GP (bypassing all pharmacokinetic barriers) similarly fails to cause mortality (see details in lines 536-552 of the revised manuscript).

      (8) Structural evidence that GPI binds PxGP comparably to its mammalian target.

      This has been comprehensively addressed through molecular docking and MM/GBSA analysis (new Figure 12, new Table 1). The PxGP homodimer structure was modeled using SWISS-MODEL with the human liver GP–acyl urea co-crystal structure (PDB: 2ATI) as the template. Docking and binding free energy calculations were performed in Cresset Flare V11.

      Key findings: GPI binds at the allosteric site at the dimer interface with ΔG = −34.63 kcal/mol, engaging seven residues across both subunits — a binding mode consistent with the experimentally determined site in mammalian GP. DFB binds with lower affinity (ΔG = −29.29 kcal/mol) and its difluorobenzoyl moiety is entirely solvent-exposed. Van der Waals interactions are the dominant driver of selectivity (Δ<sub>VDW</sub> = −11.49 kcal/mol). See new Figure 12 and Table 1 for complete data.

      (9) Dietary carbohydrate compensation and feeding behavior.

      Our new data directly address this concern. Feeding rate measurements show no significant difference between dsGP and dsGFP groups at any time point (24–120 h; Figure 11B), confirming that the metabolic changes are not attributable to altered food intake. A Discussion paragraph has been added acknowledging the potential contribution of dietary carbohydrates to glucose homeostasis and noting that starvation-challenge experiments would provide additional insight (see details in lines 570-581 of the revised manuscript).

      Minor comments and suggestions

      (1) Terminology.

      “Gluconeogenolysis” has been replaced with “gluconeogenesis” throughout the manuscript.

      (2) Typographical errors.

      A thorough language revision has been performed. “Over over four decades” and other errors have been corrected. The term “dataology” has been removed.

      (3) Metabolite normalization.

      We now present GP enzyme activity using two normalization approaches (per-protein and per-larva; Figure 10A, B) and discuss the implications of protein level changes on per-protein normalization in the Results section. For metabolite data, we have added a note in the Methods explaining our normalization approach and discussing how declining protein levels may influence interpretation (see details in lines 908-920 of the revised manuscript).

      (4) Clarity of pathway descriptions.

      The description of metabolic pathways has been simplified and a revised schematic figure has been included (Figure 13).

      (5) Figure clarity.

      Figures have been added with clearer labeling and simplified schematics. To improve clarity, we used red arrows to show the blocked metabolic flow when GP is inhibited, and green arrows to depict the activated gluconeogenic pathway (Figure 13).

      Cohen E, Casida JE. Inhibition of Tribolium gut chitin synthetase. Pestic Biochem Physiol. 1980;13(2):129-36. doi: 10.1016/0048-3575(80)90064-4.

      Mayer RT, Chen AC, DeLoach JR. Chitin synthesis inhibiting insect growth regulators do not inhibit chitin synthase. Experientia. 1981;37(4):337-8. doi: 10.1007/BF01959848.

      Cohen E. Chitin synthetase activity and inhibition in different insect microsomal preparations. Experientia. 1985;41(4):470-2. doi: 10.1007/BF01966152.

      Zhang X, Yan Zhu K. Biochemical characterization of chitin synthase activity and inhibition in the African malaria mosquito, Anopheles gambiae. Insect Sci. 2013;20(2):158-66. doi: 10.1111/j.1744-7917.2012.01568.x.

      Matsumura F. Studies on the action mechanism of benzoylurea insecticides to inhibit the process of chitin synthesis in insects: A review on the status of research activities in the past, the present and the future prospects. Pestic Biochem Physiol. 2010;97(2):133-9. doi: 10.1016/j.pestbp.2009.10.001.

      Merzendorfer H. Chitin synthesis inhibitors: old molecules and new developments. Insect Sci. 2013;20(2):121-38. doi: 10.1111/j.1744-7917.2012.01535.x

      Preiss J. Bacterial glycogen synthesis and its regulation. Annual review of microbiology. 1984;38:419-58. doi: 10.1146/annurev.mi.38.100184.002223.

      Janeček Š, Svensson B, MacGregor EA. α-Amylase: an enzyme specificity found in various families of glycoside hydrolases. Cell Mol Life Sci. 2014;71(7):1149-70. doi: 10.1007/s00018-013-1388-z.

      Waterhouse A, Bertoni M, Bienert S, Studer G, Tauriello G, Gumienny R, et al. SWISS-MODEL: homology modelling of protein structures and complexes. Nucleic Acids Res. 2018;46(W1):W296-W303. doi: 10.1093/nar/gky427.

      İnak E, De Rouck S, Van Leeuwen T. Molecular mechanisms of pesticide selectivity: Insights from acaricide toxicology. Pestic Biochem Physiol. 2025;213:106537. doi: 10.1016/j.pestbp.2025.106537.

      David MD. Insecticide ADME for support of early-phase discovery: combining classical and modern techniques. Pest Manage Sci. 2017;73(4):692-9. doi: 10.1002/ps.4345.

      Haunerland NH, Bowers WS. Binding of insecticides to lipophorin and arylphorin, two hemolymph proteins of Heliothis zea. Arch Insect Biochem Physiol. 1986;3(1):87-96. doi: 10.1002/arch.940030110.

      Other revisions

      Correction of primer sequences. Upon re‑checking the primer sequences during revision, we noticed that the originally reported primers for dsRNA synthesis of the GP gene (Table S1) were inadvertently copied incorrectly. The correct sequences have now been substituted in the revised manuscript (dsPxGP-F, dsPxGP-R). This correction does not affect any of the experimental data, results, or conclusions of the study. We apologize for the oversight.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In the revised manuscript, we have clarified several points that were raised by the reviewers. First, we now state more explicitly that the presence of intact env-containing Ty3/gypsy retrotransposons does not by itself demonstrate their mechanism of transmission, tissue specificity, infectivity, or current activity. We have therefore revised the wording throughout the manuscript to distinguish intact element structure and multicopy genomic expansion from experimentally demonstrated activity.

      Second, we performed targeted host-taxonomy concordance analyses on selected clades of the POL RT tree. These analyses do not exclude local horizontal transfer, particularly between closely related hosts, but they show that horizontal transfer alone is insufficient to explain the broader host-taxonomic structure observed across the dataset.

      Third, we incorporated representative viral and retroelement-associated fusogen proteins into our F-type ENV phylogenetic analysis and HSV/gB-type ENV structural comparison. These additions place the ENV proteins associated with Ty3/gypsy elements in a broader evolutionary context and strengthen the conclusion that these ENV associations are deeply diverged rather than recent derivatives of a single sampled viral lineage.

      Fourth, we added two each of entirely new Supplementary figures (S5 and S7) and Tables (S2 and S3) and substantially modified now Supplementary figure S8. Other figures have also been modified only to increase readability. The four tables from the original manuscript have not been modified although their numbering has changed.

      We believe that the revised manuscript is substantially improved in clarity, terminology and interpretive precision, while retaining the central conclusion that the association between env-like genes and Ty3/gypsy retrotransposons is ancient in metazoan evolution. Sincerely,

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript provides a comprehensive systematic analysis of envelope-containing Ty3/gypsy retrotransposons (errantiviruses) across metazoan genomes, including both invertebrates and ancient animal lineages. Using iterative tBLASTn mining of over 1,900 genomes, the authors catalog 1,512 intact retrotransposons with uninterrupted gag, pol, and env open reading frames. They show that these elements are widespread present in most metazoan phyla, including cnidarians, ctenophores, and tunicates-with active proliferation indicated by their multicopy status. Phylogenetic analyses distinguish "ancient" and "insect" errantivirus clades, while structural characterization (including AlphaFold2 modeling) reveals two major env types: paramyxovirus F-like and herpesvirus gB-like proteins. Although bot envelope types were identified in previous analyses two decades ago, the evolutionary provenance of these envelope genes was almost rudimentary and anecdotal (I can say this because I authored one of these studies). The results in the present study support an ancient origin for env acquisition in metazoan Ty3/gypsy elements, with subsequent vertical inheritance and limited recombination between env and pol domains. The paper also proposes an expanded definition of 'errantivirus' for env-carrying Ty3/gypsy elements outside Drosophila.

      Strengths:

      (1) Comprehensive Genomic Survey:

      The breadth of the genome search across non-model metazoan phyla yields an impressive dataset covering evolutionary breadth, with clear documentation of search iterations and validation criteria for intact elements.

      (2) Robust Phylogenetic Inference:

      The use of maximum likelihood trees on both pol and env domains, with thorough congruence analysis, convincingly separates ancient from lineage-specific elements and demonstrates co-evolution of env and pol within clades.

      (3) Structural Insights:

      AlphaFold2-based predictions provide high-confidence structural evidence that both env types have retained fusion-competent architectures, supporting the hypothesis of preserved functional potential.

      (4) Novelty and Scope:

      The study challenges previous assumptions of insect-centric or recent env acquisition and makes a compelling case for a Pre-Cambrian origin, significantly advancing our understanding of animal retroelement diversity and evolution. THIS IS A MAJOR ADVANCE.

      (5) Data Transparency:

      I appreciate that all data, code, and predicted structures are made openly available, facilitating reproducibility and future comparative analyses.

      Major Weaknesses

      (1) Functional Evidence Gaps:

      The work rests largely on sequence and structure prediction. No direct expression or experimental validation of envelope gene function or infectivity outside Drosophila is attempted, which would be valuable to corroborate the inferred roles of these glycoproteins in non-insect lineages. At least for some of these species, there are RNA-seq datasets that could be leveraged.

      We added a sentence in the discussion, subsection “The survival mechanism of errantiviruses in the genome”, citing our recent work now published (PMID: 41922845), explaining that the defence mechanism against errantiviruses appears to be conserved in insects beyond Drosophila, indirectly suggesting that their biology dependent on the presence of env—may be more universal.

      (2) Horizontal Transfer vs. Loss Hypotheses:

      The discussion argues primarily for vertical inheritance, but the somewhat sporadic phylogenetic distributions and long-branch effects suggest that loss and possibly rare horizontal events may contribute more than acknowledged. Explicit quantitative tests for horizontal transfer, or reconciliation analyses, would strengthen this conclusion. It's also worth pointing out that, unlike retrotransposons that can be found in genomes, any potential related viral envelopes must, by definition, have a spottier distribution due to sampling. I don't think this challenges any of the conclusions, but it must be acknowledged as something that could affect the strength of this conclusion

      We have added a targeted host-taxonomy concordance analysis for two well-sampled POL extended RT/connection subclades: an Annelida-associated clade from tree position A6 and a Lepidoptera-associated clade from tree position I1 (new Fig S5). Rather than attempting to infer exact numbers of duplication, loss and horizontal transfer events, which is difficult across highly expanded and unevenly sampled transposon families, we tested whether host-taxonomic labels were more clustered on the observed POL extended RT/connection topology than expected by chance. In the Annelida clade, highly supported small subclades showed strong host-family and host species concordance under host-label permutation tests. The Lepidoptera clade showed a more mixed pattern, but still contained several highly supported subclades enriched for related host groups at the superfamily or broader taxonomic level. These results do not exclude rare horizontal transfer, particularly between closely related hosts, but support the conclusion that the observed POL extended RT/connection trees retain significant host-taxonomic structure and are not consistent with frequent broad horizontal transfer between distantly related animal groups. We have added a paragraph in the Results section “Multiple intact elements of env-carrying Ty3/gypsy retrotransposons are found widespread across metazoan species” describing these observations, and also revised the Discussion to more explicitly acknowledge the possibilities of lineage-specific loss and the horizontal transfer.

      (3) Limited Taxon Sampling for Certain Phyla:

      Despite the impressive breadth, some ancient lineages (e.g., Porifera, Echinodermata) are negative, but the manuscript does not fully explore whether this reflects real biological absence, assembly quality, or insufficient sampling. A more systematic treatment of negative findings would clarify claims of ubiquity. However, I also believe this falls beyond the scope of this study.

      In the revised manuscript, we have added a targeted analysis of two representative genomes from each phylum. Although we did not detect full-length GAG-POL-ENV elements in these genomes, we recovered multiple full-length, multicopy GAG-POL Ty3/gypsy elements from all four genomes, many of which were flanked by predicted LTR sequences and associated with putative tRNA primer-binding sites. This suggests that the apparent absence of env-carrying elements in these representative Porifera and Echinodermata genomes is unlikely to be due simply to poor assembly quality or a general inability to recover intact Ty3/gypsy-like retrotransposons. We have added these data as the new Supplementary table S3 and revised the Results section “Multiple intact elements of env-carrying Ty3/gypsy retrotransposons are found widespread across metazoan species” to clarify that absence in these phyla may reflect true biological absence, lineage-specific loss, or incomplete taxon sampling.

      (4) Mechanistic Ambiguity:

      The proposed model that env-containing elements exploit ovarian somatic niches is plausible but extrapolated from Drosophila data; for most taxa, actual tissue specificity, lifecycle, or host interaction mechanisms remain speculative and, to me, a bit unreasonable.

      We stressed in the Discussion section “The survival mechanism of errantiviruses in the genome” that the mere presence of env gene does not imply the mechanism of transmission of retrotransposons.

      Minor Weaknesses:

      (1) Terminology and Nomenclature:

      The paper introduces and then generalizes the term "errantivirus" to non-insect elements. While this is logical, it may confuse readers familiar with the established, Drosophila-centric definition if not more explicitly clarified throughout. I also worry about changes being made without any input from the ICTV nomenclature committee, which just went through a thorough reclassification. Nevertheless, change is expected, and calling them all errantiviruses is entirely reasonable.

      We have revised the Results section and discussion where we introduced the term "errantivirus" to clarify that we use "errantivirus" operationally to refer to env-containing Ty3/gypsy retrotransposons identified in this study, rather than as a formal taxonomic proposal. We also now state explicitly that bona fide infectivity and amplification through the Drosophila-like ovarian somatic-cell route have not been experimentally established for most non-Drosophila elements. Our use of the term is therefore intended to distinguish env-containing Ty3/gypsy elements from related non-env containing Ty3/gypsy retrotransposons, while acknowledging that their biology outside Drosophila remains to be determined.

      (2) Figures and Supplementary Data Navigation:

      Some key phylogenies and domain alignments are found only in supplementary figures, occasionally hindering readability for non-expert audiences. Selected main-text inclusion of representative trees would benefit accessibility.

      We agree that clearer navigation between the main text and supplementary figures would improve readability. Although we considered moving selected supplementary phylogenies and alignments into the main figures, the main figures are already data-dense and are intended to provide representative summaries across many host groups and ENV types. We therefore retained the detailed trees and alignments as supplementary figures, where they can be shown at readable scale, but revised the manuscript to improve navigation. Specifically, we added signposting sentences in the Results where supplementary figures are mentioned, expanded the relevant figure legends, and clarified how each supplementary tree or alignment supports the corresponding main-text conclusion.

      (3) ORF Integrity Thresholds:

      The cutoff choices for defining "intact" elements (e.g., numbers/placement of stop codons, length ranges) are reasonable but only lightly justified. More rationale or sensitivity analysis would improve confidence in the inclusion criteria. For example, how did changing these criteria change the number of intact elements?

      We agree with the reviewer that the rationale for the ORF integrity thresholds should be stated more clearly. We have revised the Methods section "Identification of intact genomic copies of Ty3/gypsy errantiviruses" to clarify that the initial length, gap and stop-codon thresholds were deliberately permissive screening criteria, designed to avoid excluding divergent or non-canonical elements at the discovery stage. These initial filters were not used alone to define the final “intact” set. Candidate elements were subsequently subjected to multiple additional curation steps, including confirmation of Ty3/gypsy POL identity, recovery of full-length RT and Integrase domains within continuous ORFs, HHpred-based domain annotation of GAG, POL and ENV, and removal of elements with large domain truncations. Thus, the final set of intact elements is substantially more refined than would be implied by the initial stop codon or length thresholds alone.

      A full sensitivity analysis varying each threshold across the entire iterative discovery and manual-curation pipeline would be difficult to interpret, because changing early permissive filters would alter the candidate pool that then undergoes downstream structural and phylogenetic validation. Instead, we have clarified in the Methods that the early thresholds were intended as inclusive prefilters, whereas final inclusion required intact domain architecture and phylogenetic/domain support.

      (4) Minor Typos/Formatting:

      The paper contains sporadic typographical errors and formatting glitches (e.g., misaligned figure labels, unrendered symbols) that should be addressed.

      We now fixed these issues in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The authors first surveyed metazoan genomes to identify homologs of Drosophila errantiviruses and classified them into two groups, "insect" and "ancient" elements, supporting the hypothesis of an early evolutionary origin for these retrotransposons. They subsequently identified two distinct types of envelope proteins, one resembling the glycoprotein F of paramyxoviruses and the other akin to the glycoprotein B of herpesviruses. Despite differences in their primary amino acid sequences, these proteins display notable structural similarity in their predicted domain architectures. The congruence between the phylogenies of the envelope and pol genes further supports the ancient origin of the envelope genes, challenging earlier hypotheses that proposed recent recombination events with baculoviruses. Additional analysis of the Pol "bridge region" corroborated the divergence among these elements, consistent with a pattern of limited cross-species recombination. Finally, by comparing these elements with non-envelope-containing Gypsy retrotransposons, the authors concluded that errantiviruses originated from multiple elements independently.

      Strengths:

      The conclusions of this study are based on a comprehensive collection of errantiviruses identified across a wide range of metazoan genomes. These findings are further supported by multiple lines of evidence, including phylogenetic congruence and the diverse evolutionary origins of envelope genes. AlphaFold2-assisted protein domain structure analyses also provided key insights into the characterization of these elements. Together, these results present a compelling case that errantiviruses arose independently through multiple evolutionary events, extending well beyond previous hypotheses.

      Weaknesses:

      It would be beneficial to emphasize in the Abstract the potential impact of this work by more clearly articulating the current knowledge gap in the field. While the second paragraph of the Introduction briefly touches on this point, highlighting the broader significance in the Abstract would better capture readers' interest. Additionally, some methodological choices would benefit from clearer justification and explanation. For instance, in Figure 6, the selection of the bridge region/RNase H domain is not explicitly explained, leaving the rationale for its choice unclear. As a minor point, some figure labels and texts are too small and difficult to read, and improving their legibility would enhance overall clarity.

      We have revised the Abstract to more clearly state the knowledge gap addressed by this study: although env-containing Ty3/gypsy elements were known from Drosophila and sporadically reported in other animals, whether their association with env-like fusogen genes reflected recent, lineage-specific acquisitions or a much deeper evolutionary relationship remained unclear. We now highlight this broader significance in the Abstract and frame our results as evidence that env-containing Ty3/gypsy elements represent deeply diverged, genome-resident retroelements rather than a recent insect-specific phenomenon.

      We have also revised the Results, Methods and Figure 6 legend to explain why the

      RNase H-containing bridge region was analysed. Specifically, we now distinguish the Pol extended RT/connection region used for phylogenetic analysis from the RNase H-containing bridge region analysed structurally in Figure 6. We define the bridge region as the canonical RNase H domain together with the C-terminal region between RNase H and Integrase, and explain that this region was selected because RNase H-related and adjacent RNase H-like domains vary among LTR retroelement lineages. The bridge region architecture therefore provides an independent structural feature for comparing the “insect errantivirus” and “ancient errantivirus” groups.

      Finally, we have revised the figures and figure legends to improve readability. In particular, we enlarged labels where possible, clarified figure annotations, corrected cross-references between main and supplementary figures, and added signposting sentences in the Results so that readers can more easily connect the main conclusions to the supporting supplementary trees and alignments.

      Reviewer #3 (Public review):

      Summary and Significance:

      In this work, Cary and Hayashi address the important question of when, in evolution, certain mobile genetic elements (Ty3/gypsy-like non-LTR retrotransposons) associated with certain membrane fusion proteins (viral glycoprotein F or B-like proteins), which could allow these mobile genetic elements to be transferred between individual cells of a given host. It is debated in the literature whether the acquisition of membrane fusion proteins by non-LTR retrotransposons is a rather recent phenomenon that separately occurred in the ancestors of certain host species or whether the association with membrane fusion proteins is a much more ancient one, pre-dating the Cambrian explosion. Obviously, this question also touches upon the origin of the retroviruses, which can spread between individuals of a given host but seem restricted to vertebrates. Based on convincing data, Cary and Hayashi argue that an ancient association of non-LTR retrotransposons with membrane fusion proteins is most probable.

      Strengths:

      The authors take the smart approach to systematically retrieve apparently complete, intact, and recently functional Ty3/gypsy-like non-LTR retrotransposons that, next to their characteristic gag and pol genes, additionally carry sequences that are homologous to viral glycoprotein F (env-F) or viral glycoprotein B (env-B). They then construct and compare phylogenetic trees of the host species and individual encoded proteins and protein domains, where 3D-structure calculations and other features explain and corroborate the clustering within the phylogenetic trees. Congruence of phylogenetic trees and correlation of structural features is then taken as evidence for an infrequent recombination and a long-term co-evolution of the reverse transcriptase (encoded by the pol gene) and its respective putative membrane fusion gene (encoded by env-F or env-B). Importantly, the env-F and env-B containing retrotransposons do not form a monophyletic group among the Ty3/gypsy-like non-LTR retrotransposons, but are scattered throughout, supporting the idea of an originally ancient association followed by a random loss of env-F/env-B in individual branches of the tree (and rather rare re-associations via more recent recombinations).

      Overall, this is valuable, stimulating, and important work of general and fundamental interest, but still also somewhat incompletely explored, imprecisely explained, and insufficiently put into context for a more general audience.

      Weaknesses:

      Some points that might be considered and clarified:

      (1) Imprecise explanations, terms, and definitions:

      It might help to add a 'definitions box' or similar to precisely explain how the authors decided to use certain terms in this manuscript, and then use these terms consistently and with precision.

      (a) In particular, these are terms such as 'vertebrate retrovirus' vs 'retrovirus' vs 'endogenized retrovirus' vs 'endogenous retrovirus' vs 'non-LTR retrotransposon' and 'Ty3/gypsi-like retrotransposon' vs 'Ty3/gypsy retrotransposon' vs 'errantivirus'.

      We agree with the reviewer. We inserted a paragraph at the end of the first Results section, explaining how we define endogenous retroviruses (ERVs), Ty3/gypsy retrotransposons and errantiviruses.

      (b) The comment also applies to the term 'env' used for both 'env-F' and 'env-B', where often it remains unclear which of the two protein types the authors refer to. This is confusing, particularly in the methods, where the search for the respective homologs is described.

      We revised the manuscript and now used F-type env/ENV and HSV/gB-type env/ENV throughout the text. We also modified the method section where we explained the tBlastn search to clarify which ENV proteins were used initially for the search and how we classified them in later analyses.

      (c) Other examples are the use of the entire pol gene vs. pol-RT for the definition of the Ty3/gypsy clade and for the generation of phylogenetic trees (Methods and Figure S1), and the names for various portions of pol that appear without prior definition or explanation (e.g., 'pro' in Figure 1A, 'bridge' in Figure S1C, 'the chromodomain' in the text and Figure 7).

      We revised the manuscript and explained ‘pro’, ‘bridge’ and ‘the chromodomain’ in the Results section or figure legends when they first appear. Please refer to other sections of the response for pol-RT definition.

      (d) It is unclear from the main text which portions of pol were chosen to define pol-RT and why. The methods name the 'palm-and-fingers', 'thumb', and 'connections' domains to define RT. In the main text, the 'connection' domain is called 'tether' and is instead defined as part of the 'bridge' region following RT, which is not part of RT.

      We agree that our previous terminology around Pol domains was imprecise and could confuse readers. We have revised the manuscript to distinguish the region used for phylogenetic analysis from the region analysed structurally in Figure 6. The phylogenetic analysis used an extended RT/connection region, comprising the RT polymerase core together with the downstream connection subdomain. This connection subdomain is treated as part of retroviral RT in structural studies, but corresponds to a partial RNase H-like fold and has been interpreted evolutionarily as a degenerated RNase H-like tether domain. It is therefore broader than the RT polymerase core alone, but it is not the complete canonical RNase H domain.

      We now define the Figure 6 “bridge region” separately as the region spanning the canonical RNase H domain and the C-terminal region between RNase H and Integrase. Figure 6 shows that the invertebrate errantiviruses analysed retain an intact canonical RNase H domain immediately downstream of the extended RT/connection region, but differ in the additional downstream RNase H-like or mini-domain structures before Integrase. We have revised the Results, Methods and figure legends accordingly. We also acknowledge that a phylogeny based strictly on the RT polymerase core alone could differ in some local branch relationships, but the major conclusions are supported independently by the Integrase tree, host-taxonomic structure, ENV-type distribution, Pol bridge-region architecture and ENV structural features.

      (2) Insufficient broader context:

      (a) The introduction does not state what defines Ty3/gypsy non-LTR retrotransposons as compared to their closest relatives (Ty1/copia retrotransposons, BEL/pao retrotransposons, vertebrate retroviruses). This makes it difficult to judge the significance and generality of the findings.

      (b) The various known compositions of Ty3/gypsi-like retrotransposons are not mentioned and explained in the introduction (open reading frames, (poly-)proteins and protein domains, and their variable arrangement, enzymatic activities, and putative functions), and the distribution of Ty3/gypsi-like retrotransposons among eukaryotes remains unclear. The introduction does not mention that Ty3/gypsi-like retrotransposons apparently are absent from vertebrates, and Figure 7 is not very clear about whether or not it includes sequences from plants ('Chromoviridae').

      We agree that the Introduction needed more context on Ty3/gypsy retrotransposons. We have revised it to briefly state that LTR retrotransposons include several major lineages, including Ty1/copia, BEL/Pao, Ty3/gypsy and retrovirus-related elements, and that Ty3/gypsy elements are classified primarily by POL similarity and domain organisation. We also now explain that Ty3/gypsy retrotransposons typically encode GAG and POL proteins, with POL providing the enzymatic activities required for reverse transcription and integration, while noting that ORF arrangement and accessory domains can vary between lineages.

      Please note that we stated that our screen did not identify intact env-containing Ty3/gypsy elements in vertebrate genomes that were homologous to the invertebrate errantiviruses analysed here. This is not to say that non-env-containing Ty3/gypsy elements are also absent in vertebrate genomes. Finally, we revised the Figure 7 legend to make clear that the comparison includes representative non-env-containing Ty3/gypsy elements from animals, fungi and plants, including chromovirus or chromovirus-related elements.

      (c) The known association of Ty3/gypsi-like retrotransposons from different metazoan phyla with putative membrane fusion proteins (env-like) genes is mentioned in the introduction, but literature information, whether such associations also occur in the context of other retrotransposons (e.g., Ty1/ copia or BEL/pao), is not provided. The abstract is somewhat misleading in this respect. Finally, the different known types of env-like genes are not mentioned and explained as part of the introduction ('env-f', 'envB', 'retroviral env', others?)

      We expanded the introduction to introduce literature information of known env-associated retroelements, including Ty1/copia and BEL/pao and explained which ENV types are known to be associated to these elements.

      (d) Some key references and reviews might be added:

      - Pelisson, A. et al. (1994) https://www.embopress.org/doi/abs/10.1002/j.1460-2075.1994.tb06760.x (next to Song et al. (1994), for the identification of env in Ty3/gypsy)

      - Boeke, J.D. et al. (1999) In Virus Taxonomy: ICTV VIIth report. (ed. F.A. Murphy),. Springer-Verlag, New York. (cited by Malik et al. (2000) - for the definition and first use of the term 'errantivirus')

      - Eickbush, T.H. and Jamburuthugoda, V.K. (2008) https://doi.org/10.1016/j.virusres.2007.12.010 (on the classification of retrotransposons and their env-like genes)

      - Hayward, A. (2017) https://doi.org/10.1016/j.coviro.2017.06.006 (on scenarios of env acquisition)

      Thank you. We included these references in the introduction.

      (3) Incomplete analysis:

      (a) Mobile genetic elements are sometimes difficult to assemble correctly from shortread sequencing data. Did the authors confirm some of their newly identified elements by e.g., PCR analysis or re-identification in long-read sequencing data?

      Most newly identified elements are found in contigs/chromosomes that are longer than 100kb. The information of the contig/chromosome size, in which the representative copy of the identified elements are found, can be found in the column “CONTIG_SIZE” in supplementary table S1.

      (b) The authors mention somewhat on the side that there are Ty3/gypsy elements with a different arrangement (gag-env-pol instead of gag-pol-env). Why was this important feature apparently not used and correlated in the analysis? How does it map on the RT phylogenetic tree? Which type of env is found with either arrangement? Is there evidence for a loss of env also in the case of gag-env-pol elements?

      We agree that the non-canonical GAG-ENV-POL arrangement is an important feature that was insufficiently integrated into the analysis. We have revised the Results and figure annotations to make this clearer. Specifically, we now indicate GAG-ENV-POL elements in the POL tree in Fig S4 and in the HSV/gB-type ENV alignment/architecture figure S8. These elements are found in Nematoda, Bryozoa and Platyhelminthes and all carry HSV/gB-type ENV. They do not form a single monophyletic group in the Pol tree, but instead occur in distinct host-associated clades. They also show different HSV/gBtype cysteine-bridge architectures. Thus, the GAG-ENV-POL arrangement is unlikely to represent a single recent rearrangement event shared by all such elements; rather, it appears to be associated with several deeply diverged HSV/gB-type errantivirus lineages.

      We have not inferred specific env-loss events for GAG-ENV-POL elements, because doing so would require a separate analysis of related non-env-containing elements.

      (c) Sankey plots are insufficiently explained. How would inconsistencies between trees (recombinations) show up here? Why is there no Sankey plot for the analysis of env-B in Figure 5?

      We agree that the Sankey plot was insufficiently explained. We have revised the Figure 4 legend to clarify that the Sankey plot was used as a qualitative visual summary of global congruence between the Pol extended RT/connection phylogeny and the F-type ENV ectodomain phylogeny. We now state that ribbon crossing alone should not be interpreted as recombination, because tree drawings can be rotated without changing topology and the relative order of clades in the two displayed trees may differ. Instead, the relevant signal is whether Pol-defined clades map mostly to corresponding F-type ENV-defined clades. Strong discordance, potentially reflecting recombination, env exchange or poor phylogenetic resolution, would be expected to appear as extensive splitting or many-to-many connections between Pol and ENV clades.

      We did not include an equivalent Sankey plot for HSV/gB-type ENV in Figure 5 because we did not construct a global HSV/gB-type ENV phylogeny comparable to the F-type ENV ectodomain tree. Instead, HSV/gB-type ENV proteins were analysed by predicted structural organisation and cysteine-bridge architecture, which are shown in Figure 5 and Supplementary Figure S8.

      (d) Why are there no trees generated for env-F and env-B like proteins, including closely related homologous sequences that do NOT come from Ty3/gypsy retrotransposons (e.g., from the eukaryotic hosts, from other types of retrotransposons (Ty1/copia or BEL/pao), from viruses such as Herpesvirus and Baculovirus)? It would be informative whether the sequences from Ty3/gypsy cluster together in this case.

      We agree that comparison with homologous fusogens outside Ty3/gypsy retrotransposons is informative. We have therefore added an expanded F-type ENV ectodomain phylogeny that includes representative viral and retroelement-associated F-like proteins, including baculovirus F proteins, paramyxovirus and pneumovirus F proteins, and the BEL/Pao-associated Drosophila Roo F-like protein. In this expanded tree, the added viral sequences formed family-level clades within the broader F-type ENV diversity. Errantivirus F-type ENV proteins did not cluster as a shallow Ty3/gypsyspecific group or as a recent derivative of a single sampled viral family; instead, they spanned a level of diversity comparable to that separating major viral F-protein groups.

      For HSV/gB-type ENV proteins, we did not generate an equivalent global phylogeny because the primary sequences and domain organisations of viral class III fusogens and errantivirus HSV/gB-type ENV proteins were too divergent for reliable full ecto domain multiple-sequence alignment. Instead, we added a structural comparison with representative viral and retroelement-associated class III fusogens, including herpesvirus gB, rhabdovirus G, orthomyxovirus GP75/GP64-like proteins, baculovirus GP64 and BEL/Pao-associated gB-like proteins. This analysis showed that viral class III fusogens often retained family-specific cysteine-bridge architectures despite low primary-sequence identity. Errantivirus HSV/gB-type ENV groups showed a comparable pattern, retaining lineage-specific cysteine-bridge architectures despite extensive sequence divergence. We have revised the Results, Methods and supplementary figure legends to clarify these analyses and to distinguish the phylogenetic analysis of F-type ENV from the structural comparison of HSV/gB-type ENV.

      (e) Did the authors identify any other env-like ORFs (apart from env-F and env-B) among Ty3/gypsy retrotransposons? Did they identify other, non-env-like ORFs that might help in the analysis? It is not quite clear from the methods if the searches for env-F and envB - containing Ty3/gypsy elements were done separately and consecutively or somehow combined (the authors generally use 'env', and it is not clear which type of protein this refers to).

      We agree that this was not sufficiently clear. We have revised the Methods to clarify that the search was designed to identify Ty3/gypsy elements carrying ORFs structurally resembling known envelope/fusogen proteins. In the iterative tBLASTn searches, bait sequences representing both F-type ENV and HSV/gB-type ENV were included together in each round, rather than being searched as two entirely separate pipelines. Candidate elements were then annotated and classified by ORF structure, HHpred/domain similarity and structural prediction.

      Among intact Ty3/gypsy candidates recovered by this strategy, we identified two recurrent classes of env-like ORFs: F-type env and HSV/gB-type env. We did not identify an additional recurrent class of env-like ORF among the intact Ty3/gypsy elements analysed here. We also did not identify other recurrent non-env accessory ORFs that were informative for the phylogenetic analyses beyond the GAG, POL and ENV features described in the manuscript.

      (f) Why was the gag protein apparently not used to support the analysis? Are there different, unrelated types of gag among non-LTR retrotransposons? Does gag follow or break the pattern of co-evolution between RT and env-F/env-B?

      We agree that the role of GAG in the analysis should be clarified. GAG ORFs were used during element annotation to identify intact GAG-POL-ENV or GAG-ENV-POL retrotransposon architectures, but we did not use GAG as a major phylogenetic marker because GAG proteins are less conserved and less reliably alignable across deeply diverged Ty3/gypsy elements than the enzymatic POL domains. Our central question was the acquisition and long-term retention of env-like ORFs by POL-defined Ty3/gypsy retrotransposons. We have revised the Methods to clarify that GAG was used for structural annotation and intactness assessment, whereas phylogenetic analyses were based on the Pol extended RT/connection region and Integrase domain.

      (g) Data availability. The link given in the paper does not seem to work (https://github.com/RippeiHayashi/errantiviruses_2025/tree/main). It would be useful for the community to have the sequences of the newly identified Ty3/gypsy retrotransposons listed readily available (not just genome coordinates as in table S1), together with the respective annotations of ORFs and features.

      The GitHub repository that contains suggested data is made public. Please check the link again.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Additional Analyses That Could Strengthen Claims (but I concede might be well beyond the scope of this study):

      (1) Reconciliation and Gene Tree-Species Tree Analysis: Implementing explicit genetree-species-tree reconciliation (e.g., using Notung or ALE) could formally test the frequency of horizontal transfers vs. vertical inheritance in errantivirus evolution.

      We performed targeted host-taxonomy concordance analyses in representative well-sampled clades to address this question, as described in the revised manuscript.

      (2) Functional Validation in Non-Insect Hosts: RNA-seq or proteomics from representative non-insect hosts could reveal whether env genes are expressed, and-if possible-experimental assays for envelope function would move beyond computational inference. This is the one thing that might be done in a revision.

      We agree that direct functional validation in non-Drosophila species would substantially strengthen the conclusions. Such experiments are beyond the scope of the present revision. However, we now cite our recent work (PMID: 41922845) showing conserved piRNA-mediated defence against errantiviruses across insect orders, providing indirect evidence that these elements remain biologically active outside Drosophila.

      (3) LTR Age Dating: Estimating insertion ages using LTR divergence across clades would contextualize the timing of expansion events and help test hypotheses about ancient vs. recent proliferation.

      We agree that LTR divergence-based insertion dating would be informative for estimating the timing of recent expansion events. However, our main evolutionary conclusions concern the deeper history of env acquisition and long-term retention across metazoan lineages, rather than the precise insertion age of individual genomic copies. Because the dataset includes elements from highly divergent genomes with variable assembly quality and many multicopy families, systematic LTR dating across all clades would require additional curation and is beyond the scope of the present revision.

      (4) Comparative Host Defense Analysis: Surveying host antiviral or transposon defense systems (e.g., piRNA, APOBEC) in lineages rich in errantiviruses could test for signatures of recurrent molecular arms races.

      We agree that comparative analysis of host defence pathways, including piRNA and antiviral systems, is an important future direction. However, a systematic survey of host-defence evolution across all errantivirus-rich lineages is beyond the scope of the present revision.

      Reviewer #2 (Recommendations for the authors):

      Apart from my comments in the Public Review, I have a few additional minor recommendations:

      (1) The authors frequently use the term "active" to describe complete retrotransposons. However, in transposon biology, "active" implies recent or ongoing transpositional activity and should therefore be used with caution. Terms such as "complete" or "fulllength" would be more appropriate in this context.

      We agree that intact ORF structure should not be over-interpreted as evidence of recent mobilisation. We have therefore revised the manuscript to distinguish element intactness from evidence of recent expansion. Specifically, we now use “intact” or “full-length” to describe element structure. We also clarify that the presence of multiple highly similar copies in the same genome, defined as >98% nucleotide identity across >98% of the three ORFs, is evidence consistent with recent or ongoing genomic expansion, but not definitive proof of current transposition. These changes have been made in the Abstract, Results, Methods and Discussion.

      (2) In the third paragraph of the Results, the authors conclude that many identified errantiviruses were mobilized recently. However, since the search strategy specifically targets complete and uninterrupted elements, this may introduce a bias toward younger elements, making the conclusion somewhat circular.

      We agree and have revised the wording to avoid over-interpreting completeness as evidence of recent mobilisation. We now distinguish between intact/full-length element structure, which was part of our search strategy, and independent evidence for recent or ongoing mobilisation, such as the presence of multiple highly similar copies. We have therefore softened statements that previously implied that intact ORFs alone demonstrate recent activity.

      (3) The first paragraph of the Introduction requires appropriate references.

      We have now added references to enhance the readability of the first paragraph of the introduction.

      (4) It would benefit readers if the retrotranspositional process were briefly explained, including a description of what the tRNA primer binding site (PBS) is and its role in retrotransposition.

      In the same part of the introduction, we now describe the role of the tRNA primer binding site in retrotransposition.

      Reviewer #3 (Recommendations for the authors):

      Suggestions and requests for clarification currently are part of the public review, as they also point the reader to critical open questions if the authors decide not to amend their version of the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this important study, the authors have performed a zebrafish drug screen to identify suppressors of atherogenic lipoproteins. They utilize a well-established LipoGlo assay to find molecules that modulate these lipoproteins, identifying 49 potential hits. They perform some validation experiments, including studies linking enoxolone to its likely inhibitory effect on a specific transcription factor, HNF4alpha. Overall, the results are convincing and robust, and will open up new areas of exploration for those investigators interested in in vivo lipid biology.

      We appreciate the manuscript assessment and provide new data (see below) that significantly increases the “strength of the evidence”.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The authors performed a whole-organism chemical screen with over 3000 agents. Such screens are challenging, and the authors used strict criteria for determining hits. The conclusions of this study are well supported by the presented data.

      Weaknesses:

      There are areas within the study and writing that can be improved and extended, specifically within the gene expression studies.

      We appreciate the Reviewer recognizing the strength of the data and the challenging nature of performing a whole animal small molecule screen. With regard the “Weaknesses”, as the Reviewer suggested we improved and clarified the text and “extended” the study by adding analysis and discussion of six additional hits (see Figure 3 and Sup. Fig.1) and modified the results section with the addition of over a page of text describing the additional phenotyping results:

      “Secondary characterization of selected validated B-lp lowering hits.

      To further evaluate the biological relevance and potential mechanisms of selected validated hits, we performed secondary analyses assessing total B-lp levels, larval morphology, and lipoprotein size distribution. Our lab previously defined the method to measure total B-lp levels from whole animals by using the homogenate of a single zebrafish larva [35]. We confirmed that treatment of animals with 4 µM pomiferin significantly reduced total B-lp levels measured from homogenates collected from whole animals after treatment (p = 8.6x10<sup>-4</sup>; Figure 3A). However, pomiferin treatment (4 µM) produced animals with reduced body length and lethality at higher doses, suggesting developmental toxicity may confound interpretation of its B-lp-lowering effect.

      Treatment of animals with riboflavin tetrabutyrate (Figure 3B) and calcipotriene (Figure 3C) reduced total B-lp levels (p < 2x10<sup>-16</sup>) but did not affect larval morphology. A key feature of B-lps is their size, often a proxy for the total amount of lipid in the particle [35]. Particle size can impact the particle's lifetime (e.g. in metabolically healthy humans, small particles are cleared rapidly by the liver) [54–56]. Thus, we also assessed whether these compounds alter B-lp size distribution. Animals were treated for 48 h with vehicle, 5 µM lomitapide, or a drug of interest, and whole-animal homogenates were prepared and subjected to native polyacrylamide gel electrophoresis followed by chemiluminescent imaging. B-lps were classified into four classes based on gel migration: zero mobility (ZM), very low-density lipoproteins (VLDL), intermediate-density lipoproteins (IDL), and low-density lipoproteins (LDL) as previously described [35]. Lomitapide treatment effectively reduces VLDL particles and increases LDL particles [35] (Figure 3E), whereas riboflavin and calcipotriene did not affect lipoprotein classes. Thus, riboflavin tetrabutyrate and calcipotriene reduce total B-lp levels without overt developmental toxicity or changes in lipoprotein subclass distribution, suggesting they may act through mechanisms that decrease overall particle abundance rather than altering lipoprotein turnover or catabolism.

      Although doxycycline treatment lowered B-lp levels in whole fixed animals in the primary screen and validation studies, we did not observe a reduction in total B-lps in whole-animal homogenates (Figure 3D). However, we detected a slight increase in VLDL levels (p < 2x10<sup>-16</sup>; Figure 3E), suggesting that doxycycline may alter lipoprotein composition or distribution rather than total particle abundance.

      Alternatively, two structurally related compounds, thiethylperazine and prochlorperazine, at 4 µM significantly reduced (p < 1.4x10<sup>-10</sup> and p < 1.2x10<sup>-6</sup> respectively), B-lp levels measured from whole-animal homogenates (Figure 3F and 3H). Furthermore, both 8 µM thiethylperazine and 8 µM prochlorperazine increased relative LDL (p = 1.2x10<sup>-3</sup> and p = 2.3x10<sup>-3</sup>, respectively) and decreased relative VLDL levels (p = 8.4x10<sup>-4</sup> and p = 9.1x10<sup>-4</sup>, respectively; Figure 3G and 3I) suggesting a shift toward smaller lipoprotein particles and a potential alteration in lipid processing or clearance pathways. Together, these results highlight the diversity of mechanisms among validated hits, ranging from compounds that reduce total B-lp abundance without affecting B-lp class composition to those that shift B-lp class distribution, while also underscoring the importance of secondary assays to distinguish true B-lp modulators from those that likely produce a B-lp effect through generalized toxicity.

      Enoxolone significantly reduces B-lps in the larval zebrafish.

      Hit compounds were prioritized for follow-up studies based on reproducible dose-dependent responses, minimal toxicity as indicated by normal morphology over development, lack of direct NanoLuciferase inhibition, and the presence of literature suggesting potential links to lipid metabolism. One compound meeting these criteria was enoxolone, also known as 18β-Glycyrrhetinic acid, (Figure 2 Drug 20, Supplemental Table 1, Supplemental Figure 1T, Supplemental Figure 2L).”

      The Discussion now has the following additional text:

      “Further validation of these hits demonstrated a wide range of potential mechanisms of lipoprotein regulation. We identified hits that affected larval development, some hits that reduced total B-lp levels, and several structurally related compounds that directly reduced B-lp particle size (Figure 3).”

      Reviewer #2 (Public review):

      Strengths:

      The study was methodical and robust, using a published and well-validated zebrafish LipoGlo model. The authors validated the hits from the screen independently and considered the possibility that some drugs may have been detected as false positive results due to effects on the enzymatic activity of NanoLuciferase; only one hit, verteporfin, was shown to be a false positive. Using LipoGlo-Electrophoresis, the authors are able to obtain extra insights into the ApoB-lipoprotein size/subclass distribution. They showed that while enoxolone treatment reduces total B-lps, there are no overt changes in B-lp size distribution compared to vehicle-treated animals, other than a slight increase in the zero mobility (ZM) fraction, which contains very large particles and/or tissue aggregates. In contrast, the positive control, lomitapide, does show a change in B-lp size distribution compared to vehicle-treated animals - an increase in frequency of LDLs (low-density lipoprotein), but a decrease in VLDLs (very low density lipoprotein). This study also assesses the LipoGlo-Electrophoresis profile of HNF4⍺ inhibitors. Work in the zebrafish larvae means that the effect on overall development and an entire vertebrate organism can also be assessed. Finally, the authors applied a thorough statistical measure to define a hit, using the Strictly Standardized Mean Difference (SSMD) method.

      We appreciate that the Reviewer valued the rigour and robustness of our approach.

      Weaknesses:

      While the screen was thorough and well-validated, the authors missed a chance to provide a lot of extra significance to a wide range of readership. While the hits were thoroughly validated and displayed, the authors could have also presented the LipoGlo-Electrophoresis for all validated hits or at least a number of them. This would hugely increase the insights into these compounds. Also, the authors chose to validate and follow up a mechanism for Enoxolone, yet this hit was already known to modulate lipid metabolism through HNF4⍺, therefore, hugely limiting the impact of the paper. So what the authors have shown that is novel is only subtly added to this - consistent in vertebrate models, RNA sequencing of pathways, further validation of the HNF4⍺ pathway, and a profile of resulting B-lp size distribution. It seemed an easy way out to pick such a candidate, and they could have followed up by validating more thoroughly a completely novel drug. Also, the authors' prior paper showing the methodology also depicted complementary EM and LipoGlo-microscopy approaches. The microscopy especially, would have been an easy complementary add-on to the screen to really give extra insights into B-lp metabolism in a whole organism for all candidates. This felt like a missed opportunity.

      Here we agree and added Fig. 3 describing the phenotyping of 6 additional compounds including some LipoGlo-Electrophoresis analyses as suggested by the Reviewer. The text of the Results section was modified as described for Reviewer 1 (see above).

      Reviewer #3 (Public review):

      Strengths:

      The study uses a well-validated in vivo stain (LipoGlo) for measuring lipoproteins in the context of a developing whole organism with a quantitative read-out on a high-throughput platform, allowing for screening of thousands of compounds altering the complex metabolic/physiologic functions necessary for lipoprotein production.

      The use of genetic mutant HNF4alpha to assign the mechanism of action to the prime candidate compound studied (enoxolone) is a powerful approach for this challenging aspect of chemical genetics studies.

      We appreciate that the Reviewer understands how challenging it can be to assign a mechanism to any small molecule and thereby recognizes the power of the zebrafish model combined with our unique lipoprotein phenotyping tools.

      Weaknesses:

      As shown in Figure 5A, the HNF4alpha mutant homozygous -/- already lowers lipoproteins. Is it just that the mutant level is already at a minimum in this homozygous mutant (and thus enoxolone cannot induce even lower lipoprotein levels), or is it true that the enoxolone molecule is primarily acting through this TF (i.e. HNF4alpha homozygous mutant is truly epistatic to enoxolone function) as favored in the text.

      While it is definitely interesting to study enoxolone effects during whole embryo development, the link to HNF4alpha had previously been described in the literature, as pointed out by the authors. The generalizability of the approach to identify truly novel pathways remains to be fully realized, but sharing this available screen data to date will invite further inquiry and be very valuable to the community.

      Here too we agree that a link between enoxolone was proposed in the literature. However, we added quite a lot of additional insight regarding the transcriptional targets shared by HNF4alpha and enoxolone. The goal of identifying the mechanism(s) of action of other novel small molecule hits from the screen is important and that work is ongoing.

      Figure 5 - The same allele of HNF4alpha loss of function/hypomorph (rdu14) is used in both 5A and 5B, but labeled differently in each subpanel. This is explained in the figure legend, but could be updated to use the same nomenclature in both panels to clarify the Figure presentation.

      We thank the Reviewer for catching this and have modified the Figure (now Fig. 6) so the subpanels are labeled identically to avoid any confusion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors describe statistical methods to improve the calling of hits. In the results section, they discuss the use of fold change and of strictly standardized mean difference as criteria to account for the variability that occurs within in vivo screens. The combination of both criteria resulted in a significant reduction in the total number of hits, from 487 (16%) to 49 (1.6%), which is a large decrease in total hits. The authors should comment on whether some of the 487 compounds were randomly tested individually to confirm that they were true negatives and compared to the true hits of 49. For example, did the authors independently test calcipotriene and diphenylboric acid to confirm that these are truly negative?

      Initially, we did not directly retest any of the 487 compounds outside of the top 49 hits with the intention to compare their effect size to the top compound. While we do suspect that there is a chance that some of these compounds may still have a significant effect on lipoprotein levels, we prioritized compounds with the strongest overall biological effect size and largest statistically significant effect. We agree with the Reviewer that it would be interesting to continue to test more of these compounds and examine whether the lipoprotein reduction phenotype correlated with the primary screen effect, this would be a large experimental undertaking, though a randomized subset could be tested. Validation studies of Calcipotriene (Supplemental Figure 2E) showed a small, but significant, lipoprotein reduction at only the 0.25 µM dose across 3 independent experiments. Because the effect is small and not a dose-response, it is not a highly prioritized hit. We did not validate diphenlylboric acid in this study but will in the future.

      (2) The transcriptomics profile studies are not well described. The authors do not provide a detailed list of differential genes for each of the conditions in their dataset. This should be included.

      Supplemental Table 2 includes the fold change and p-values of all genes measured in differential expression analysis. We agree with the Reviewer and added new supplementary tables (Supplemental Table 5 and 7) that contain the differentially expressed genes at each treatment duration.

      (3) The Gene Ontology analysis results are not accompanied by a list of genes that match the GO terms. The authors should include these results. This part was confusing as it was not clear if the cholesterol pathway affected was due to an abundance of genes that were up- or down-regulated.

      We agree with Reviewer 1 and added Supplemental Table 6 that contains each gene ontology term with the associated DE genes that contribute to the significance of that term as well as explicit directionality information.

      (4) Figure 6A shows heat maps of differential gene expression, but there is no key provided in the figure or legend. Are they color-coded for fold change or log2 fold change?

      We agree that Figure 6A can be approved and added a key to the legend as suggested by the Reviewer.

      (5) The overlap of genes that are changed in hnf4a mutants and with enoxolone is provided as percentages, but the actual genes are not listed. Do these genes represent cholesterol biosynthesis pathways? There are other bioinformatic tools that the authors could apply to their dataset for further analyses. For example, enrichr (https://maayanlab.cloud/Enrichr/) is one tool to query for GO Biological processes, cell/tissue type, and even overlap of genes with genetic models and other disease states. An extension of these bioinformatic studies will be useful to determine if other pathways are relevant.

      We agree with the Reviewer and added a table of the genes that overlap between drug treatments and hnf4a mutants (Supplemental Table 7). We also appreciated the Reviewer’s suggestion to deploy Enrichr, which we have done and modified the Results section to now read:

      “Of the 439 differentially expressed genes from 12, 16, and 24 hours post-treatment, 34 differentially expressed genes are shared between all three treatment durations and are associated with gene ontology terms related to carbohydrate metabolism and signaling pathways (Figure 7D). We expanded this analysis using the bioinformatic tool Enrichr [68–70], which largely recapitulated gene ontology results described above. However, Enrichr analysis revealed significant overlap between several late enoxolone-responsive gene sets and transcriptional signatures associated with prochlorperazine, another compound identified in our screen (Figure 2, Figure 3, Supplemental Figure 1AO, Supplemental Figure 2Z). Enrichment of prochlorperazine-associated signatures was observed at 12 (prochlorperazine MCF7 up, 4/58 genes [INSIG1;IRF7;ISG15;ATF3], adjusted p-value = 0.006), 16 (prochlorperazine MCF7 up, 6/58 genes [INSIG1;DDIT4;IRF7;PMAIP1;ISG15;ATF3], adjusted p-value = 0.00008; prochlorperazine PC3 up, 3/29 genes [INSIG1;DDIT4;ATF3], adjusted p-value = 0.006), and 24 hours post-treatment (prochlorperazine PC3 up, 6/29 genes [DUSP5;INSIG1;DDIT3;TRIB3;SQSTM1;ATF3], adjusted p-value = 0.0001; prochlorperazine MCF7 up, 7/58 genes [DDIT3;INSIG1;IRF7;PMAIP1;ISG15;SAT1;ATF3], adjusted p-value = 0.0005). These results suggest that enoxolone and prochlorperazine may perturb overlapping molecular pathways, an observation that warrants further investigation. Ultimately, these data demonstrate distinct early and late responses to enoxolone treatment, and the early response modulates key lipid metabolism pathways.”

      (6) Figures 1E and 1F would benefit if the exact spots/points where enoxolone and other hits mentioned in the text were labelled.

      Great idea, we modified Figure 1F as suggested.

      (7) Figure 3 and Figure 4 graphs should state Fold change on the Y-axis title.

      We are thankful the Reviewer noticed this typo and we have fixed both Figures.

      Reviewer #2 (Recommendations for the authors):

      (1) To boost the impact for more readers, the authors should include the LipoGloElectrophoresis and LipoGlo-Microscopy results from a few more of the validated hits, especially ones that are completely novel (unlike Enoxolone, which already had a known role in lipid metabolism). Results on enoxolone are useful as a validation of the assay, mostly with some minor additional insights.

      We agreed with Reviewer 2 and added more validation testing of for a few hits (see response to Reviewer 1, Strengths and Reviewer 2, Weaknesses). We did not perform these experiments on each drug for technical reasons, mainly because these experiments are low-throughput (especially the Microscopy). Nonetheless, we performed many additional experiments to add phenotyping data for 6 new drugs that included multiple LipoGlo-Electrophoresis panels to an entirely new Figure.

      (2) The authors should include raw data from the screen from all drugs tested in the supplementary and then for which SSMD was calculated for, providing in an excel sheet or similar the values and how these were calculated, i.e. the 487 unique drugs that lower B-lp levels with an SSMD cutoff of < -1.0.

      This information was provided in the supplemental file as separate .csv files with the associated R script, which can be run locally and contains the SSMD functions. Considering the Reviewer comments, we ensured this information in provided in Supplemental Tables 1 and 2.

      (3) Page 3, lines 23-24: What does the 2 to 4 fold chance mean? Perhaps rewrite: Genetic mutations in Lipoprotein(a) increase the chance of heart attack or stroke 2-4 fold greater than without the mutation.

      We agree and the sentence now reads: Patients with genetic mutations in the Lipoprotein(a) encoding gene have a 2 to 4 fold increased risk of sudden heart attack or stroke.

      (4) Figure 1, for C and D, label some of the most significant hits and definitely show where exonolone lies.

      We agree see response to Reviewer 1 Pt6

      (5) Page 6, lines 1-9: I'm a bit confused why this is here if you do not present the data.

      We thought this was relevant information to share for researchers that running drug screens with positive controls and defining hit cutoffs. In light of the Reviewer’s comment, we removed the last sentence from this paragraph.

      (6) Page 6, line 8: This needs better justification of why you are validating enoxolone rather than other hits; otherwise, it could seem like cherry picking. Especially as enoxolone is known to affect lipid metabolism. Otherwise present more details of a couple of validated candidates.

      We agree. As the Reviewer requested, we validated more compounds (described above) and modified the text of the results to elaborate on our justification for selecting enoxolone for further study. The text of the results now reads: Hit compounds were prioritized for follow-up studies based on reproducible dose-dependent responses, minimal toxicity as indicated by normal morphology over development, lack of direct NanoLuciferase inhibition, and prior reports the presence of literature suggesting potential links to lipid metabolism. One compound meeting these criteria was enoxolone, also known as 18β-Glycyrrhetinic acid, (Figure 2 Drug 20, Supplemental Table 1, Supplemental Figure 1T, Supplemental Figure 2L).

      (7) Supplementary Figures 1 and 2: The resolution is too low, and the reader cannot even see the charts or the text.

      We agree and now have uploaded higher resolution images

      (8) Page 7, lines 31-35: Needs a higher resolution and magnified image to merit this 'offhand' statement. Also cite reference [59] here.

      We agree with the Reviewer and added magnified insets of the heart. As far as the suggestion of adding Ref 59, we do not see the connection to that paper (A point mutation decouples the lipid transfer activities of microsomal triglyceride transfer protein PLOS Genetics 16:e1008941)

      (9) Figure 3E: In addition to the proportions graph would be useful to also have an absolute amount of lipoprotein in each class graph.

      While there may be changes in total luminescence values from lane to lane in these gels, we have not fully validated the absolute quantitation of a full lane. We typically use plate-based whole-animal assays to determine total absolute lipoprotein levels and calculate the proportion of the whole lane for each lipoprotein class, as described in our prior publication detailing the assay. We do expect that there is some additional variation incorporated into the native PAGE assay due to sample freeze/thaw, dilution, and loading.

      (1) Page 10 lines 27-28: "Continual statin use for more than 1 year reduced circulating Blps and all-cause mortality by ~30% in individuals with high B-lp levels." This sentence doesn't seem right, intimates continual statin use causes death - I don't think that's right.

      We thank the Reviewer for catching this and have corrected the sentence. It now reads: “Continual statin use for more than 1 year in individuals with high B-lp levels reduced circulating Blps and lowered all-cause mortality by ~30% [12,13].”

      (11) Page 12, line 25: "canlikely" is a typo, should be can likely.

      We fixed that sentence and now reads: “Further, the drug screening paradigm we developed using the LipoGlo system is highly scalable and can be deployed to screen large novel drug libraries to identify many additional B-lp-lowering compounds.”

      (12) Figure 1 legend: "An ordered plot of each SSMD score measured from 5 μM lomitapide treated animals from each 96-well plate (n = 1381) relative to respective vehicle treatment." This comes across as though it's 5uM Iomitapide/vehicle. But it's the SSMD score of each drug compared to Iomitapide and relative to the respective vehicle (I think) - make it clearer.

      We agree and clarified the legend so it now reads: “…(D) An ordered plot of each SSMD score measured from positive control (5 µM lomitapide) treated animals from each 96-well plate (n = 1381) relative to respective vehicle treatment. The solid black line at y = 0 represents the divide in increased and decreased SSMD score, the solid blue line at y = -1.41 represents the curve's inflection point, and the dashed black line at y = -1 represents the SSMD (open circles) cutoff used to define a hit. “

      (13) Figure 1E: What is the x axis?

      Each data point on the x-axis represents each drug at every dose tested, we will clarify the test. The legend now reads: “(E) A plot of SSMD scores measured from each drug at each dose tested, each open circle represents the SSMD score of an individual drug at an individual dose.” In addition, “Compound (each dose tested)” was added to the x-axis of the figure panel.

      (14) Figure 2: Would you not have space to put the drug names in the figures? Where, for example, is enoxolone?

      We agree and have updated the figure accordingly.

      (15) Figure 3A-C: label enoxolone on the x axis.

      We agree and added this text to what is now Figure 4.

      (16) Figure 3D: Looks like delayed development with enoxolone, if left to grow, would the embryos develop normally?

      We did not examine if animals treated from 3-5 dpf develop normally beyond 5 dpf.

      (17) Figure E. Are stars all compared to vehicle control? Perhaps useful to have lines to indicate what are the significantly different relationships.

      Comparison in these experiments are always to the vehicle (negative control) and clarified the legend considering the Reviewer’s comment we modified the legend to now read,”… * <0.05 as compared to vehicle.”

      (18) Figure 3E: As well as the proportion of total lipoprotein, it would also be beneficial to see absolute lipoprotein levels.

      See above response to Reviewer 2 Pt9.

      (19) Figure 4: Does the overall health or size of the animal correlate with the luminescence score?

      While we do know that, in untreated animals, lipoprotein levels vary with age (and, thus, size), we have not examined this more granularly than in 24-hour time points after treatment. Further, we have not examined this in the context of a drug treatment.

      (20) Figure 4D: Again the absolute in each fraction would be meaningful, also the 5078 looks brighter?

      See above response to Reviewer 2 Pt9.

      (21) Figure 4A, 5A: Would the traces (line plots) not be useful here to see the overall dynamics over time?

      We considered presenting Figure 5A this way but decided to keep the plots as is because we wanted to show the individual points which make the line blots very difficult to read. Further, our analysis evaluates individual animals at each time point as it is not possible to follow the same animal over time.

      Reviewer #3 (Recommendations for the authors):

      Figure 4: Consider using standard scientific notation for the very small p values in some of the figure legends.

      We agree and adjusted the p-value notation as suggested.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments following re-submission:

      Overall, I think the authors have done a satisfactory job of addressing most of the points I raised.

      There’s one final issue which I think still needs better discussion.

      I think reviewer 2 articulated better than I have the point I was concerned about: the relationship between JNDs and metamers as depicted in the schematics and indeed in the whole conceptualization.

      I think the issue here is that there seems to be a conflating of two concepts- ’subthreshold’ and ’metamer’-and I’m not convinced it is entirely unproblematic. It’s true that two stimuli that cannot be discriminated from one another due to the physical differences being too small to detect reliably by the visual system are a form of metamer in the strict definition ’physically different, but perceptually the same’.

      However, I don’t think this is the scientifically substantial notion of metamer that enabled insights into trichromacy. That form of metamerism is due to the principle of univariance in feature encoding, and involves conditions in which physically very different stimuli are mapped to one and the same point in sensory encoding space whether or not there is any noise in the system. When I say ’physically very different’ I mean different by a large enough amount that they would be far above threshold, potentially orders of magnitude larger than a JND if the system’s noise properties were identical but the system used a different sensory basis set to measure them. This seems to be a very different kind of ’physically different, but perceptually the same’.

      We are in full agreement with this. Typically, the notion of metamers is about deterministic information loss, which can be modeled as a projection from a high-dimensional physical space to a lower-dimensional perceptual space. This is the topic of the paper, and is analogous to the work on color matching in the 19th century.

      In contrast, sensitivity to small differences within the perceptual space due to internal noise is usually addressed by other methods, such as signal detection theory, a topic which is not the focus of this paper. It is analogous to the work on discriminability of colors such as MacAdam ellipses (MacAdam, 1942). It is instructive to look at progress in the color field. The color matching experiment and the question of metamerism is quite well worked out, whereas the question of how to quantify discriminability within that space has been an ongoing topic of investigation for over a century.

      Here, we aim to develop and test a model of metamers, analogous to the color matching experiments, but we do not attempt to develop a model of discriminability. Nevertheless, while the two types of information loss are conceptually distinct, they are both present in the nervous system of the observer, and both are reflected in the performance vs. scaling plots in our paper. We have added clarifications about this point in the introduction on page 2, starting on line 45, and the discussion, starting on page 16, line 397.

      Finally, regarding physical differences between stimuli: the differences between target images and synthesized metamers are quite large (high mean squared error), as shown in Appendix 5. In no condition did we present subjects with stimulus pairs that were physically similar.

      I do think the notion of metamerism can obviously be very usefully extended beyond photoreceptors and photon absorptions. In the interesting case of texture metamers, what I think is meant is that stimuli would be discriminable if scrutinised in the fovea, but because they have the same statistics they are treated as equivalent.

      The notion of “texture metamers” is perhaps a reference to the work by Freeman and Simoncelli (2011), whose stimuli are similar to ours: when synthesized using a model with sufficiently small scaling, they are indiscriminable, and therefore metamers. The reviewer is of course correct that the stimulus pairs are not metamers when the observers move their eyes due to differences in spatial encoding as a function of eccentricity. That is, they are only metameric under a specific set of viewing conditions, and they are not metameric when those conditions are violated. The same is true for color metamers, as the spectral sensitivity of the cones also differ with eccentricity (Stockman and Sharpe, 2000).

      I think the discussion of this could still be clearly articulated in the manuscript. It would benefit from a more thorough discussion of the difference between metamerism and subthreshold, especially in the context of the Voronoi diagrams at the beginning.

      We agree that a more thorough discussion of the diagrams could help clarify the issues to the reader. We have modified the caption of figure 1 with the goal of clarifying interpretation of the diagrams, and see also our discussion earlier in this note about noise and discriminability.

      It needs to be made clear to the reader why it is that two stimuli that are physically similar (e.g., just spanning one of the edges in the diagram) can be discriminable, while at the same time, two stimuli that are very different (e.g., at opposite ends of a cell) can’t.

      Do the cells include BOTH those sets of stimuli that cannot be discriminated just because of internal noise AND those that can’t be discriminated because they are projected to literally the same point in the sensory encoding space? What are the strengths and limits of models that involve the strict binarization of sensory representations, and how can they be integrated with models dealing with continuous differences? These seem like important background concepts that ought to be included in either the introduction of discussion sections. In this context it might also be helpful to refer to the notion of ’visual equivalence’ as described by:

      This is an important point and we appreciate the reviewer raising it. In brief, as one traverses a region in one of the Voronoi diagrams, the images are changing physically but subject to the constraint that they all project to the same single point in the reduced perceptual space. When one crosses from one region to another, the images now project to a different point in the perceptual space. Whether or not that the two locations in the perceptual space are distant enough to be distinguishable given the internal noise is a question pertaining to the topic of JNDs in the perceptual space, rather than the mapping from physical space to the perceptual space. We do not address that question in detail in this paper, though we do now reference it in the caption of figure 1, as well as in the new sections in the introduction and discussion mentioned earlier in this response.

      We do note that the perceptual space is not discrete: the model outputs are real-valued. The apparent discretization is a limitation of the simplified 2-D schematics.

      Ramanarayanan, G., Ferwerda, J., Walter, B., & Bala, K. (2007). Visual equivalence: towards a new standard for image fidelity.ACM Transactions on Graphics (TOG), 26(3), 76-es.

      Other than that, I congratulate the authors on a very interesting study, and look forward to reading the final version.

      Reviewer #2 (Public review):

      Summary:

      The authors have improved clarity overall and have spoken to most of the issues raised by the reviewers. There are still two outstanding problems however, where issues raised during the review were inappropriately dismissed in the manuscript. These should be explicitly addressed as limitations to the results presented (no eye tracking), and early pilot experiments that informed the experiments as presented (pink noise) rather than brushed off as ’unnecessary’ and ’would be uninformative’.

      Eye tracking:

      It is generally accepted that experiments testing stimuli presented at specific locations in peripheral vision require eye tracking to ensure that the stimulus is presented as expected, in particular, in the correct location. As I stated in the previous round of review, while a stimulus presentation time of 200ms does help eliminate some saccades, it does not eliminate the possibility that subjects were not fixating well during stimulus onset. I am also unclear what the authors mean by ’trained observer’ in this context, though the authors state that an author subject in a different portion of the paper is an ’expert observer’. Does this mean the ’trained observers’ are non-expert recruited subjects?

      Given the conditions tested differ from previous work (Freeman & Simoncelli, 2011) ‘these differences are a main contribution of the paper!’ which DID include eye tracking in a subset of subjects, it is entirely possible to get similar results to this work in the context of non eye-tracking controlled stimulus presentation. The reasons now in the manuscript are not reasons that make eye tracking ’considered unnecessary’.

      I appreciate that the authors now state the lack of eye tracking explicitly, but believe the paper needs to at least state that this is a limitation of the results reported, and eyetracking being ’considered unnecessary’ is unreasonable, nor a norm in this subfield.

      By “trained” observers, we mean people who were recruited from the community of vision science labs at NYU and who have participated in many visual psychophysics experiments. All of the participants are “trained” in this sense, and are thus used to maintaining fixation while performing peripheral tasks. One of these participants, an author, was also an expert in the specific content area of the paper. By “expert”, we mean high familiarity with the stimulus types and models employed in the paper.

      We have now further clarified this in the text in the subsection of the methods on Observers, on page 22.

      We also discuss the issue at greater length in the methods subsection Apparatus, on page 26. We removed the word "unnecessary" and make it clear that while we don’t think our results are undermined, the lack of eye tracking is nonetheless a limitation.

      N=1: The authors now state clearly the limitations of a single subject in the manuscript, and state the expertise level of this subject.

      Large number of trials: The authors now address this and include an enumeration of the large number of trials.

      Simple Models / Physiology comparison: I support the choice to reduce claims regarding tight connections to physiology, and appreciate the explanation of the luminance model.

      Previous Work: I appreciate the author’s changes to the introduction, both in discussing previous work and citation fixes.

      Blurred White, Pink Noise: While the authors now address pink noise, the explanation for such stimuli being expected to be uninformative is confusing to me. The manuscript now first states that pink noise is a natural choice, then claims it would be uninformative, while also stating in the rebuttal (not the manuscript) that they tried it and it indeed reduced the artifacts they note. The logic of the experiments indeed relies on finding the smallest critical scaling value, which is measured by subjects determining if a synthesis is similar or different to a target or second synth. A synthesis free from artifacts would surely affect the subjects responses and the smallest critical scaling measured.

      The statement that the authors experimented with pink noise early on and found this able to address the artifacts should be stated in the manuscript itself, not just in the rebuttal, and the blanket statement that this experiment would be ’uninformative’ is incorrect. Surely this early pilot the authors mention in the rebuttal was informative to designing the experiments that appear in the final paper, and would be an informative experiment to include.

      First, we did render some test stimuli with pink noise seeds, but we did not collect psychophysical data, hence there are no results we could add. Visual inspection of these stimuli was indeed clarifying in the following sense. The pink noise stimuli had fewer high-frequency “artifacts”. If our goal was to synthesize stimuli that are indistinguishable from the original stimulus, as one might do to save compute power when in a device that for foveated rendering, then starting with pink noise would be better than starting with white noise. Our purpose was just the opposite. For our experiments, the artifacts were just what we wanted: the more artifacts, the better. The reason is that a strongest test of a metamer model is whether two stimuli that are as physically different from one another as possible, are nonetheless indistinguishable when their model representations are the same. Stimuli synthesized from pink noise seeds are harder to discriminate from the target stimulus, not easier. Thus using them in an experiment would result in a larger estimate of critical scaling. Since our explicit goal was to estimate the smallest critical scaling window, these stimuli would not bring us closer to our goal. As the reviewer points out, these metamers were “informative” in the sense that they informed our experimental design, but they are “uninformative” (relative to white noise seeds) for estimating the critical scaling.

      We have updated our description in the discussion starting on page 19, line 449, and included a new appendix to demonstrate this point (appendix 2 on page 35).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Typo: p. 19, l. 439: ’Why does asymptotic performance, but not critical scaling, depends on image content?’": remove ’s’ from ’depends’.

      Fixed.

      Reviewer #2 (Recommendations for the authors):

      Recommendations: State that the lack of eye tracking to control stimulus presentation is a limitation of the results presented.

      Remove the claim that pink noise or filtered white noise seeds would be uninformative, and mention the fact that the authors in fact experimented with pink noise seeds in an early version of the experiments (which was surely informative to the experimental setup as presented here).

      Addressed as described above.

      References

      Freeman J, Simoncelli EP. Metamers of the ventral stream. Nature Neuroscience. 2011 aug; 14(9):1195–1201. doi: 10.1038/nn.2889.

      MacAdam DL. Visual Sensitivities To Color Differences in Daylight*. Journal of the Optical Society of America. 1942 may; 32(5):247. http://dx.doi.org/10.1364/josa.32.000247, doi: 10.1364/josa.32.000247.

      Stockman A, Sharpe LT. The Spectral Sensitivities of the Middle- and Long-Wavelength-Sensitive Cones Derived From Measurements in Observers of Known Genotype. Vision Research. 2000 jun; 40(13):1711–1737. doi: 10.1016/s0042-6989(00)00021-3.

    1. Author response:

      We greatly appreciate the positive and constructive comments from the reviewers, which recognized the intellectual contributions of our manuscript and also captured its limitations. Below is our response to the three major points from Reviewer #1 and the comments from Reviewer #2.

      Response to Reviewer #1

      (1) Use of transient transfection in non-hematopoietic cells. We appreciate the reviewer’s point regarding physiological relevance. Our goal in these cellular microscopy experiments was to dissect the biophysical principles of NUP98::KDM5A (including its mutants) condensate formation under controlled expression levels (concentration). While AML model systems driven by NUP98::KDM5A are available, they do not offer this possibility because of pre-existing NUP98::KDM5A expression. The suspension culture of hematopoietic cells also brings practical challenges for correlative FISH+IF and high-resolution microscopy analysis. We agree that validating our observed behaviors in AML models would be valuable, e.g. by creating HSPC lines with inducible expression of tagged NUP98::KDM5A and its mutants, but such experiments fall outside the scope of the current study. We will add text acknowledging this limitation and clarifying that our mechanistic conclusions are grounded in biophysical principles that should generalize across cell types.

      (2) Correlation between H3K4me3 and gene activation. We agree that active transcription correlates with H3K4me3, and that this baseline relationship must be considered. Our analysis explicitly uses fold-change between patient cells and healthy controls as the readout. This comparison inherently accounts for the activating effect of H3K4me3 itself. Regarding the reviewer’s comment that “many H3K4me3-positive genomic regions do not show NUP98::KDM5A binding”, we would like to clarify that this point is exactly what our manuscript aims to explain. Our cell line studies demonstrate that, at a patient-relevant expression level, NUP98::KDM5A condensates form preferentially at H3K4me3 locus with high local mark density. Considering that NUP98::KDM5A concentration in the nucleus is lower than the K_D between KDM5A PHD3 and H3K4me3, this means that H3K4me3 loci with lower mark density will not see NUP98::KDM5A binding without the high local concentration of the fusion protein (as a result of condensate formation). This is consistent with our observation that genes with the highest local density show disproportionately stronger upregulation. We will further clarify this point in the revised manuscript.

      (3) Generalization to other NUP98 fusions lacking PHD domains. We appreciate this important conceptual question. We will expand the discussion to note that many NUP98 fusions, despite diverse partner domains, produce similar transcriptional programs. Our current hypothesis is that the highly active status of the HOX cluster genes in HSPC attracts NUP98 oncofusions targeting H3K4me3 (e.g. KDM5A and PHF23), while NUP98::HOXA9 and other transcription factor fusions directly target the HOX cluster via DNA sequence recognition. This convergence and the broad targets of the dysregulated HOX transcription factors explain the transcriptional program similarity, sustained by the previously described enrichment of transcriptional co-activators by NUP98 oncofusion condensates. We will explicitly discuss this hypothesis in our revised manuscript.

      Response to Reviewer #2

      We thank the reviewer for the thoughtful evaluation and agree that the causal link between H3K4me3 recognition, condensate formation, and gene activation remains partly correlative. Experiments such as qPCR or ChIP following expression of WT versus mutant constructs would indeed strengthen the causal chain. However, performing these assays across multiple constructs and loci in a physiologically relevant system is not feasible within the current revision cycle. We have added text acknowledging this limitation and clarifying that our study focuses on establishing the biophysical mechanism of targeting, while functional consequences are inferred from patient datasets rather than new perturbation experiments.

    1. Author response:

      Reviewer 1:

      (1) Major Concerns (highest priority):

      There is a lack of detail about the methodology in the results section/figure legends, which makes it difficult to interpret the data. Sometimes, adequate information is also not included in the methods themselves. For example, Figure 1A: how many doses did each mouse receive? How long after dosing were animals sacrificed? Figures 1E and 2A: is this RNA-seq analysis?

      Thank you for pointing this out. We will provide additional details about the methodology in the results, figure legends, and methods to clarify our findings.

      Using a second EGFR inhibitor for some of the key experiments would increase the rigor of the studies shown.

      Thank you for this suggestion. We agree that a second EGFR inhibitor may increase the scientific rigor of the findings. However, at the time the experiments were completed, osimertinib was the clear standard of care first-line therapy for EGFR-mutated lung cancer and the clinical relevance of using other inhibitors in these experiments is not clear.

      (2) Nice to have experiments:

      Using CRISPR KO of CDK4 in the CDK4-amplified HCC827 and testing response to Osi and presence of replication stress would also increase the rigor of the studies.

      Thank you for this suggestion. We agree that this would increase the rigor of the studies, but these experiments are currently beyond the scope of this manuscript.

      The authors show that in their patient data, some cell cycle regulators which are amplified in NSCLC at similar rates to CDK4/6, such as CCNE1, had no increase in FGA. Overexpressing CCNE1 and testing Osi response in their cell models would be a nice test of their proposed mechanism that it is the genomic instability and FGA that are driving resistance. This wouldn't need to be done in vivo, but could be done using cell culture-based methods.

      Thank you for this suggestion. For this manuscript, we have chosen to focus on the role of CDK4 and CDK6 in osimertinib resistance as they have shown the clearest correlation with decreased responsiveness to EGFR TKI treatment in other studies. We agree that assessing the role of CCNE1 in osimertinib resistance is important and will address this in future studies.

      Similarly, testing the overexpression of some of the proposed target genes, such as STEAP1 and AGR2, on the therapeutic response to Osi in cell culture would also be a nice test of the mechanism proposed.

      Thank you for this suggestion. We agree that these experiments are important, but they are currently beyond the scope of this study.

      Reviewer 2:

      Major Comments:

      (1) Reliance on overexpression models over loss-of-function

      The mechanistic studies mainly rely on overexpression of CDK4 and CDK6 to simulate the amplified state. Although the authors argued that the level of overexpression mimics that observed in resistant tumors, a complementary study in which CDK4/CDK6 were suppressed in a model where CDK4/CDK6 is amplified (such as HCC827 or TH116), and replication stress and osimertinib sensitivity tested would greatly strengthen their observations. Indeed, there is mention of CDK4 constructs to perform knockdown studies in the methods, but those studies are not included in this submission.

      Thank you for this suggestion. We agree that loss of function studies are an important complement to over expression studies. However, we have chosen to focus on the use of pharmacologic inhibitors of CDK4/6, which are more clinically relevant than knockdown studies. We demonstrate that the CDK4/6 inhibitor palbociclib is able to prevent replication stress and restore osimertinib sensitivity in CDK6 amplified TH116 patient-derived xenografts (Figure 5 and Supplementary Figure 10).

      (2) Mechanistic link between genomic instability and specific gene amplifications

      The authors highlight that CDK4/6 activation leads to recurrent copy number gains and transcriptional upregulation of specific pro-tumor genes including AGR2, ASNS, and STEAP1. While the paper establishes that CDK4/6 overexpression causes general genomic instability (increased FGA), it does not mechanistically explain why these specific genes are consistently amplified. The authors should investigate or discuss whether these specific loci are inherently fragile under replication stress, if they are direct downstream targets of the E2F transcriptional program, or if this is a result of random genomic instability followed by strong positive selection under osimertinib pressure.

      Thank you for this suggestion. We will provide additional discussion about whether these specific loci are likely to be inherently fragile under replication stress, if they are direct downstream targets of the E2F transcriptional program, or if it is more likely the result of random genomic instability followed by strong positive selection under osimertinib pressure.

      (3) Discrepancies in tumor mutational burden (TMB) reporting

      There is a slight contradiction regarding the TMB data that needs to be clarified for readers. The manuscript states that in the clinical datasets, "EGFR-mutant LUAD harboring cell cycle gene alterations exhibited significantly elevated FGA and TMB relative to cell cycle-negative tumors" (Line 265-266). However, in the next section, the authors say, "Notably, no corresponding increase in TMB was observed with CDK4 or CDK6 CNA, similar to our findings in preclinical models" (Line 271-273). The authors should clarify or discuss why broad cell cycle alterations correlate with high TMB, while CDK4/6-specific alterations drive structural instability (FGA) without increasing TMB.

      Thank you for this suggestion. We will clarify our findings and provide further discussion about why there may be a difference between the effects of broad cell cycle alterations and CDK4/6-specific alterations on TMB and structural genomic instability.

    1. Author response:

      Reviewer #1 (Public Review):

      Summary:

      In this manuscript, the authors investigated the effect of chronic activation of dopamine neurons using chemogenetics. Using Gq-DREADDs, the authors chronically activated midbrain dopamine neurons and observed that these neurons, particularly their axons, exhibit increased vulnerability and degeneration, resembling the pathological symptoms of Parkinson's disease. Baseline calcium levels in midbrain dopamine neurons were also significantly elevated following the chronic activation. Lastly, to identify cellular and circuit-level changes in response to dopaminergic neuronal degeneration caused by chronic activation, the authors employed spatial genomics (Visium) and revealed comprehensive changes in gene expression in the mouse model subjected to chronic activation. In conclusion, this study presents novel data on the consequences of chronic hyperactivation of midbrain dopamine neurons.

      Strengths:

      This study provides direct evidence that the chronic activation of dopamine neurons is toxic and gives rise to neurodegeneration. In addition, the authors achieved the chronic activation of dopamine neurons using water application of clozapine-N-oxide (CNO), a method not commonly employed by researchers. This approach may offer new insights into pathophysiological alterations of dopamine neurons in Parkinson's disease. The authors also utilized state-of-the-art spatial gene expression analysis, which can provide valuable information for other researchers studying dopamine neurons. Although the authors did not elucidate the mechanisms underlying dopaminergic neuronal and axonal death, they presented a substantial number of intriguing ideas in their discussion, which are worth further investigation.

      We thank the reviewer for these positive comments.

      Weaknesses:

      Many claims raised in this paper are only partially supported by the experimental results. So, additional data are necessary to strengthen the claims. The effects of chronic activation of dopamine neurons are intriguing; however, this paper does not go beyond reporting phenomena. It lacks a comprehensive explanation for the degeneration of dopamine neurons and their axons. While the authors proposed possible mechanisms for the degeneration in their discussion, such as differentially expressed genes, these remain experimentally unexplored.

      We thank the reviewer for this review. We do believe that the manuscript has a mechanistic component, as the central experiments involve direct manipulation of neuronal activity, and we show an increase in calcium levels and gene expression changes in dopamine neurons that coincide with the degeneration. However, we agree that deeper mechanistic investigation would strengthen the conclusions of the paper. We have planned several important revisions, including the addition of CNO behavioral controls, manipulation of intracellular calcium using isradipine, additional transcriptomics experiments and further validation of findings. We anticipate that these additions will significantly bolster the conclusions of the paper.

      Reviewer #2 (Public Review):

      Summary:

      Rademacher et al. present a paper showing that chronic chemogenetic excitation of dopaminergic neurons in the mouse midbrain results in differential degeneration of axons and somas across distinct regions (SNc vs VTA). These findings are important. This mouse model also has the advantage of showing a axon-first degeneration over an experimentally-useful time course (2-4 weeks). 2. The findings that direct excitation of dopaminergic neurons causes differential degeneration sheds light on the mechanisms of dopaminergic neuron selective vulnerability. The evidence that activation of dopaminergic neurons causes degeneration and alters mRNA expression is convincing, as the authors use both vehicle and CNO control groups, but the evidence that chronic dopaminergic activation alters circadian rhythm and motor behavior is incomplete as the authors did not run a CNO-control condition in these experiments.

      Strengths:

      This is an exciting and important paper.

      The paper compares mouse transcriptomics with human patient data.

      It shows that selective degeneration can occur across the midbrain dopaminergic neurons even in the absence of a genetic, prion, or toxin neurodegeneration mechanism.

      We thank the reviewer for these insightful comments.

      Weaknesses:

      Major concerns:

      (1) The lack of a CNO-positive, DREADD-negative control group in the behavioral experiments is the main limitation in interpreting the behavioral data. Without knowing whether CNO on its own has an impact on circadian rhythm or motor activity, the certainty that dopaminergic hyperactivity is causing these effects is lacking.

      This is an important point. Although we show that CNO does not produce degeneration of DA neuron terminals, we do not exclude a contribution to the behavioral changes. We agree that this behavioral control is necessary, and will address it in revision with a CNO-only running wheel cohort.

      (2) One of the most exciting things about this paper is that the SNc degenerates more strongly than the VTA when both regions are, in theory, excited to the same extent. However, it is not perfectly clear that both regions respond to CNO to the same extent. The electrophysiological data showing CNO responsiveness is only conducted in the SNc. If the VTA response is significantly reduced vs the SNc response, then the selectivity of the SNc degeneration could just be because the SNc was more hyperactive than the VTA. Electrophysiology experiments comparing the VTA and SNc response to CNO could support the idea that the SNc has substantial intrinsic vulnerability factors compared to the VTA.

      We agree that additional electrophysiology conducted in the VTA dopamine neurons would meaningfully add to our understanding of the selective vulnerability in this model, and will complete these experiments in revision.

      (3) The mice have access to a running wheel for the circadian rhythm experiments. Running has been shown to alter the dopaminergic system (Bastioli et al., 2022) and so the authors should clarify whether the histology, electrophysiology, fiber photometry, and transcriptomics data are conducted on mice that have been running or sedentary.

      We will explicitly clarify which mice had access to a running wheel in our revision. Briefly, mice for histology, electrophysiology, and transcriptomics all had access to a running wheel during their treatment. The mice used for photometry underwent about 7 days of running wheel access approximately 3 weeks prior to the beginning of the experiment. The photometry headcaps sterically prevented mice from having access to a running wheel in their home cage.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Rademacher and colleagues examined the effect on the integrity of the dopamine system in mice of chronically stimulating dopamine neurons using a chemogenetic approach. They find that one to two weeks of constant exposure to the chemogenetic activator CNO leads to a decrease in the density of tyrosine hydroxylase staining in striatal brain sections and to a small reduction of the global population of tyrosine hydroxylase positive neurons in the ventral midbrain. They also report alterations in gene expression in both regions using a spatial transcriptomics approach. Globally, the work is well done and valuable and some of the conclusions are interesting. However, the conceptual advance is perhaps a bit limited in the sense that there is extensive previous work in the literature showing that excessive depolarization of multiple types of neurons associated with intracellular calcium elevations promotes neuronal degeneration. The present work adds to this by showing evidence of a similar phenomenon in dopamine neurons.

      We thank the reviewer for the careful and thoughtful review of our manuscript.

      While extensive depolarization and associated intracellular calcium elevations promotes degeneration generally, we emphasize that the process we describe is novel. Indeed, prior studies delivering chronic DREADDs to vulnerable neurons in models of Alzheimer’s disease did not report an increase in neurodegeneration, despite seeing changes in protein aggregation (e.g. Yuan and Grutzendler, J Neurosci 2016, PMID: 26758850; Hussaini et al., PLOS Bio 2020, PMID: 32822389). Further, a critical finding from our study is that in our paradigm, this stressor does not impact all dopamine neurons equally, as the SNc DA neurons are more vulnerable than the VTA, mirroring selective vulnerability characteristic of Parkinson’s disease. This is consistent with a large body of literature that SNc dopamine neurons are less capable of handling large energetic and calcium loads compared to neighboring VTA neurons, and the finding that chronically altered activity is sufficient to drive this preferential loss is novel.

      In addition, we are not aware of prior studies that have chronically activated DREADDs to produce neurodegeneration. Other studies have shown that acute excitotoxic stressors can produce neuronal degeneration, but the chronic increase in activity is central to our approach.

      In terms of the mechanisms explaining the neuronal loss observed after 2 to 4 weeks of chemogenetic activation, it would be important to consider that dopamine neurons are known from a lot of previous literature to undergo a decrease in firing through a depolarization-block mechanism when chronically depolarized. Is it possible that such a phenomenon explains much of the results observed in the present study? It would be important to consider this in the manuscript.

      As discussed in greater detail in the results section below, our data suggests this may not be a prominent feature in our model. However, we cannot rule out a contribution of depolarization block, and will expand on the discussion of this possibility in the revised manuscript.

      The relevance to Parkinson's disease (PD) is also not totally clear because there is not a lot of previous solid evidence showing that the firing of dopamine neurons is increased in PD, either in human subjects or in mouse models of the disease. As such, it is not clear if the present work is really modelling something that could happen in PD in humans.

      We completely agree that evidence of increased dopamine neuron activity from human PD patients is lacking and the existing data are difficult to interpret without human controls. However, as we outline in the manuscript, multiple lines of evidence suggest that the activity level of dopamine neurons almost certainly does change in PD. Therefore, it is very important that we understand how changes in the level of neural activity influence the degeneration of DA neurons. In this paper we examine the impact of increased activity. Increased activity may be compensatory after initial dopamine neuron loss, or may be an initial driver of death (Rademacher & Nakamura, Exp Neurol 2024, PMID: 38092187). Beyond what is already discussed in the manuscript, additional support for increased activity in PD models include:

      - Elevated firing rates in asymptomatic MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488)

      - Increased frequency of spontaneous firing in patient-derived iPSC dopamine neurons and primary mouse dopamine neurons that overexpress synuclein (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060)

      - Increased spontaneous firing in dopamine neurons of rats injected with synuclein preformed fibrils compared to sham (Tozzi et al., Brain 2021, PMID: 34297092)

      We will include and further discuss these important examples in our revision.

      Similarly, in future studies, it will also be important to study the impact of decreasing DA neuron activity. There will be additional levels of complexity to accurately model changes in PD, which may differ between subtypes of the disease, the disease stage, and the subtype of dopamine neuron. Our study models the possibility of chronically increased pacemaking, and interpretation of our results will be informed as we learn more about how the activity of DA neurons changes in humans in PD. We will discuss and elaborate on these important points in the revision.

      Comments on the introduction:

      The introduction cites a 1990 paper from the lab of Anthony Grace as support of the fact that DA neurons increase their firing rate in PD models. However, in this 1990 paper, the authors stated that: "With respect to DA cell activity, depletions of up to 96% of striatal DA did not result in substantial alterations in the proportion of DA neurons active, their mean firing rate, or their firing pattern. Increases in these parameters only occurred when striatal DA depletions exceeded 96%." Such results argue that an increase in firing rate is most likely to be a consequence of the almost complete loss of dopamine neurons rather than an initial driver of neuronal loss. The present introduction would thus benefit from being revised to clarify the overriding hypothesis and rationale in relation to PD and better represent the findings of the paper by Hollerman and Grace.

      We agree that the findings of Hollerman and Grace support compensatory changes in dopamine neuron activity in response to loss of dopamine neurons, rather than informing whether dopamine neuron loss can also be an initial driver of activity. We will clarify this point in our revision. In addition, the results of other studies on this point are mixed: a 50% reduction in dopamine neurons didn’t alter firing rate or bursting (Harden and Grace, J Neurosci 1995, PMID: 7666198; Bilbao et al, Brain Res 2006, PMID: 16574080), while a 40% loss was found to increase firing rate and bursting (Chen et al, Brain Res 2009. PMID: 19545547) and larger reductions alter burst firing (Hollerman & Grace, Brain Res 1990, PMID: 2126975; Stachowiak et al, J Neurosci 1987, PMID: 3110381). Importantly, even if compensatory, such late-stage increases in dopamine neuron activity may contribute to disease progression and drive a vicious cycle of degeneration in surviving neurons. In addition, we also don’t know how the threshold of dopamine neuron loss and altered activity may differ between mice and humans, and PD patients do not present with clinical symptoms until ~30-60% of nigral neurons are lost (Burke & O’Malley, Exp Neurol 2013, PMID: 22285449; Shulman et al, Annu Rev Pathol 2011, PMID: 21034221).

      Other lines of evidence support the potential role of hyperactivity in disease initiation, including increased activity before dopamine neuron loss in MitoPark mice (Good et al., FASEB J 2011, PMID: 21233488), increased spontaneous firing in patient-derived iPSC dopamine neurons (Lin et al., Acta Neuropath Comm 2021, PMID: 34099060), and increased activity observed in genetic models of PD (Bishop et al., J Neurophysiol 2010, PMID: 20926611; Regoni et al., Cell Death Dis 2020,  PMID: 33173027).

      It would be good that the introduction refers to some of the literature on the links between excessive neuronal activity, calcium, and neurodegeneration. There is a large literature on this and referring to it would help frame the work and its novelty in a broader context.

      We agree that a discussion of hyperactivity, calcium, and neurodegeneration would benefit the introduction. While we briefly discuss calcium and neurodegeneration in the discussion, we will expand on this literature in both the introduction and discussion sections. We will carefully review and contextualize our work within existing frameworks of calcium and neurodegeneration (e.g. Surmeier & Schumacker, J Biol Chem 2013, PMID: 23086948; Verma et al., Transl Neurodegener 2022, PMID: 35078537). We believe that the novelty of our study lies in 1) a chronic chemogenetic activation paradigm via drinking water, 2) demonstrating selective vulnerability of dopamine neurons as a result of altering their activity/excitability alone, and 3) comparing mouse and human spatial transcriptomics.

      Comments on the results section:

      The running wheel results of Figure 1 suggest that the CNO treatment caused a brief increase in running on the first day after which there was a strong decrease during the subsequent days in the active phase. This observation is also in line with the appearance of a depolarization block.

      The authors examined many basic electrophysiological parameters of recorded dopamine neurons in acute brain slices. However, it is surprising that they did not report the resting membrane potential, or the input resistance. It would be important that this be added because these two parameters provide key information on the basal excitability of the recorded neurons. They would also allow us to obtain insight into the possibility that the neurons are chronically depolarized and thus in depolarization block.

      We do report the input resistance in Supplemental Figure 1C, which was unchanged in CNO-treated animals compared to controls. We did not report the resting membrane potential because many of the DA neurons were spontaneously firing. However, we will report the initial membrane potential on first breaking into the cell for the whole cell recordings in the revision, which did not vary between groups. This is still influenced by action potential activity, but is the timepoint in the recording least impacted by dialyzing of the neuron by the internal solution. We observed increased spontaneous action potential activity ex vivo in slices from CNO-treated mice (Figure 1D), thus at least under these conditions these dopamine neurons are not in depolarization block. We also did not see strong evidence of changes in other intrinsic properties of the neurons with whole cell recordings (e.g. Figure S1C). Overall, our electrophysiology experiments are not consistent with the depolarization block model, at least not due to changes in the intrinsic properties of the neurons. Although our ex vivo findings cannot exclude a contribution of depolarization block in vivo, we do show that CNO-treated mice removed from their cages for open field testing continue to have a strong trend for increased activity for approximately 10 days (S1E).  This finding is also consistent with increased activity of the DA neurons. We will add discussion of these important considerations in the revision.

      It is great that the authors quantified not only TH levels but also the levels of mCherry, co-expressed with the chemogenetic receptor. This could in principle help to distinguish between TH downregulation and true loss of dopamine neuron cell bodies. However, the approach used here has a major caveat in that the number of mCherry-positive dopamine neurons depends on the proportion of dopamine neurons that were infected and expressed the DREADD and this could very well vary between different mice. It is very unlikely that the virus injection allowed to infect 100% of the neurons in the VTA and SNc. This could for example explain in part the mismatch between the number of VTA dopamine neurons counted in panel 2G when comparing TH and mCherry counts. Also, I see that the mCherry counts were not provided at the 2-week time point. If the mCherry had been expressed genetically by crossing the DAT-Cre mice with a floxed fluorescent reported mice, the interpretation would have been simpler. In this context, I am not convinced of the benefit of the mCherry quantifications. The authors should consider either removing these results from the final manuscript or discussing this important limitation.

      We thank the reviewer for this insightful comment, and we agree that this is a caveat of our mCherry quantification. Quantitation of the number of mCherry+ DA neurons specifically informs the impact on transduced DA neurons, and mCherry appears to be less susceptible to downregulation versus TH. As the reviewer points out, it carries the caveat that there is some variability between injections. Nonetheless, we believe that it conveys useful complementary data. As suggested, we will discuss this caveat in our revision. Note that mCherry was not quantified at the two-week timepoint because there is no loss of TH+ cells at that time.

      Although the authors conclude that there is a global decrease in the number of dopamine neurons after 4 weeks of CNO treatment, the post-hoc tests failed to confirm that the decrease in dopamine number was significant in the SNc, the region most relevant to Parkinson's. This could be due to the fact that only a small number of mice were tested. A "n" of just 4 or 5 mice is very small for a stereological counting experiment. As such, this experiment was clearly underpowered at the statistical level. Also, the choice of the image used to illustrate this in panel 2G should be reconsidered: the image suggests that a very large loss of dopamine neurons occurred in the SNc and this is not what the numbers show. A more representative image should be used.

      We agree that the stereology experiments were performed on relatively small numbers of animals. Combined with the small effect size, this may have contributed to the post-hoc tests showing a trend of p=0.1 for both the TH and mCherry dopamine cell counts in the SN at 4 weeks. As part of the planned experiments for our revision, we will perform an additional stereologic analysis to further assess the loss of SNc dopamine neurons. We will also review and ensure the images are representative.

      In Figure 3, the authors attempt to compare intracellular calcium levels in dopamine neurons using GCaMP6 fluorescence. Because this calcium indicator is not quantitative (unlike ratiometric sensors such as Fura2), it is usually used to quantify relative changes in intracellular calcium. The present use of this probe to compare absolute values is unusual and the validity of this approach is unclear. This limitation needs to be discussed. The authors also need to refer in the text to the difference between panels D and E of this figure. It is surprising that the fluctuations in calcium levels were not quantified. I guess the hypothesis was that there should be more or larger fluctuations in the mice treated with CNO if the CNO treatment led to increased firing. This needs to be clarified.

      We thank the reviewer for this comment. We understand that this method of comparing absolute values is unconventional. However, these animals were tested concurrently on the same system, and a clear effect on the absolute baseline was observed. We will include a caveat of this in our discussion. Panel D of this figure shows the raw, uncorrected photometry traces, whereas panel E shows the isosbestic corrected traces for the same recording. In panel E, the traces follow time in ascending order. We will also include frequency and amplitude data for these recordings.   

      Although the spatial transcriptomic results are intriguing and certainly a great way to start thinking about how the CNO treatment could lead to the loss of dopamine neurons, the presented results, the focusing of some broad classes of differentially expressed genes and on some specific examples, do not really suggest any clear mechanism of neurodegeneration. It would perhaps be useful for the authors to use the obtained data to validate that a state of chronic depolarization was indeed induced by the chronic CNO treatment. Were genes classically linked to increased activity like cfos or bdnf elevated in the SNc or VTA dopamine neurons? In the striatum, the authors report that the levels of DARP32, a gene whose levels are linked to dopamine levels, are unchanged. Does this mean that there were no major changes in dopamine levels in the striatum of these mice?

      We will review the expression of activity-related genes in our dataset, although we must keep in mind that these genes may behave differently in the context of chronic activation as opposed to acutely increased activity. We will also include experiments assessing striatal dopamine levels by HPLC in the revision.

      The usefulness of comparing the transcriptome of human PD SNc or VTA sections to that of the present mouse model should be better explained. In the human tissues, the transcriptome reflects the state of the tissue many years after extensive loss of dopamine neurons. It is expected that there will be few if any SNc neurons left in such sections. In comparison, the mice after 7 days of CNO treatment do not appear to have lost any dopamine neurons. As such, how can the two extremely different conditions be reasonably compared?

      Our mouse model and human PD progress over distinct timescales, as is the case with essentially all mouse models of neurodegenerative diseases. Nonetheless, in our view there is still great value in comparing gene expression changes in mouse models with those in human disease. It seems very likely that the same pathologic processes that drive degeneration early in the disease continue to drive degeneration later in the disease. Note that we have tried to address the discrepancy in time scales in part by comparing to early PD samples when there is more limited SNc DA neuron loss. Please note the numbers of DA neurons within the areas we have selected for sampling (Figure at right). Therefore, we can indeed use spatial transcriptomics to compare dopamine neurons from mice with initial degeneration and patients where degeneration is ongoing during their disease.

      Author response image 1.

      Violin plot of DA neuron proportions sampled within the vulnerable SNV (deconvoluted RCTD method used in unmasked tissue sections of the SNV).

      Control and early PD subjects.

      Comments on the discussion:

      In the discussion, the authors state that their calcium photometry results support a central role of calcium in activity-induced neurodegeneration. This conclusion, although plausible because of the very broad pre-existing literature linking calcium elevation (such as in excitotoxicity) to neuronal loss, should be toned down a bit as no causal relationship was established in the experiments that were carried out in the present study.

      Our model utilizes hM3Dq-DREADDs that function by increasing intracellular calcium to increase neuronal excitability, and our results show increased Ca2+ by fiber photometry and changes to Ca2+-related genes, strongly suggesting a causal relation and crucial role of calcium in the mechanism of degeneration. However, we agree that we have not experimentally proven this point, as we acknowledged in the text. Additionally, we have planned revision experiments involving chronic isradipine treatment to further test the role of calcium in the mechanism of degeneration in this model.

      In the discussion, the authors discuss some of the parallel changes in gene expression detected in the mouse model and in the human tissues. Because few if any dopamine neurons are expected to remain in the SNc of the human tissues used, this sort of comparison has important conceptual limitations and these need to be clearly addressed.

      As discussed, we can sample SN DA neurons in early PD (see figure above), and in our view there is great value for such comparisons. We agree that discussion of appropriate caveats is warranted and this will be clearly addressed in the revision.

      A major limitation of the present discussion is that it does not discuss the possibility that the observed phenotypes are caused by the induction of a chronic state of depolarization block by the chronic CNO treatment. I encourage the authors to consider and discuss this hypothesis.

      As discussed above, our analyses of DA neuron firing in slices and open field testing to date do not support a prominent contribution of depolarization block with chronic CNO treatment. However, we cannot rule out this hypothesis, therefore we will include additional electrophysiology experiments and add discussion of this important consideration.  

      Also, the authors need to discuss the fact that previous work was only able to detect an increase in the firing rate of dopamine neurons after more than 95% loss of dopamine neurons. As such, the authors need to clearly discuss the relevance of the present model to PD. Are changes in firing rate a driver of neuronal loss in PD, as the authors try to make the case here, or are such changes only a secondary consequence of extensive neuronal loss (for example because a major loss of dopamine would lead to reduced D2 autoreceptor activation in the remaining neurons, and to reduced autoreceptor-mediated negative feedback on firing). This needs to be discussed.

      As discussed above, while increases in dopamine neuron activity may be compensatory after loss of neurons, the precise percentage required to induce such compensatory changes is not defined in mice and varies between paradigms, and the threshold level is not known in humans. We also reiterate that a compensatory increase in activity could still promote the degeneration of critical surviving DA neurons, whose loss underlies the substantial decline in motor function that typically occurs over the course of PD. Moreover, there are also multiple lines of evidence to suggest that changes in activity can initiate and drive dopamine neuron degeneration (Rademacher & Nakamura, Exp Neurol 2024). For example, overexpression of synuclein can increase firing in cultured dopamine neurons (Dagra et al., NPJ Parkinsons Dis 2021, PMID: 34408150) while mice expressing mutant Parkin have higher mean firing rates (Regoni et al., Cell Death Dis 2020,  PMID: 33173027). Similarly, an increased firing rate has been reported in the MitoPark mouse model of PD at a time preceding DA neuron degeneration (Good et al., FASEB J 2011, PMID: 21233488). We also acknowledge that alterations to dopamine neuron activity are likely complex in PD, and that dopamine neuron health and function can be impacted not just by simple increases in activity, but also by changes in activity patterns and regularity. We will amend our discussion to include the important caveat of changes in activity occurring as compensation, as well as further evidence of changes in activity preceding dopamine neuron death.

      There is a very large, multi-decade literature on calcium elevation and its effects on neuronal loss in many different types of neurons. The authors should discuss their findings in this context and refer to some of this previous work. In a nutshell, the observations of the present manuscript could be summarized by stating that the chronic membrane depolarization induced by the CNO treatment is likely to induce a chronic elevation of intracellular calcium and this is then likely to activate some of the well-known calcium-dependent cell death mechanisms. Whether such cell death is linked in any way to PD is not really demonstrated by the present results. The authors are encouraged to perform a thorough revision of the discussion to address all of these issues, discuss the major limitations of the present model, and refer to the broad pre-existing literature linking membrane depolarization, calcium, and neuronal loss in many neuronal cell types.

      While our model demonstrates classic excitotoxic cell death pathways, we would like to emphasize both the chronic nature of our manipulation and the progressive changes observed, with increasing degeneration seen at 1, 2, and 4 weeks of hyperactivity in an axon-first manner. This is a unique aspect of our study, in contrast to much of the previous literature which has focused on shorter timescales. Thus, while we will revise the discussion to more comprehensively acknowledge previous studies of calcium-dependent neuron cell death, we believe we have made several new contributions that are not predicted by existing literature. We have shown that this chronic manipulation is specifically toxic to nigral dopamine neurons, and the data that VTA dopamine neurons continue to be resilient even at 4 weeks is interesting and disease-relevant. We therefore do not want to use findings from other neuron types to draw assumptions about DA neurons, which are a unique and very diverse population. We acknowledge that as with all preclinical models of PD, we cannot draw definitive conclusions about PD with this data. However, we reiterate that we strongly believe that drawing connections to human disease is important, as dopamine neuron activity is very likely altered in PD and a clearer understanding of how dopamine neuron survival is impacted by activity will provide insight into the mechanisms of PD.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents a potentially important integrative model linking spontaneous retinal waves, apoptosis, microglial activity, and vascular development during postnatal retinal maturation. Its significance lies in proposing a mechanistic framework that could reshape understanding of how neural activity and tissue remodeling are coordinated in the developing central nervous system. The evidence is strengthened by the use of multiple complementary techniques, including Ca++ imaging, high-throughput electrophysiology, transcriptomics, histology, and pharmacology.

      Strengths:

      (1) Multimodal Validation: The authors correlate large-scale functional imaging (calcium imaging and MEA) with high-resolution structural and molecular data (scRNA-seq and IHC), providing strong topographical evidence for the "centrifugal expansion" pattern.

      (2) The primary significance lies in identifying apoptotic Retinal Ganglion Cells (RGCs) as the physiological "pacemakers" for stage II retinal waves. By linking programmed cell death directly to neural activity and subsequent angiogenesis, the authors propose a self-regulating developmental loop.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      (1) While the PANX1 pharmacological data provide compelling functional support, extending these conclusions to the broader CNS may be premature. Additional direct mechanistic validation would further strengthen the claim of causality.

      We agree with the reviewer that the conclusions would be greatly solidified with more direct mechanistic validation. However, we are unable to conduct more experimentation as the grant is finished and the Sernagor lab is in the process of being shutdown, after the unexpected passing of the PI.

      In order to make clearer that this mechanism was found in retinal tissue, not CNS, we have moved any mention of the implications of our work to a broader CNS mechanism to the discussion section. We have also added text into the discussion highlighting the need for more mechanistic investigation to uncover the full extent of the developmental processes described herein, see Line 413.

      (2) While the manuscript beautifully illustrates the co-occurrence of events during retinal development, strengthening the distinction between correlation and direct causation would enhance the impact of the findings.

      We have been clear to only present our findings as correlational as we were unable to fully explore the causational nature within the mechanisms presented. In the discussion, we have used published evidence and experimental papers to bolster our understanding of the causal aspects of this research. We have also included sections of text to address what experimentation is be required to examine the causal interactions more directly, see Line 413.

      Reviewer #2 (Public review):

      Summary:

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Using scRNAseq, the authors classify autofluorescence cell clusters (ACCs) at the leading edge of vasculature outgrowth as Hmox-1+ microglia. From here, they show microglia engulfment of apoptotic RGCs, and the potential release of ATP may contribute to Ca2+ wave generation. The authors demonstrate these mechanisms through the use of two pharmacological agents to either block the ATP release from Panx-1 or block receptor binding to ATP. Furthermore, while previous studies have described the site of initiation of retinal Ca2+ waves as random, this study shows that the initiation of Ca2+ waves is biased to the leading edge of vascular growth in the developing retina. To do this, the authors use a combination of wide-field Ca2+ imaging and multi-electrode arrays to pinpoint the sites of Ca2+ wave initiation in the developing retina.

      Strengths:

      The authors use several techniques to interrogate these mechanisms, including single-cell RNAseq, wide-field Ca2+ imaging, and multi-electrode arrays. With these experiments, this manuscript proposes several novel ideas, such as ATP as the Ca2+ wave-initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth.

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Weaknesses:

      The main weakness of the manuscript is the overreliance on only two pharmacological agents to test the central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP release (i.e., BAX-KO, Panx-1 KO).

      We thank the reviewer for their insightful suggestions for further experimentation to bolster the research. Initially, we utilised pharmacological interventions as they provided acute and quick answering of the research question. At the outset of the research, we were not certain that purinergic release through PANX-1 channels was the mediator for the developmental mechanisms described. We tested a wide variety of specific agonists and blockers before seeing any profound effects on wave generation. These agonists and antagonists have been used before and are proven to deliver reliable results. In addition, since the ACCs had never been reported before we were unsure if a knockout animal would display the same anatomical phenotype. Furthermore, it is known that knockout mouse lines, especially connexin and hemichannel pores, do not lose function but rather have other isoforms or compensation mechanisms which can substitute the original function. For the retina, for example, it was shown that Cx36 can functionally replace Cx45 after Cx45 KO (Frank et al, 2010).

      We agree that while direct mechanistic validation would significantly reinforce the arguments, we are limited in conducting further experiments since the grant has been completed and the Sernagor lab is in the process of shutting down following her passing.

      In order to address the omission of mechanistic validation in the paper we have added text into the discussion highlighting the need deeper investigation in the causality of the developmental processes described herein, see Line 413.

      M. Frank et al., Neuronal connexin-36 can functionally replace connexin-45 in mouse retina but not in the developing heart, J. Cell Sci. 123, 3605 (2010).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      General and major comments

      (A) Introduction

      (1) The introduction is currently quite extensive. I recommend streamlining the background information to more directly frame the study's core objectives.

      We have reduced the background information contained in the introduction to better align with the direct outputs of the study. However, as this research paper examines the interactions of multiple complex developmental processes a relatively in-depth introduction is needed to inform the reader of the salient points.

      (2) To improve clarity, it would be highly beneficial to conclude the introduction with a sequential summary of key observations in the order they are presented in the study. This would provide a clearer roadmap for the reader.

      We have reworked the end of the introduction to better align with the key observations in the order they are presented through the figures.

      (B) Results

      (3) The characterization of ACCs (Apoptotic Cell Clusters) would be more effective if separated from the description of their spatiotemporal occurrence with SVPs. Consider moving the microscopic description to a dedicated section or merging it with the subsequent paragraph on molecular identification.

      We thank the reviewer for their suggestion to reorganise the results section. However, we believe that highlighting the integrated nature of the ACC positioning and development of the vascular plexus is an important stepping off point for the reader. It allows us to highlight the original serendipitous discovery of the ACCs and their highly cohesive role in bridging multiple developmental processes.

      (4) Please explicitly state the specific retinal developmental stages (e.g., P3-P6) within the section regarding RNA sequencing, as this timing is critical for interpreting the transcriptomic data.

      We have amended the text to specify the postnatal days used for RNA sequencing, see Line 123

      (5) Regarding the scRNA-Seq data: were ACC-negative samples (isolated via FACS) also processed? A direct comparison between ACC+ and ACC- sampled microglia would significantly strengthen the claim that microglia are specifically attracted to ACCs. If these data are available, they would make an elegant and compelling addition to the manuscript.

      We thank the reviewer for this important suggestion and agree that direct comparison of ACC+ and ACC− microglia would further strengthen the study. Unfortunately, ACC− populations were not processed for scRNA-seq in the current study because of the prioritisation of the rare ACC+ population. Nevertheless, several independent observations support the conclusion that microglia are preferentially associated with ACCs, including: (i) the enrichment of microglia within ACC-containing regions observed histologically, (ii) the spatial proximity analyses shown in Figure 3, and (iii) the distinct transcriptional profile of ACC-associated microglia identified by scRNA-seq.

      We have also added a section to the discussion to highlight the need for a direct comparison of the ACC+ and ACC- transcriptomics profiles in future work, see Line 344.

      (6) For the Ca2+ -imaging experiments, please briefly describe the staining protocol and specify which cell types (e.g., RGCs) were labeled within the Results text to assist the reader's immediate understanding.

      We have added a short description of the labelling technique in the results, see Line 228

      (7) The manuscript notes that MEA waves are evident in the graphs of Figure 7, but the raw wave data or representative traces are not shown. Including these (similar to those of the imaging waves) would provide necessary visual verification of the physiological phenomena described.

      We have added a supplementary figure 3 which details stage 2 retinal waves recorded using MEAs.

      (C) Interpretations and Logic

      (8) The finding ' ...Wholemount staining revealed a broad centro-peripheral gradient of apoptosis; however, this apoptotic annulus was positioned more peripherally than the ACCs, SVP, and Hmox1-positive microglia (Figure 4B)' seems to contradict the HMOX1/Yo-Pro-1 stained microglia. It is not clear whether the authors aim to prove that this particular set of microglia phagocytise RGCs, or another set that lines up better with dying cells and does not show up on the HMOX-1 label. I believe the authors intend to show the time difference of the two events - cells dying and HMOX-1 microglia appear at the site later. I believe the logic is good; it may need a sentence pointing this out at the end of this paragraph.

      We have added a statement in Line 182 which clarifies our intent to show that the Hmox1 microglia phagocytose the dying RGCs after they initiate apoptotic mechanisms.

      (9) If the authors intend to demonstrate a temporal lag between cell death and the appearance of HMOX1+ microglia as evidence of causality, a concluding sentence to this effect would greatly clarify the logic of this paragraph.

      We have added a concluding sentence to the paragraph in Line 190 which indicates a causal link between RGC cell death and appearance of hmox1 positive microglia.

      (D) Figures and Presentation

      (10) The blood vessel staining in the final panel of Figure 1B is currently quite faint. Increasing the brightness/contrast for this panel would allow the reader to better appreciate the underlying architecture.

      We have updated the panel in Figure 1B to match the brightness of the others of that series.

      (11) Given that the peripherality of events in Figure 7 suggests a specific sequence, the authors should consider adding a summary timeline (P3-P6). A plot using curves (mean or median values), color-coded to match the corresponding events, would provide a much-needed visual synthesis of the data.

      We agree with reviewer that the D1/2 metrics would benefit from more clarification to show the timeline of development more clearly. We have added another panel to Figure 7, which shows mean/standard deviation plots for each developmental measure using D1/2 as timelines. This allows the reader to better compare the progression of centrifugal spread more clearly.

      (12) Please ensure that graph labels and axis titles are uniform in size across all figures. e.g., the labels in Figure 6 and several other graphs are currently too small to be legible in the PDF; these should be enlarged for better accessibility.

      We have fixed the labels and axis titles to maintain readability across the paper

      Minor comments

      (1) In line 125, there is a missing closing parenthesis after the reference to Figure 1C.

      We have fixed this error

      (2) The specific algorithm used for the unsupervised cluster analysis has not been identified in the text. Please specify whether k-means, Louvain, or another method was employed to ensure reproducibility.

      We have reworked the section detailing the cluster analysis to make clear we used Louvain-based clustering, see line 475.

      (3) While it is appropriate to leave comprehensive technical details for the Methods section, a brief conceptual explanation of the D1, D2, and D3 metrics should be included in the Results text to aid general comprehension.

      We have added a brief description of the D1/2 and D1/3 metrics when they are first mentioned in the results section, see line 243.

      (4) In the Figure 7 schematic, the representation of the starburst amacrine cell should be revised to more accurately reflect its well-characterized morphology (e.g., thin primary and gradually thickening higher-order dendrites).

      We have changed the SAC representation to better match the characteristics of that cell type.

      (5) Throughout the manuscript (e.g., in lines 293-294), it is claimed that RGC apoptosis promotes the expression of PANX-1 hemichannels. While the data effectively demonstrate the release of purinergic molecules (e.g., ATP) via PANX-1 from dying cells, the evidence for an actual upregulation or increase in PANX-1 protein/mRNA levels is not explicitly shown. Please clarify whether the findings suggest increased activity of existing channels or a true increase in expression. If the latter is not empirically supported, the phrasing should be adjusted to reflect functional activation rather than de novo expression.

      We agree with the author that our research shows a functional increase in PANX-1 and we have adjusted the language to match. In the introduction and discussion, we provide published evidence and experimental papers which describe the upregulation of the PANX-1 molecule in dying RGCs.

      Reviewer #2 (Recommendations for the authors):

      Savage et al. investigate the synchronization of retinal Ca2+ waves with developmental cell death, microglia activation, and vascular outgrowth. These developmental processes occur through a mechanism where apoptotic cells release ATP through Panx-1 channels to stimulate both Ca2+ retinal waves and microglia activation. Furthermore, the authors demonstrate the initiation of Ca2+ waves occurs at the leading edge of vascular growth in the developing retina. This manuscript proposes several novel ideas, such as ATP as the Ca2+ wave initiating cue, and the localization of the Ca2+ wave initiation to the leading edge of vascular growth. The main weakness of the manuscript is the overreliance on only two pharmacological agents to test their central hypotheses. These conclusions would be strengthened if, in addition to their pharmacological manipulations, they used genetic knockout models to perturb programmed cell death or ATP releases (i.e., BAX-KO, Panx-1 KO). In addition, the following comments should also be addressed:

      We thank the reviewer for their nice summary and for highlighting the strengths of this work.

      Specific comments:

      (1) Line 128: Why would ACCs be involved in SVP guidance if they are trailing the leading edge of the vasculature? Would it be the other way around, where the leading edge of the vasculature would be trailing the ACCs?

      At this point in the paper we are suggesting that the highly stereotyped position of the ACCs under the leading edge of the SVP indicates that they have a mechanistic involvement in SVP growth. Not that they are the direct cause of the expansion. As the paper progresses, we make clear that contrary to our original hypotheses which state the ACCs may cause or control the integrated development of the retina, they are a hallmark of the apoptotic RGCs in the periphery which are the chemogenic beacons for vascular growth being ‘decommissioned’ by the microglia which fine-tune vascular growth and create the ACCs.

      (2) Line 137: Please state in the text and figure legend, at what age ACCs were isolated from the retina.

      We have added the relevant information to Line 123

      (3) Line 154: Why didn't RGCs form their own cluster? Why do the RBPMS+ cells appear across the entire dataset (Figure 2G)? Have previous investigations also shown engulfed cell transcriptomes appearing in the microglia clusters using scRNAseq?

      Previous transcriptomic studies have demonstrated that phagocytic microglia can encapsulate transcripts originating from neurons and other neural cell types, which are detectable by RNA‑seq despite not belonging to a common microglial genetic signature. For example, Solga et al. showed that CNS microglia contain neuronal and oligodendrocyte‑specific mRNAs that localise within microglia but are not translated. This research group interpreted that this RNA is acquired through phagocytosis or macropinocytosis of surrounding neural cells (Solga et al., 2015). Similarly, in zebrafish, synapse‑engulfing microglia identified in situ display neuronal and synaptic gene expression in single‑cell RNA‑seq profiles, consistent with engulfed neuronal material contributing to the detected transcriptome (Sliva et al., 2021). In line with these observations, and given our FACS strategy enriching autofluorescent ACCs rather than intact RGCs, we interpret the widespread Rbpms expression across ACC‑associated clusters as an expected consequence of microglial engulfment of apoptotic RGCs, rather than evidence for a distinct population of viable RGCs that failed to form a separate cluster.

      Solga, A.C., Pong, W.W., Walker, J., Wylie, T., Magrini, V., Apicelli, A.J., Griffith, M., Griffith, O.L., Kohsaka, S., Wu, G.F. and Brody, D.L., 2015. RNA‐sequencing reveals oligodendrocyte and neuronal transcripts in microglia relevant to central nervous system disease. Glia, 63(4), pp.531-548.

      Silva, N.J., Dorman, L.C., Vainchtein, I.D., Horneck, N.C. and Molofsky, A.V., 2021. In situ and transcriptomic identification of microglia in synapse-rich regions of the developing zebrafish brain. Nature communications, 12(1), p.5916.

      We have incorporated this information into the discussion in Line 329

      (4) Figure 4C-H: The data in the figure would be strengthened if the authors added quantification for their co-localization images.

      We thank the reviewer for this important suggestion and agree that quantification of the co-localisation would further strengthen the study. Unfortunately, we are limited in conducting further experiments since the grant has been completed and the Sernagor lab is in the process of shutting down following her passing.

      (5) Figure 5A-I: The error bars are quite large. The figure legend says they represent SEM, but are you sure they don't represent standard deviation (SD)?

      The large SEM bars reflect substantial biological variability across retinas and across individual microglia, which is expected for morphometric measures such as circularity, perimeter, branch number and total skeleton length, as well as for counts of rare cell populations (Hmox1+ microglia, YO‑PRO‑1+ cells and double‑positive cells). Importantly, despite this variability, the effects of probenecid and PSB‑0739 on microglial morphology and on the frequencies of apoptotic and double‑positive cells remain statistically robust in our non‑parametric ANOVA and post‑hoc tests, as indicated by the reported P‑values.”

      For panels where the distributions were clearly non‑Gaussian, we used non‑parametric statistics (Kruskal‑Wallis ANOVA), reporting medians and 95% confidence intervals, and we retained SEM in Figure 5 for consistency with the original plotting routine while clarifying this choice in the legend and Methods.

      (6) Figure 5: What are these measurements made at? Please add the age to the results section and the figure legend.

      We have added the postnatal day of the animals used.

      (7) Line 236-239: There is no mention of the use of Probenecid in the text and in Figure 6B. This treatment of Ca2+ waves should be mentioned before line 245.

      We have added a brief description of probenecid application in Line 227.

      (8) Figure 6: The order of this figure and the results section may be clearer if panels C-F were switched with panels G-J.

      We thank the reviewer for their suggests to improve the flow of the results section and figure 6. We have swapped the panels as suggested and amended the text to fit the new flow.

      (9) For the discussion section: Why does probenecid only affect the wave initiation at the P3 timepoint, but not later timepoints, P4-P6.

      We have added some text in the discussion to explain our findings of differential effects of PANX-1 blockade across the P3-6 timeline. See Line 404.

      Editorial revisions:

      (1) Line 92-96: Citation needed.

      We have added appropriate citations for this section.

      (2) Figure 1F: Please consider changing the color scheme from green/red to green/magenta for colorblind readers.

      We have changed the image LUT

      (3) Figure 3A: This figure may benefit from separating out each channel separately (Iba1 and ACC) and then having a "merge" panel.

      We have separated out the panels in Figure 3A

      (4) Figure 4E, H: There is no label for the immunomarkers in Figure 4H. There is no inset in Figure 4E as mentioned in the figure legend. However, it appears that Figure 4H is a magnified image of Figure 4G and not Figure 4E.

      We have fixed the error in the figure legend and included an inset indicator in Figure 4G. We have added immunomarkers to Figure 4H

      (5) Line 198: "SAC" should be "SACs".

      We have fixed this error

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work addresses a question of practical importance that had never been systematically analysed in the cryo-ET field: when collecting tilt-series data, what is the optimal angular step size between successive tilt images? Due to the upper limit in electron exposure (100 - 150 e<sup>-</sup>/Å<sup>2</sup>), this question is important, since finer angular sampling improves attainable reconstruction resolution (Crowther criterion) but reduces the signal-to-noise ratio of each individual image, potentially compromising both image quality and the ability to computationally align successive frames. To address this, the authors designed a thorough benchmarking study comparing five tilt increments (1°, 2°, 3°, 5°, and 10°) while keeping the total dose and tilt range constant. They evaluated the consequences at every stage of the cryo-ET workflow - from raw image quality and tilt-series alignment, through template matching for ribosome detection, to high-resolution subtomogram averaging - with the goal of providing the community with an evidence-based recommendation for data acquisition.

      The manuscript is well written, and the experimental design is carefully thought out. The work provides valuable practical insights into cryo-ET data acquisition by demonstrating that balancing two competing demands - sufficient dose per individual tilt image and fine angular sampling - is essential to achieve high-quality tomographic reconstructions. The identification of a practical optimum at 3° tilt increment is the key contribution of the work. It will be interesting to see in the future whether this optimum shifts for smaller molecular targets, and how emerging tilt interpolation strategies such as cryoTIGER may interact with the choice of experimental angular increment.

      The conclusions of this paper are mostly well supported by data, but some aspects of data analysis need to be clarified and/or extended, including:

      (1) Line 109: The authors state that the tilt range was kept at ± 60° relative to the lamella plane. Assuming a typical lamella pre-tilt of ~10°, the absolute stage tilt would approach its mechanical limit. Two clarifications would be appreciated: (a) What was the average pre-tilt across all lamellae? (b) How many dark tilt images, if any, were excluded during tomogram reconstruction?

      We thank the reviewer for asking for further clarification. For all our datasets, the pre-tilt of the stage was +8° with the lamella untilted under the e-beam, resulting in a tilt range of -52° to + 68°, thereby not reaching the mechanical limit, which is 70° for our microscope stage.

      Regarding “dark tilt images”, for most datasets, we did not need to remove many tilt images. However, we now noticed notably more absence of images from higher tilt values for the 1° dataset (see SFig 1). When analysing further, we noticed that for this dataset, we did not actively remove many images prior to tomogram reconstruction, but rather that they were not acquired in the first place by SerialEM. During acquisition, SerialEM performs various safeguarding checks that can abort the acquisition of a tilt series (or of a single branch). As this seems predominantly a problem for the 1-degree tilt-increment dataset, we have decided to add this to the manuscript as follows, including the figure as new SFig 1.

      In the main text:

      “For most datasets, image acquisition was largely complete, with the exception of the 1-degree dataset, which showed a markedly higher proportion of missing images at high tilt angles (SFig. 1). Closer inspection revealed that many of these images were not acquired, as SerialEM applies built-in safeguards (e.g. autofocus inconsistency or insufficient image counts) that can abort a tilt-series branch before completion.”

      (2) Line 148: "When analysing tomographic volumes, we found that tomograms from data with a smaller increment displayed higher SNR values (see Fig. 2B)." It would be helpful to specify which comparisons are statistically meaningful (e.g. Mann-Whitney U test?). While the difference between 1° and 2° appears pronounced, the differences between 2°, 3°, and 5° seem minimal. From my point of view, reporting the mean SNR values +/- standard deviations for each condition would already indicate some significance. Furthermore, since SNR is expected to depend on lamella thickness, it should be clarified whether the average lamella thickness is comparable across the five datasets.

      We have now calculated the mean and standard deviation of the tomogram SNR, as follows:

      Author response table 1.

      Furthermore, we performed a statistical significance test. Kruskal-Wallis test confirmed significant differences in SNR across tilt increments (H=270.97, p<0.001). Pairwise Mann-Whitney U tests with Bonferroni correction revealed significant differences between all pairs except 2° and 3° (p=0.093), suggesting these two conditions indeed yield comparable SNR.

      Lastly, we have now added data regarding local lamella thickness for all tilt-series used in the study, as displayed in SFig. 4.

      We incorporated this in the manuscript as follows.

      In the main text:

      “When analysing tomographic volumes, we found that tomograms from data with a smaller increment displayed higher SNR values (see Fig. 2B and Supplementary Note), whilst showing a similar lamella thickness distribution (see SFig. 4A).”

      and:

      “Firstly, we selected ca. 20 tomograms per condition, based on tomogram content and local lamella thickness [31] (for more details, see Methods and SFig. 4B).”

      As a supplementary note:

      “As the tomogram SNR distribution of particularly the 2° and 3° dataset showed similar SNR distributions, we performed formal significance testing for the data in this panel (see Fig. 2B). Kruskal-Wallis test confirmed significant differences across conditions (H=270.97, p<0.001); pairwise Mann-Whitney U tests with Bonferroni correction revealed all pairs were significantly different except 2° vs. 3° (p=0.093), indicating comparable SNR for these two tilt increments.”

      And, adding the test in the Methods:

      “To quantify differences in signal-to-noise ratio (SNR) across tilt increment conditions, a non-parametric Kruskal-Wallis test was performed as an omnibus test of the null hypothesis that all groups are drawn from the same distribution. Because SNR distributions were not assumed to be normal, and sample sizes differed across conditions, non-parametric tests were used throughout. Following the omnibus test, all 10 pairwise comparisons between conditions were assessed using two-sided Mann-Whitney U tests. To control for multiple comparisons, raw p-values were adjusted using the Bonferroni correction (multiplied by the number of comparisons, n=10, capped at 1.0). Statistical significance was defined as a Bonferroni-corrected p-value below 0.05. All analyses were performed in Python using the scipy.stats module.”

      (3) Line 167: "Indeed, the variation in maximum resolution correlates with lamella thickness across all datasets (see Fig. 2F)." The reported R<sup>2</sup> values of 0.30 (1°), 0.38 (2°), 0.66 (3°), 0.61 (5°), and 0.60 (10°) reveal a notably weak linear relationship for the finer tilt increments. It is also difficult to assess whether the lamella thickness distributions are comparable across conditions from the current figures - visually, the 1° dataset appears to be based on thinner lamellae, while the 10° dataset appears to include thicker samples. A histogram of lamella thickness distributions for each condition, provided as supplementary material, would greatly aid interpretation. Given this thickness dependency, reporting mean +/- standard deviation of lamella thickness per condition is highly appreciated.

      We have added the full lamella thickness distribution per dataset now in SFig. 4.

      The apparent weaker relationship between resolution fit and local lamella thickness for the 1 dataset seems to be largely apparent to the few very thin data points in this data (for more clarity, see the same data plotted separately in Author response image 1). We speculate that this is due to even less signal in these very thin and very low-dose images.

      Author response image 1.

      (4) Figure 4: It should be specified which tomogram subsets were used for the Rosenthal-Henderson analysis, whether lamella thickness was taken into account in the subset selection, and whether ribosomes too close to the lamella edges were excluded. Finally, linear fits should be displayed across the full x-axis range for all tilt increments to facilitate direct visual comparison.

      We have described the process of tomogram subset selection in detail in the Methods section Template matching and 3D classification. To further add clarity, we have incorporated the local lamella distribution plots for the full data, as well as specifically for the tomograms subjected to TM and STA in SFig. 4.

      Regarding the linear fits, we respectfully disagree with this suggestion. Displaying the linear fits only over the range used for their calculation avoids implying that the linear relationship extends beyond the measured data, and in our view produces a clearer figure.

      (5) General: Were ribosomes located at the lamella edges excluded from the analysis? As demonstrated in the authors' own prior work (Tuijtel et al., Science Advances, 2024), Ga-FIB milling induces structural damage at the lamella surfaces. To exclude the influence on the STA results, particles near the lamella edges should be removed prior to analysis, and the criteria for this exclusion should be stated explicitly.

      We have not excluded any ribosomes from close to the surface. As we still treated all data the same for each condition, we anticipate that the results of the comparison reported here still hold true.

      The aim of the authors was to provide the cryo-ET community with an evidence-based recommendation for the choice of tilt increment, and they largely succeeded in this goal. The identification of 3° as a practical optimum - balancing sufficient dose per tilt image for effective per-particle refinement with fine enough angular sampling for accurate tilt-series alignment - is well supported by the data and consistent across the multiple quality metrics employed. The conclusion that coarser increments (5° and 10°) compromise tomogram quality, template matching accuracy, and STA resolution is robust and clearly demonstrated. However, the conclusion rests entirely on a single biological system using ribosomes as the sole molecular target, which are exceptionally favourable due to their abundance, size, and electron contrast. Whether the identified optimum holds for smaller, lower-abundance, or lower-contrast targets remains an open question.

      In future, it would be particularly interesting to test whether emerging tilt interpolation strategies, such as cryoTIGER, which is particularly intriguing, can effectively compensate for coarser experimental angular sampling in post-processing. Here, the optimal experimental increment may shift, and the interaction between these two approaches represents a promising direction for future work. More broadly, as cryo-ET datasets grow larger and public repositories expand, the practical tradeoffs between acquisition time, data storage, and structural quality identified here will become increasingly relevant to the field.

      We agree with the reviewer and thank them for this positive assessment. An interesting note to the use of cryoTIGER in particular is that it uses already aligned tilt-series as an input, and it therefore is unlikely to overcome severe alignment issues associated with large tilt-increments.

      Reviewer #2 (Public review):

      The determination of macromolecular structures directly within their native cellular environment is becoming increasingly routine, making standardized data collection strategies essential. In this manuscript, Tuijtel et al. provide a timely and valuable contribution by benchmarking key acquisition parameters and establishing practical guidelines for in situ cryo-electron tomography (cryo-ET). Critically, the authors present a systematic framework for optimizing data collection to achieve the highest attainable resolution.

      Using Dictyostelium cells as a model system, the authors generate multiple datasets at a constant total dose while varying the tilt increment. They demonstrate that tilt-series acquired with finer increments (1-3 degrees) yield superior alignment accuracy and improved template-matching performance, resulting in higher-quality reconstructions than those collected with coarser increments (5 degrees or above). Furthermore, the authors show that for subtomogram averaging, a 3-degree tilt increment outperforms all other conditions tested, particularly after per-particle refinement as implemented in M.

      Overall, the manuscript is clearly written, and the conclusions are well supported by the data presented. I have no major concerns. There are some minor points that the authors should address, including:

      (1) The phrase "electron optical density distribution" (line 31, Introduction) should be revised to "electrostatic potential" or "Coulomb potential distribution," which more accurately reflects what is measured in cryo-EM/ET.

      We thank the reviewer for this correction and have adjusted it in the text:

      “It captures the 3-dimensional (3D) electrostatic potential of the specimen under scrutiny and enables the structural analysis of macromolecular complexes within their native context.”

      (2) The authors state that the maximum tolerable electron dose is approximately 100-150 e<sup>-</sup>/Å<sup>2</sup> (line 34, Introduction). This is an oversimplification, as bacterial specimens, for example, have been shown to tolerate doses of 200 e<sup>-</sup>/Å<sup>2</sup> or higher (see Breigel et al., PNAS, 2009; https://www.pnas.org/doi/10.1073/pnas.0905181106#T1). The statement should be revised to reflect this variability.

      We adjusted this statement to now read:

      “One of these is the maximum electron dose that can be applied to biological specimens before irreversible damage occurs, which is about 100-150 e-/Å2 for most eukaryotic cells.”

      (3) Lines 56-57: The authors do not cite their own prior work benchmarking tilt-series acquisition strategies on in vitro samples. This earlier study provides important context and should be referenced and briefly discussed.

      We assume the reviewer meant this study: Turonova et al., Nat. Comms. (2020). We have now added and discussed this reference as follows:

      “Since accumulated radiation dose progressively degrades high-resolution information, this motivated the development of the dose-symmetric tilt scheme, which prioritizes acquisition of low-tilt images early to better preserve high-resolution information [11, 16].”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 159: "Surprisingly though, the resolution to which the CTF was fitted was similar for all conditions, despite an 8-fold increase in dose (see Fig. 2B, E)." The reference to Figure 2B at this point is unclear.

      We thank the reviewer for pointing this out, we have removed the reference to panel B.

      (2) Line 162: "As the data shown in Fig. 2D-F pertains to images of the untilted specimen, ..." For clarity, this should also be stated explicitly in the figure caption.

      We have added this to the figure legend; “both estimated with Gctf on projection images of the sample at effective zero-tilt position.”

      (3) Figure 3: These are compelling results, and the 3D classification outcomes provide an excellent visual representation of the quantitative data shown in Figure 3. Including (some or all) of the initial five classes in the figure would further strengthen this already convincing presentation. Additionally, applying a uniform extraction threshold (e.g., z-score of 3.75 or 5) across all tilt increments would facilitate a more direct comparison. But this is really a minor remark, the authors and the editors may judge if the current presentation is already sufficient.

      We have adjusted Figure 3 according to the reviewer’s recommendation:

      We referenced this in the main text as:

      “To ensure similar data processing strategies for all conditions, extraction thresholds were also lowered for the 1-, 2- and 3-degree conditions, and 3D classification was performed to filter out the junk particles (see Fig. 3 A and SFig. 8). “

      In order to directly compare the particle extraction, we have already carried out such a uniform extraction threshold, with a z-score threshold of 5 (apart for the 10-degree data, where this was not possible). This extraction was then used for the 3D classification that led to the TM analysis and further STA investigations.

      Reviewer #2 (Recommendations for the authors):

      (1) Supplementary Figure 1: To improve accessibility for a broader readership, the authors should annotate or highlight the key organelles and protein complexes visible in the tomographic slices.

      We thank the reviewer for this suggestion, but adding arrowheads made this figure too crowded in our opinion. Furthermore, for most of the panels, the mentioned features of interest is centred in the image panel, which should make identification straightforward.

      (2) 'In situ' should be in italics throughout the text.

      We have changed this.

    1. Author response:

      We sincerely thank the reviewers for their time, helpful critiques and overall positive evaluation of our work. We plan to investigate the points raised by the reviewers and look forward to submitting a revised manuscript that addresses concerns about therapeutic relevance, functional distinctions between the CI-intact and CI-suppressed adapted states, and PC regulation in this system.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This study describes motor cortical activity patterns during food handling in mice, investigating whether the hand/s used is reflected in distinct neural activity. The experiments focus on forelimb M1 and M2 (fM1, fM2) and an oral-manual region LOM. The main findings are that fM1 and fM2 have largely similar relationships with forelimb control, and LOM neurons are more broadly tuned. These conclusions are reached using a variety of analyses spanning straightforward firing rate analyses, selectivity metrics, PCA, and GLM decoding methods to assess tuning generalizability. The study's significance is strengthened by including analyses of bimanual control, and in this sphere, there are descriptive data and analyses that aficionados of cortical control of dexterous behaviors will find useful. The use of unimanual control is useful as a point of comparison here, but less novel overall. There are a number of places where the descriptions of what is being analyzed, what is being concluded, and data reporting should be strengthened and clarified. Additionally, the study could be greatly improved by consolidating figures and the analyses shown, since many are redundant. Many of the analyses need clearer reporting of means and effect sizes in the text, rather than just statistical outcomes. Overall, at this juncture, the study presents analyses of a unique dataset that may seed future investigations of mechanisms of bimanual coordination.

      Strengths:

      There are relatively few studies that compare neural activity across bimanual and unimanual control. This study uses a naturalistic food handling task to explore neural relationships to forelimb kinematics under these conditions. The uniqueness of the task and analysis target is a strength of the study.

      The authors remain fairly conservative and make few strong claims in the study, which may be warranted given the diversity of tuning profiles they observed.

      Weaknesses:

      There are a number of statistical tests that were accompanied by too little information to critically evaluate. Means and effect sizes needed to be better reported; some details of analyses were difficult to parse, making the strength of the conclusions difficult to evaluate.

      We will improve these aspects of the presentation and reporting of statistical analyses in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Barrett et al. examine how neural activity in the mouse motor cortex varies when a movement is performed with the ipsilateral or contralateral forelimb. First, they train animals to grasp and manipulate a pellet of food with either the left forepaw, the right forepaw, or both. Next, they measure activity in the primary and secondary forelimb motor areas (fl-M1 and fl-M2) and in the classical tongue-jaw area (tj-M1 / LOM). While responses in the forelimb areas are diverse, with some neurons preferring ipsilateral or bilateral movements, a plurality of cells prefer the contralateral limb. In LOM, by contrast, little limb selectivity is observed. At the neural population level, structure is preserved across conditions in LOM, but not in the forelimb areas. Finally, paw position can be decoded from activity in all three areas, and the LOM decoder generalized across limbs.

      Strengths:

      While previous studies in macaques have compared motor cortical activity during movement (and perturbation) of the contralateral and ipsilateral arms, no analogous work has been undertaken in rodents. This paper closes this knowledge gap by showing, for the first time, moderate-to-strong lateralization in the forelimb motor cortical areas of mice transporting grasped food pellets to the mouth, and a relative absence of lateralization in the classical tongue-jaw area. On the whole, I think this is a solid paper that reports novel observations of interest to the motor systems community.

      Weaknesses:

      The central question posed is whether cortical activity depends on the effector(s) used (ipsi forelimb, contra forelimb, or both). The corresponding hypotheses (Figure 1) are somewhat coarse-grained and are not mutually exclusive. One might expect to see condition-independent, limb-selective, and uni-/bimanual-selective signals in motor cortex (though their magnitudes could differ substantially), and to find these signals intermingled at the level of single neurons. The authors may wish to consider setting up a more focused question. For example, can bimanual responses be explained as a sum of the unimanual responses from the left and right limbs?

      In the revised manuscript, we will clarify that the possibilities illustrated in figure 1 are not intended as mutually exclusive. We indeed find all of these signals intermingled. In terms of a more focused question amenable to hypothesis testing, this can be expressed as: for each dimension (laterality vs manuality) are the activity patterns in each area closer to those predicted by invariance or dependence, as compared to the other areas? The various statistical analyses in the paper all essentially boil down to testing this question. Broadly, the answer is yes: we see activity closer to the invariant prediction LOM, and activity closer to the dependent prediction in fl-M1 and fl-M2. Testing whether bimanual activity can be explained as a simple linear sum of the left and right unimanual might provide additional insight into this question, and this analysis will be presented in the revised manuscript.

      In the area usually identified as tongue-jaw motor cortex (here referred to as LOM), unit and population activity look quite similar for ipsilateral, contralateral, and bilateral forelimb reaches. The most parsimonious explanation is that the activity is related mostly to mouth and tongue movements, rather than limb movements. Systematic mapping studies with microstimulation in the rat (Neafsy et al., Brain Res. Rev. 1986) and optogenetic stimulation in the mouse (Mayrhofer et al., Neuron 2019) tend to support the idea that tjM1/LOM is specialized for control of the tongue and mouth. Thus, I'm not entirely convinced that it "encodes ingestion-related forelimb parameters necessary for oromanual coordination." The authors could say more about this issue: what specific limb-related parameters do they think are encoded, why would these parameters be effector-independent, what evidence for this encoding is presented here, and how can limb- and mouth-related components be distinguished? The problem could potentially be addressed experimentally, as well, by delivering food pellets directly to the mouth while preventing manipulation with the paws, but this experiment isn't strictly necessary.

      The issue of orofacial movement confounds is an important one that we made a point of addressing in the discussion. There are three main points that we believe cast doubt on this as the most likely explanation for the effector-invariant representation in LOM.

      First, while we do not disagree that LOM has an important role in tongue and jaw control, there is plenty of evidence from mapping and behavioural studies (which we cite in the introduction) that it also plays a role in forelimb motor control as well.

      Second, while we cannot observe all orofacial movements, we have previously shown that the jaw is less active when the hands and LOM are most active (Barrett et al., 2024). Conversely, LOM firing is much lower during chewing, when the tongue and jaw are very active.

      Finally, the correlation between LOM firing and forelimb movements is not merely a coarse-grained one on the timescale of active manipulation vs passive holding phases. LOM firing closely tracks the position of the forelimb(s) on fast timescales and with near-zero lag, giving better decoding than from fl-M1 or fl-M2, as we have shown here and previously (Barrett et al., 2022). If we assume that LOM only encodes orofacial movements, then this result implies that orofacial movements correlate with forelimb position better than fl-M1 or fl-M2 firing correlates with forelimb position.

      The revised manuscript will include an expanded discussion to clarify these and related points.

      Because the corticospinal tract is strongly lateralized, cortical activity presumably has a smaller effect on ipsilateral than contralateral motor output. Somatosensory feedback should also be relatively lateralized for the forelimb areas. The authors could say a bit more about this issue and how it relates to their data and conclusions in the Discussion.

      We will discuss this in the revised manuscript.

      An important limitation of the behavioral task is that it involves only a single stereotyped movement for each limb, instead of multiple directions, speeds, or loads. This issue and its consequences for the analyses (especially those in Figures 7-10) and conclusions could be discussed.

      This limitation applies to the analyses relating to the transport-to-mouth movement (Figures 3-7). The population correlation structure and decoding analyses (Figures 8-10) consider activity throughout the full duration of food handling, which involves a much greater variety of movements (Barrett et al., 2020). Indeed, this was a major motivation for including these analyses. The revised manuscript will clarify this point.

      Reviewer #3 (Public review):

      Summary:

      Barrett et al. compare the responses of different parts of the mouse primary and secondary motor cortex in the context of a task where the animals manipulate and eat food using either or both hands. They find that roughly half the activity is conserved when reaching with one hand vs. the other hand, or with both. Similarity of activity was somewhat higher in the "lateral oral and manual" (LOM) part of the motor cortex, consistent with notions of a more generalized oromanual function there.

      Strengths:

      This work aims at addressing two worthwhile questions in a mouse model of motor control: (1) what specializations do we have for controlling feeding movements, and (2) how are the arms and hands coordinated with one another? The authors develop a simple but innovative apparatus to block either hand during food handling, track the behavior at high temporal fidelity, and record a sizable neural dataset. The analyses come from numerous angles to take good advantage of the data, and succeed in showing multiple lines of evidence for greater invariance in LOM than in the forelimb parts of M1 and M2.

      Weaknesses:

      There are several limitations of the current study. Most importantly, the behavior presents an inherent challenge: there is only one type of movement for each of the three conditions (contra hand, ipsi hand, and bimanual).

      See our response to reviewer #2 above regarding the variety of movements. We agree that this a limitation for the unit-level and PCA analyses, but one that is alleviated by the population correlation and decoding analyses, which relate to complex ongoing movements.

      This is entirely reasonable from the perspective that this is the ethological behavior when feeding, but it limits what analyses are possible. In particular, it precludes disentangling the neural relationship with many correlated aspects of behavior, and limits identifying population-level features of the neural activity meaningfully. This means that there are a number of alternative possible sources of the neuron-level area differences found here, and the population-level features may not be reliable.

      Important behavioural confounds include orofacial movements, non-specific movement initiation signals, and arousal. Orofacial movements we have discussed above in our response to Reviewer 2. Movement initiation signals would likely be transient and well-timed to movement onset, hence this may explain some of the effector-independent activity in fl-M1 and fl-M2 (consistent with e.g. (Kaufman et al., 2016)). However, we do not believe this to be the case in LOM as its activity is delayed and sustained relative to movement initiation. Regarding arousal, the mouse is actively engaged in consuming the food even when the hands are stationary, so there is no a priori reason to believe that arousal varies rapidly during the behaviour. Consistent with this, measurements of noradrenergic activity in the locus coeruleus during consumption suggest that arousal varies on slow timescales, on the order of seconds to tens of seconds (Sciolino et al., 2022). Such slow variation in arousal would not explain the rapid but condition-invariant changes in firing in any of the cortical areas studied here. The updated manuscript will include more detailed discussion of these points.

      Second, the behavior tracking was used at a relatively coarse level, and thus the relationships to various behavioral variables were left less distinguishable than they might have been.

      Behavior tracking was performed with kilohertz temporal resolution and submillimeter spatial resolution.

      Finally, there may be an issue with the coordinates of what is being called forelimb M1 here, which may include some hindlimb M1.

      Our recording coordinates are based on the territory of corticospinal neurons retrogradely labeled from C6 spinal cord, medial to any layer 4 labelling, as reported in our previous study (Yamawaki et al., 2021). Thus we are confident in calling this area forelimb M1.

      References:

      Barrett, J. M., Martin, M. E., Gao, M., Druzinsky, R. E., Miri, A., & Shepherd, G. M. G. (2024). Hand-jaw coordination as mice handle food is organized around intrinsic structure-function relationships. The Journal of Neuroscience, 44(42), e0856242024.

      Barrett, J. M., Martin, M. E., & Shepherd, G. M. G. (2022). Manipulation-specific cortical activity as mice handle food. Current Biology, 32(22), 4842-4853.e6.

      Barrett, J. M., Tapies, M. G. R., & Shepherd, G. M. G. (2020). Manual dexterity of mice during food-handling involves the thumb and a set of fast basic movements. PLOS ONE, 15(1), e0226774.

      Kaufman, M. T., Seely, J. S., Sussillo, D., Ryu, S. I., Shenoy, K. V., & Churchland, M. M. (2016). The Largest Response Component in the Motor Cortex Reflects Movement Timing but Not Movement Type. eNeuro, 3(4).

      Sciolino, N. R., Hsiang, M., Mazzone, C. M., Wilson, L. R., Plummer, N. W., Amin, J., Smith, K. G., McGee, C. A., Fry, S. A., Yang, C. X., Powell, J. M., Bruchas, M. R., Kravitz, A. V., Cushman, J. D., Krashes, M. J., Cui, G., & Jensen, P. (2022). Natural locus coeruleus dynamics during feeding. Science Advances, 8(33), eabn9134.

      Yamawaki, N., Raineri Tapies, M. G., Stults, A., Smith, G. A., & Shepherd, G. M. (2021). Circuit organization of the excitatory sensorimotor loop through hand/forelimb S1 and M1. eLife, 10, e66836.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study is methodologically solid and introduces a compelling regulatory model. However, several mechanistic aspects and interpretations require clarification or additional experimental support to strengthen the conclusions.

      Strengths:

      (1) The manuscript presents a compelling structural and biochemical analysis of human glutamine synthetase, offering novel insights into product-induced filamentation.

      (2) The combination of cryo-EM, mutational analysis, and molecular dynamics provides a multifaceted view of filament assembly and enzyme regulation.

      (3) The contrast between human and E. coli GS filamentation mechanisms highlights a potentially unique mode of metabolic feedback in higher organisms.

      Weaknesses:

      (1) The mechanism underlying spontaneous di-decamer formation in the absence of glutamine is insufficiently explored and lacks quantitative biophysical validation.

      (2) Claims of decamer-only behavior in mutants rely solely on negative-stain EM and are not supported by orthogonal solution-based methods.

      We thank the reviewer for the summary and noting of the strengths. We agree that the evolutionary divergence of metabolic feedback in GS homologs is a fruitful avenue for future studies. With regard to the weaknesses, the di-decamer in the absence of glutamine only forms under high (higher than physiological) concentrations of enzyme. Our primary evidence for the mutant behavior was the lack of crosslinking (Figure 1E), with supplementary support from the negative stain. In the revised version we will soften the language to say “reduced” rather than “did not support” filament formation.

      Reviewer #2 (Public review):

      The authors set out to resolve the high-resolution structure of a glutamine synthetase (GS) decamer using cryo-EM, investigate glutamine binding at the decamer interface, and validate structural observations through biochemical assays of ATP hydrolysis linked to enzyme activity. Their work sits at the intersection of structural and functional biology, aiming to bridge atomic-level details with biological mechanisms - a goal with clear relevance to researchers studying enzyme catalysis and metabolic regulation.

      Strengths and weaknesses of methods and results:

      A key strength of the study lies in its use of cryo-EM, a technique well-suited for resolving large, dynamic macromolecular complexes like the GS decamer. The reported resolutions (down to 2.15 Å) initially suggest the potential for detailed structural insights, such as side-chain interactions and ligand density. However, several methodological limitations significantly undermine the reliability of the results:

      (1) Cryo-EM data processing: The absence of critical details about B-factor sharpening - a standard step to enhance map interpretability - is a major concern. For high-resolution maps (<3 Å), sharpening is typically applied to resolve side-chain features, yet the submitted maps (e.g., those in Figures 1D, 2D, and supplementary figures) appear unprocessed, with density quality inconsistent with the claimed resolutions. This makes it difficult to evaluate whether observed features (e.g., glutamine binding) are genuine or artifacts of unsharpened data.

      (2) Modeling and density consistency: The structural models, particularly for glutamine binding at the decamer interface, do not align with the reported resolution. The maps shown in Figure 2D and Supplementary Figure S7 lack sufficient density to confidently place glutamine or even surrounding residues, conflicting with claims of 2.15 Å resolution. Additionally, fitting a non-symmetric ligand (glutamine) into a symmetry-refined map requires justification, as symmetry constraints may distort ligand placement.

      (3) Biochemical assay controls: While the enzyme activity assays aim to link structure to function, they lack essential controls (e.g., blank reactions without GS or substrates, substrate omission tests) to confirm that ATP hydrolysis is GS-dependent. The use of TCEP, a reducing agent, is also not paired with experiments to rule out unintended effects on the PK/LDH system, further limiting confidence in activity measurements.

      Achievement of aims and support for conclusions:

      The study falls short of convincingly achieving its goals. The claimed high-resolution structural details (e.g., side-chain densities, ligand binding) are not supported by the provided maps, which lack sharpening and show inconsistencies in density quality. Similarly, the biochemical data do not robustly validate the structural claims due to missing controls. As a result, the evidence is insufficient to confirm glutamine binding at the decamer interface or the functional relevance of the observed structural features.

      Likely impact and utility:

      If these methodological gaps are addressed, the work could make a meaningful contribution to the field. A well-resolved GS decamer structure would advance understanding of enzyme assembly and ligand recognition, while validated biochemical assays would strengthen the link between structure and function. Improved data processing and clearer reporting of validation steps would also make the structural data more reliable for the community, providing a resource for future studies on GS or related enzymes.

      We disagree with the reviewer’s overall assessment.

      With regard to sharpening and resolution: we examined sharpened maps and in a revised version will present additional supplementary figures showing these maps side by side. We note that the resolutions reported are global and that the most interesting features are, of course, in the periphery and subject to conformational and compositional heterogeneity. We will include supplementary figures of core side chain densities that are more like what are expected by the reviewer in the revision. With regard to modeling: the apo filament and turnover filament datasets were handled nearly identically. The additional density is therefore likely not artefactual to the symmetry operator - however, the lower resolution in this region noted by the reviewer is worthy of further exploration. The maps are public and we think this is the most plausible interpretation of the density, which we based primarily on the biochemical data and will include more speculation in the version.

      With regard to the biochemical controls: we point the reviewer to Figure S1, which shows that omission of ammonia or glutamate in the wild-type (tagless) system removes any coupling of the reactions. We will perform the additional controls to publication quality in the revised version along with the TCEP control. We note that the reducing agent is present across all experiments, ruling out an effect on any specific result. The inclusion of TCEP is also very standard in other published uses of the Coupled ATPase assay (e.g. PMID: 31778111 and PMID: 32483380 by our first author)

      Additional context:

      Cryo-EM has transformed structural biology by enabling high-resolution analysis of large complexes, but its success hinges on rigorous data processing and validation steps that are critical to ensuring reproducibility. The challenges highlighted here are not unique to this study; they reflect broader issues in the field where incomplete reporting of methods can obscure the reliability of results. By addressing these points, the authors would not only strengthen their current work but also set a positive example for transparent and rigorous structural biology research.

      All the data is public and the reviewer or anyone is free to reinterpret the maps and models - and we encourage that rather than just an interpretation of our static figures. In addition, we will upload the raw micrograph data for the apo filament and turnover filament datasets to EMPIAR prior to submitting the revision.

      Reviewer #3 (Public review):

      In this manuscript, the authors propose a product-dependent negative-feedback mechanism of human glutamine synthetase, whereby the product glutamine facilitates filament formation, leading to reduced catalytic specificity for ammonia. Using time-resolved cryo-EM, the authors demonstrate filament formation under product-rich conditions. Multiple high-quality structures, including decameric and di-decameric assemblies, were resolved under different biochemical states and combined with MD simulations, revealing that the conformational space of the active site loop is critical for the GS catalysis. The study also includes extensive steady-state kinetic assays, supporting the view that glutamine regulates GS assembly and its catalytic activity. Overall, this is a detailed and comprehensive study. However, I would advise that a few points be addressed and clarified.

      (1) In Figure 2D and Supplementary Figure 7, the extra density observed between the two decamers does not appear to have the defining features of a glutamine. A less defined density may be expected given the nature of the complex, but even though mutagenesis assays were performed to support this assignment, none of these results constitutes direct and conclusive evidence for glutamine binding at this site. I would thus suggest showing the density maps at multiple contour thresholds to allow readers to also better evaluate the various small molecules under turnover conditions that cannot be well fitted based on this density map, helping to provide a more balanced interpretation of the results.

      (2) On the same point regarding the density for the enzyme under turnover conditions, more details should be provided about the symmetry expansion and classification performed, and also show the approximate ratio of reconstructions that include this density. Did you try symmetry expansion followed by focused classification, especially on the interface region?

      (3) The interface between the two decamers of the model needs to be double-checked and reassigned, especially for the residues surrounding the fitted glutamine. For example, the side chain of the Lys residue shown in the attached figure is most likely modeled incorrectly.

      We thank the reviewer for the feedback. As noted above, we will include supplemental figures that show maps at multiple thresholds and sharpening schemes. We noted in the manuscript and above that our interpretation here is based on integrating biochemical evidence alongside the density and will make that even more clear in the revised manuscript. The filaments +/- the putative glutamine density were processed nearly identically, but we will attempt various schemes of focused classification/symmetry expansion in the revision as well. However, we point out that there is extensive averaging there that makes modeling a bit trickier than expected given the global resolution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments

      (1) Limitation to Di-decamer Formation:

      Could the authors clarify why hGS, when visualized by cryo-EM, predominantly forms di-decamers rather than extended filaments? Since glutamine bridges the two decameric rings, one would expect this to promote further polymerization. It remains unclear why longer filaments are not observed under the conditions used. Moreover, data in Figure 1B-1D are not mixed with Gln, and the mechanism of the apo-form filament formation is not clearly discussed. If K52 and C53 are in charge of filamentation, it is supposed to form a long filament, not a stack of 2 decamers. The kinetics of wild type and mutations, including K52A and C53A, are different. This confused me as K52 and C53 don't participate in the reactions of glutamine synthesis. Does data reduce catalytic efficiency with K52A or C53A mutation suggest that di-decameric GS exhibits a greater catalytic turnover rate than the pure decameric GS?

      We thank the reviewer for pointing this out and it was indeed filament length that was a point of curiosity during the study. Prior to preprint, we repeated the freezing conditions under identical turnover conditions and with high protein concentration as reported and indeed found much longer filaments - pointing to the capacity of the system to form larger complexes. However, we only captured a couple of ‘screening’ images and did not collect a full second dataset. Therefore, we hypothesize that filament length may be stochastic based on the subtle differences of individual grid vitrification based on the assumption that decamers within a filament can freely and quickly exchange. However, as shown in the time-resolved cryoEM experiment, the fraction of particles that are characterized as participating in a filament form (length-agnostic metric), does not appear to be subject to individual grid vitrification conditions but rather by experimental conditions (concentration of reaction-derived glutamine).

      The filament formation in apo state was not further explored because the concentrations required to achieve filament formation in this case were supraphysiological.

      We have edited the discussion to emphasize this point more clearly:

      “While the enzyme concentrations to achieve robust filamentation in the absence of glutamine are much higher than observed in cells, the protein concentrations used in our time-resolved cryoEM experiments where filamentation is correlated with accumulation of glutamine are within the range of intracellular GS concentrations in S. cerevisiae (Engel et al. 2025) and human cell lines (Wiśniewski et al. 2014).”

      Regarding the ability of apo-GS to form filaments - we identified that these residues were important based on their structural location at the decamer: decamer interface (Figure 1D; away from the active site as pointed out) and because their individual mutation to alanine attenuated the ability to form higher order filaments (Figure 1E). Therefore, these mutants were crucial controls in the steady-state kinetic experiments reported (Figure 2E). Here, we used exogenous glutamine to seed/stabilize filaments because we identified glutamine as serving this function and, crucially, because we did not observe glutamine occupancy in the active site under turnover conditions (which would suggest an orthosteric feedback inhibition mechanism; Supplementary Figure 10). Under these conditions, the wild-type, filament-competent protein displayed a ~3-fold K<sub>M, ammonia</sub> increase with glutamine addition compared to no glutamine, which when considered in the context of the greater E-305 flap conformational heterogeneity speaks to a model of allosteric feedback inhibition. Importantly, K52A and C53A show no difference in K<sub>M, ammonia</sub> between the glutamine and no glutamine condition, suggesting that attenuation of filament formation at this interface via mutation, eliminates the kinetic deficit. Therefore, K52A and C53A are not product-inhibited in the same manner as wild-type GS.

      We have clarified the discussion to emphasize this result:

      “Importantly, point mutations of the interfacial residues do not show a K<sub>M, ammonia</sub> defect in the presence of glutamine, indicative of the importance of the filament form for product feedback.”

      (2) Origin of Di-decamer Formation in the Absence of Glutamine:

      While the manuscript demonstrates glutamine-stabilized filamentation, the spontaneous formation of di-decamers under apo conditions is not mechanistically explained. The observation of 10-mer, 20-mer, and 40-mer species in Figure 1B should be validated against molecular weight standards or through SEC-MALS. The inference of higher-order oligomers based solely on migration is insufficient. Additional characterization (e.g., SEC-MALS, AUC, or mass photometry) would clarify whether these assemblies are biologically relevant or incidental.

      We thank the reviewer for pointing out the low precision of preparatory size exclusion chromatography assignments of GS molecular weight filament depicted in Figure 1. We have included calibration standards and assignment in Supplementary Figure 1 and updated Figure 1 to include the ambiguity of these assignments in panel B. It is important to note that GS has historically been underestimated in size via these methods and was originally assigned as an octamer for which there was previous consensus (PMID: 10708854). The low concentration requirements of mass photometry preclude its use for this purpose and we are not in a position to do SEC-MALS or AUC for this. Hopefully, the negative stain, crosslinking, and cryo-EM results are sufficient to indicate that we have correlated signals with the correct species!

      (3) Validation of Interface Mutants as Decamer-only Species:

      K52A and C53A mutants are used to disrupt di-decamer formation and are shown by negative-stain EM to exist as decamers. While supportive, this is qualitative. The inclusion of quantitative biophysical data (e.g., SEC-MALS or mass photometry) would more convincingly demonstrate that these mutants do not transiently assemble into higher-order oligomers. Furthermore, molecular measurements describing the spatial relationship of interface residues - such as the distance between K52 and E55 or between C53 residues of opposing decamers - would aid interpretation. The use of the term "adjacent" (line 530) is vague and should be made more precise.

      We thank the reviewer for the thoughtful comments. Beyond negative-stain EM we also performed a biochemical validation of the filament interface through bi-functional crosslinking based on the premise that the new filament interface, as defined by the apo-filament structure, presented new/unique pairs of nucleophilic amino acid R-groups in close proximity. We used Bis-sulfosuccinimidyl glutarate (BSG) or bis-maleimoethane (BMOE) to covalently link adjacent primary amines and sulfhydryls respectively (Figure 1E). This experiment defines two key principles of GS filament formation in the absence of glutamine:

      (1) It is concentration dependent. In the wild-type case there is a protein-dependent increase on crosslinking efficiency for both crosslinkers.

      (2) It is dependent on C53 and K52. Mutation of C53 or K52 significantly attenuated crosslinking efficiency.

      To make sure that these results are more prominent, we have now included a table of these intersubunit distances between epsilon amine groups of lysines and gamma sulfhydroxyl of cysteines groups based on the apo-filament structure and labeled this as either participating in the filament interface or not. Furthermore, in line with multiple reviewers comments, we have updated Figure 1D to include a sharpened representation of the map that shows strong side chain density for the amino acid side chains to further support these reported side chain distance measurements.

      We thank the reviewer for pointing out the low precision of the SEC chromatogram interpretation of Figure 1B and the figure has been amended to show filaments of variable length, instead of defined length. We also included a calibration curve to Supplemental Figure 1B and estimated molecular weights. While these estimates of size are lower than ground truth it is important to note that GS has historically displayed smaller than predicted molecular weights via size exclusion chromatography and analytical ultracentrifugation where initial characterization papers defined the oligomeric state as an octamer rather than decamer (PMID: 10708854). These studies and the present indicate potential adherence to resin and/or other factors about the shape of GS that lead to longer retention. Lastly, these SEC procedures were performed as a preparative step rather than for analytical purposes, so resolution was not the ultimate goal.

      (4) Terminology: "Scarless" hGS:

      The term "scarless human glutamine synthetase" is unconventional and potentially confusing. If it refers to the wild-type sequence lacking N- or C-terminal tags or mutations, I recommend using the term "native hGS" for clarity.

      We usually reserve “native” for proteins isolated from the original species and not recombinantly expressed (as here). So we will leave the term scarless in the document.

      (5) Helical Parameters of Filament Assembly:

      The manuscript states a ~26{degree sign} rotation (clockwise or counterclockwise?) between decamers in the filament, yet does not describe how this value was derived. Given that hGS filaments form helices, this parameter could be assessed via helical reconstruction. Is it possible the actual helical twist is ~30{degree sign}, implying 12 stacked decamers per full turn? Please elaborate on how the rotational angle was determined.

      Rotation was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was measured. Helical reconstruction was not pursued in this work owing to the typically short filaments observed in micrographs and the relative ease by which a 20-mer species could be selected via traditional 2D and 3D classification/reconstruction methods.

      We have added to the methods the following to better illustrate this measurement:

      “Decamer: Decamer rotation across the filament interface was determined through inspection of the apo-filament cryoEM map in ChimeraX where an outline of a pentamer from one decameric unit was rotated with respect to the outline of a pentamer across the filament interface and the rotation was depicted in Figure 1D.”

      (6) Time-Resolved Cryo-EM and Filament Growth:

      The use of time-resolved cryo-EM is innovative; however, the accessible timescales are relatively short. I suggest complementing this approach with techniques such as dynamic light scattering (DLS) or mass photometry, which allow extended real-time monitoring of filament assembly over longer durations (e.g., hours). These methods can also provide higher temporal resolution and particle size distributions.

      We thank the reviewer for this suggestion and agree that understanding the kinetics of filament formation is critical. While DLS and mass photometry are excellent for monitoring assembly over hours, our data indicates that GS filament formation occurs on a much faster timescale.

      As shown in Figure 2C, when we added ATP and Glutamine directly to GS and vitrified the sample after only 5 minutes, the majority of particles had already formed filaments, indicating that the interaction had reached saturation. This contrasts with our time-resolved experiment, where the kinetics of filament formation were likely rate-limited by the enzymatic generation of glutamine rather than the assembly process itself.

      Consequently, we anticipate that filament assembly occurs on the order of seconds or less—a timescale we interpret as a necessary prerequisite for a rapid and effective cellular feedback mechanism. Therefore, we believe the current cryo-EM data accurately captures the biologically relevant window of assembly.

      (7) Cryo-EM Symmetry Imposition and Loop Flexibility:

      The use of D5 symmetry in cryo-EM reconstructions may obscure conformational heterogeneity in flexible elements, such as the E305 loop. Since the authors used MD simulations to characterize loop dynamics, it would strengthen the study to also perform symmetry expansion followed by non-uniform refinement and alignment-free 3D classification of individual subunits. This could provide experimental validation of the proposed conformational variability.

      We agree with the reviewer that symmetry enforcement can mask conformational heterogeneity, particularly for flexible elements like the E305 loop (the E-flap). To address this, we followed the reviewer’s suggestion and performed symmetry expansion on both the turnover decamer and filament consensus maps. This was followed by focused, alignment-free 3D classification on the asymmetric unit containing the E-flap.

      Our analysis revealed a clear distinction: while 4 out of 12 turnover decamer classes showed partial density for the E-flap (class 3, 7, 8, and 12)—consistent with the flexibility observed in our MD simulations—none of the turnover filament classes demonstrated similar density. To ensure a direct comparison, we utilized a C5-expanded turnover decamer map to maintain an identical asymmetric unit to the D5-expanded turnover filament map. We note here that the turnover decamer consensus volume is different from the deposited map for which no symmetry was applied.

      Despite the different initial symmetries (D5 for filaments vs. C5 for decamers), we utilized a C5-expanded decamer map to maintain an identical asymmetric unit. For transparency, we have uploaded this C5-refined consensus map and all resulting 3D classification maps to Zenodo. We agree that the text is now strengthened given this result and we have added the following to the main text:

      “The differential loop density between turnover-decamer and turnover-filament species was further supported by 3D classification of symmetry-expanded particles, which recovered partial E305-loop density in 4/12 turnover-decamer classes (C5; Supplemental Figure 18) compared to 0/12 classes for the turnover-filament (D5; Supplemental Figure 19).”

      (8) Crosslinking Gel Analysis (Figure 1E):

      The SDS-PAGE gels shown in Figure 1E have molecular weight ladders cropped. For proper interpretation, please include full ladders with size markers and labels. In addition, clarify whether the crosslinked samples were denatured in reducing buffer. Crosslinking efficiency and specificity using BMOE or BSG require verification under reducing conditions (e.g., DTT, β-mercaptoethanol, or TCEP) to confirm covalent linkage between decamers.

      We have added in Figure 1E molecular weight markers estimates to aid in gel interpretation and have included the uncropped gels in Supplementary Figure 1D that contain the full MW ladder. The methods were clarified to indicate that reducing reagent was used in both the crosslinking reaction and all SDS-PAGE samples.

      “Protein samples were diluted to concentrations noted in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.1 mM TCEP) and, reacted with crosslinker to a final concentration of 0.5 mM for 10 mins at room temperature followed by quench in 5X SDS-PAGE sample buffer (225 mM Tris pH 6.8, 50% glycerol, 0.05 % SDS, 0.2 mg/mL bromophenol blue, 1M DTT) supplemented with 100 mM of either NH<sub>4</sub>Cl (to quench BSG reactions only) or DTT (to quench BMOE). Protein concentrations were normalized after quench prior to SDS-PAGE analysis.”

      (9) Missing Reference for NADH-Coupled Assay:

      Line 678-679 refers to an NADH-coupled assay described "previously" without citing a source. Please provide a proper reference to ensure reproducibility.

      The original paper describing the implementation of a coupled-assay to measure ADP production from glutamine synthetase was written by Bennett Shapiro and Eric Stadtman in 1970 and has been included. We will note that the conditions of this assay have been much improved since this time with better buffers, commercially available reagents of combined lactate dehydrogenase and pyruvate kinase, and modern plate readers. We added the following reference:

      “Shapiro, B.M. and Stadtman, E.R., 1970. [130] Glutamine synthetase (Escherichia coli). In Methods in enzymology (Vol. 17, pp. 910-922). Academic Press.”

      (10) Unclear Description of the NADH Assay:

      The stability of NADH is influenced by pH and light exposure. Please specify the pH range used in the assay and whether precautions (e.g., light shielding) were taken. NADH autoxidation at high pH or degradation at low pH could impact assay reliability and should be addressed in the Methods section.

      For clarity and transparency the following text was added to the Methods section.

      “Stocks of ATP, NADH, and phosphoenolpyruvate were made in base buffer (60 mM HEPES pH 7.6, 50 mM NaCl, 50 mM KCl, 10 mM MgCl<sub>2</sub>, 0.5 mM TCEP) and the pH was adjusted until it reached 7.5 on ice prior to aliquoting, flash freezing, and storage at -80°C in the dark. NADH was only exposed to light upon thawing and assay set-up and no appreciable change in absorbance of control experiments were noted.”

      (11) Ligand Density in Figure 2 and Supplementary Figure 7:

      The density attributed to glutamine, ADP, and phosphate appears broader than expected. Please include cross-correlation (CC) values, estimated occupancies, and Q-factors for ligand fitting. Varying the contour level to assess density consistency would clarify whether the observed volume represents multiple conformations, partial occupancy, or overfitting. A similar concern applies to the cysteine sidechain density.

      We have updated Supplemental Figures to include:

      (1) Globally refined map in comparison to locally refined map where both are sharpened per previous feedback.

      (2) Ligand placement now also include Q-scores and CC values

      We did not include multiple contour levels because these are included in the resolution representative Supplemental Figure and because alternative contours do not influence Q-scores. From this analysis it is apparent that phosphate and ADP are both worse fits to the density.

      Moreover, we have now included Supplementary Table 2 that includes all ligand validation statistics for the reader to evaluate the range of B-factor, CC values, and Q-scores for all ligands in all models.

      (12) Missing Ligand B-factors in Supplementary Table 1:

      The ligand refinement statistics in Supplementary Table 1 are incomplete. Please include B-factors and occupancy values for all ligands.

      We have updated the PDB depositions to include B-factors in .cif files that are now available. We have also included Supplementary Table 2 in the manuscript detailing the ligand statistics for all models including cross-correlation, Qscore, and Bfactor.

      (13) Style and Formatting Issues: format consistently throughout.

      (a) Kinetic Parameters: Please follow the IUPAC and IUBMB-recommended formatting:

      kcat should be italic with subscript.

      KM should be italic K with upright M.

      Use lowercase s<sup>-1</sup>, not uppercase S<sup>-1</sup>.

      Refer to:

      IUBMB enzyme nomenclature guidelines https://iubmb.org/wp-content/uploads/2021/01/Current_IUBMB_recommendations_on_enzyme_nome nclature.pdf

      IUPAC Green Book https://publications.iupac.org/books/gbook/green_book_2ed.pdf

      We have corrected the abbreviations according to the reviewers recommendations.

      (b) Inconsistent Terminology and Typography:

      cryo-EM vs. cryoEM are used inconsistently - standardize throughout.

      FSC 0.143 appears with and without subscript formatting-please unify.

      Line 123: "X-ray" should be capitalized.

      Line 266: CryoEM should cryoEM, lowercase "c"

      Line 571: "100 μg ml<sup>-1</sup>"-use superscript minus; ensure consistency with "mg ml<sup>-1</sup>".

      Lines 605, 606, 626, 628: MgCl<sub>2</sub>-ensure the <sub>2</sub> is subscripted throughout.

      Temperature units (lines 607, 630, 640): Write as "4 {degree sign}C" instead of "4C".

      Microliters (lines 660, 681, 697): Replace "uL" with "μL".

      Line 797: Use superscripts: K<sup>+</sup>, Cl<sup>-</sup>.

      Line 862: CO<sub>2</sub> should appear with subscript.

      We have made all terminology and typography consistent throughout based on these suggestions.

      Reviewer #2 (Recommendations for the authors):

      To strengthen the manuscript and address the methodological and interpretational gaps identified, we recommend the following revisions and additions:

      (1) Data processing and cryo-EM map quality

      (a) B-factor sharpening: Reprocess all cryo-EM maps using standardized B-factor sharpening workflows (e.g., the autoSharpen tool in cryoSPARC or similar methods) to enhance side-chain and ligand density visibility.

      We have updated main and supplementary figures to include sharpened maps. All maps were sharpened using the Autosharpen feature of Phenix, specifically, by half-maps. We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      (b) Document the specific parameters used (e.g., B-factor values, solvent content estimates) in the Methods section to improve transparency.

      See above regarding the additions made to the methods section.

      (c) Map replacement and reanalysis: Replace all figures and supplementary panels displaying raw (unsharpened) maps (e.g., Figures 1D, 2D, 3A/B, 4A, 5B, and Supplementary Figs. 2D, S7B, S10) with the newly sharpened versions. Reanalyze density features (e.g., glutamine binding sites, ATP triphosphate groups) using these revised maps and update results to reflect any changes in interpretation.

      As requested, we have updated figures with sharpened maps and found our original analyses to hold. In particular, we have included multiple metrics of ligand model scoring in Supplementary Figure 7B including Q-score and CC. Additionally, we have included all ligand model statistics in Supplementary Table 2.

      (2) Structural modeling and validation

      (a) Ligand fitting justification: Provide high-resolution ({less than or equal to}3 Å) density slices or side-chain density close-ups (e.g., for phenylalanine rings or glutamine-binding regions) to validate claims of atomic-level detail. For non-symmetric ligands (e.g., glutamine) fitted into symmetry-refined maps, explicitly describe how symmetry constraints were adjusted or applied during fitting (e.g., local symmetry refinement, manual adjustment of ligand orientation) and include validation metrics (e.g., cross-correlation scores, density fit plots) to support the placement.

      We have supplied 5 new supplementary figures to demonstrate the resolution of our sharpened cryo-EM maps (most notably Supplementary Figures 3, 4, 8, 12, 22 and panels in others) . Of particular note is the sharpened map features of R298A decamer under turnover conditions which demonstrates multiple instances of a ring density for aromatic residues.

      See discussion above regarding the placement of glutamine in the interface density and updated handling of symmetry during refinement.

      (b) Ligand density supplements: Include supplementary figures showing representative ligand-density fits (e.g., ATP, glutamine) with clear side-chain or functional group annotations, as is standard in structural biology publications.

      In our revision we have included the following updated figures and figure panels demonstrating ligand density into sharpened maps:

      Turnover Filament Glutamine Ligand: Figure 2D-E (updated representation) and Supplementary Figure 7 (new and updated representations).

      Turnover Filament ATP and Mg(II): Supplementary Figure 10 (updated representation)

      Turnover Decamer ADP and Mg(II): Supplementary Figure 5 (new figure panel)

      Turnover R298A ADP and Mg(II): Supplementary Figure 14 (new figure panel)

      (3) Biochemical assay rigor

      (a) Control experiments: Perform and report the following controls to strengthen enzyme activity claims:

      - A blank control (reaction mixture without GS, ammonia, or glutamate) to quantify background ATP hydrolysis.

      - Substrate omission controls (reactions lacking ammonia or glutamate) to confirm that ATP hydrolysis depends on both substrates and GS catalysis.

      - A TCEP effect control (compare ATP hydrolysis rates with and without TCEP) to rule out reducing agent interference with the PK/LDH coupled assay.

      We have provided blank, substrate omission, and TCEP controls in Supplementary Figure 1. These results demonstrate negligible ATP hydrolysis without complete substrate inclusion and do not indicate any impact from TCEP inclusion.

      (b) Direct activity validation: Consider supplementing the coupled assay with a more direct measure of GS activity (e.g., quantifying inorganic phosphate release via malachite green assay) to cross-validate results.

      On the merits of the PK/LDH coupled assay being used for >55 years to measure steady-state activity of glutamine synthetases and that it is a robust assay as supported by the additional control experiments presented above in Supplementary Figure 1 we have elected not to pursue tedious cross-validation with a non-continuous assay and believe our interpretation of the enzyme kinetic results hold.

      (4) Writing and presentation clarity

      (a) Methods detail: Expand the Methods section to explicitly describe:

      - Cryo-EM data processing steps, including B-factor sharpening parameters, map reconstruction workflows, and any post-processing (e.g., filtering, masking).

      - Criteria used to validate ligand fitting (e.g., density threshold values, manual vs. automated docking).

      See above the revisions made in response to critique from review #1 which we will briefly summarize here:

      We have included in the methods section the following to reflect this change:

      “Final cryo-EM maps were sharpened in Phenix using the Autosharpen feature by half-maps.”

      Focused masks are represented in Figure 2E, Supplementary Figure 18, and Supplementary Figure 19. The details around focused mask utilization are included in the revised figure captions and the following was included in the Methods.

      “Focused masks were generated in ChimeraX (v.1.7 and above). Focused refinement and 3D classification (3 Å filter resolution, PCA initialization) were performed in cryoSPARC. Strategy of class picking and refinement are noted in Supplementary Figures 18 and 19.”

      Map reconstruction workflows are present in the relevant Supplementary Figures. No post-processing steps beyond map sharpening in Phenix were carried out. In general, human GS represents a straightforward protein to reconstruction via cryo-EM.

      Ligand identification criteria was described throughout the results section. Supplementary Figure 7 was revised to show sharpened density for either globally refined or locally refined maps fit with all three products of the glutamine synthetase reaction (ADP, Pi, and glutamine) individually showing the best CC and Qscore for glutamine. Beyond Supplementary Figure 7 we also combined both biochemical experiments and cryoEM to make this ligand assignment supported by:

      (1) Time-resolved cryo-EM experiments that show increasing filament particles over reaction time (Figure 2B and Supplementary Figures 9 and 10)

      (2) Glutamine+ATP cryo-EM screening (Figure 2C) showing long filaments

      (3) Supplementary Table 2 showing reasonable ligand statistics for glutamine

      To clarify this in the Methods sections we include the following statement:

      “Ligands were placed with ISOLDE (v1.7) and those with >0.5 Qscore and supporting biochemical and/or literature precedent were built.”

      (b) Results framing: In the Results, clearly distinguish between observations supported by sharpened maps and preliminary/unvalidated features. Avoid over interpreting density in unprocessed maps (e.g., referring to "glutamine binding" in Figure 2D without noting current density limitations).

      We have updated our discussion of Figure 2D (and now also Figure 2E) to include discussion of only sharpened maps and noted current density limitations to the interpretation.

      (5) Data and material availability

      (a) Ensure all supporting data are publicly accessible:

      - Upload raw cryo-EM movies, particle stacks, and processed maps to the Electron Microscopy Data Bank (EMDB) with appropriate accession codes.

      - Deposit final atomic models in the Protein Data Bank (PDB) and reference these accession codes in the manuscript.

      We deposited maps and models with accession codes in advance of review. The PDB and EMDB codes are available in Supplementary Table 1.

      Furthermore, for the focused maps and focused classifications that were generated during the review, and for the benefit of not cluttering the PDB/EMDB, we have included these more specific analyses in Zenodo: 10.5281/zenodo.20298855.

      - Provide detailed protocols for biochemical assays (e.g., TCEP handling, enzyme purification) in the Methods or as supplementary information to enable reproducibility.

      See updates above to reviewer #1

      (b) By implementing these revisions, the manuscript will better align with eLife's standards for methodological rigor, transparency, and reproducibility, allowing readers to confidently evaluate the study's contributions to structural and functional biology.

      We agree!

      Reviewer #3 (Recommendations for the authors):

      (1) In line 252, it would be helpful to show negative-stain EM images for each SEC peak, further probing whether any peaks correspond to partially aggregated, as this could affect the measured Kcat and Km.

      We aren’t in a position to do this experiment. We routinely check for aggregation by noting Absorbance at 340nm for non-specific scattering indicative of aggregation and observed no evidence of aggregation in our fractions.

      (2) In Supplementary Figure 6, many of the classes in the "Selected Filament Classes" inset appear to be averages of closely spaced particles, which may bias the calculation and should be excluded. In the "Selected Decamer Classes", I would suggest removing the top-view particle classes, as these particles not only have significantly different ice penetration rates, but are also more difficult to distinguish in 2D classification.

      We agree and have provided an additional, more strenuous cutoff, analysis of the tr-cryo-EM data wherein only classes that show clearly aligned decamers are included and all top views are omitted (Additional Supplementary Figure 6). We are happy to say that even with the more strenuous cutoffs that our main conclusions that filaments increase with forward reaction time holds.

      (3) In line 445, "Figure 5A" should be corrected to "Figure 5B".

      We thank the reviewer for pointing this out and have made the correction.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors performed seqFISH in 26 gastruloids and performed a variety of computational analyses on these novel spatial data sets. Whilst the data is valuable and the computational concepts useful (exposure index, L-metric, ...), the article falls short on novelty and is written using a very clunky language, often with contradictory conclusions.

      We thank the reviewer for their comments about the value of our data and computational concepts. We agree with the reviewer’s critical comments and have endeavored to address them all. We believe the resulting manuscript is greatly clarified and improved.

      Major issues:

      (1) The authors did well in explaining and detailing the provenance of data and the individual experiments performed. However, their 26 gastruloid data still constitute a very limited sampling from their total organoids: one experiment pooled 4 plates at an 80-94% success rate; 6 different aggregation experiments were done, making a total of 1843 gastruloids, sampled 26 (~1-2%). A simple IF stain of 2-3 markers in a bigger sample could have given a more accurate picture of specific domains of interest and their proximity. Regardless, more information should be given about the existing samples: variation across experimental batches, differences between 300-cell vs 100-cell gastruloids that were used.

      This omission was an oversight on our part and we thank the reviewer for catching it. We did the following to address this point:

      (1) Added date labels to Figure S1.2d (now S1.1a) so that the proportion correct for each separate experiment is clear:

      (2) We added the raw images of the gastruloids used in the study taken before fixation.

      (3) We segmented these images and quantified metrics of the masks to address differences across samples in morphology. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation. Elongation was measured as 1-(the ratio of the width and the height of the segmented gastruloid area); the code can be found here:

      https://github.com/arjunrajlaboratory/ImageAnalysisProject/blob/1b2f2119f77083c27f58f8c36b14 c48d97ea706c/workers/properties/blobs/blob_metrics_worker/entrypoint.py

      Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300).

      The literature also supports our assumption that combining experiments with different starting numbers of cells would not dramatically affect the results (Bennabi et al. 2024). In this paper they show that only 35 genes were differentially expressed between gastruloids formed from 100 cells and those formed from 300 cells (compared to 319 for those formed from 1200 cells and 336 for those formed from 50 cells, both compared to 300 cells). The same paper also demonstrates that the positioning of gene expression (as measured by IF staining) for several representative genes (Bra and Foxc1) does not significantly differ between gastruloids formed from 100 and 300 cells when normalized for overall size and AP axis length (as we’ve done in this paper as well).

      We have updated the text and revised Figure S1.1 to reflect these changes.

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. The experiment was performed 3 times on different days, so to ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].

      To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (see also Revised Figure S1.1a-c).

      (2) Language in the manuscript should be revised. Overall the manuscript is very long, descriptive and written "impressions and beliefs" are often not adequately justified and indeed can be contradictory, e.g. in Section 1: the title states "cell types' locations ...are consistent", a few sentences down we find "there was substantial variation" and "within range of what would be considered a 'morphologically normal' gastruloid". "quite consistent", "compelling patterning", "we don't believe"... these types of expressions are best avoided and replaced with data or used and bolstered with quantitative numbers such as percentages when a given cutoff is used. Another example: "location of each cell type relative to gastruloid morphology was quite consistent the posterior region ... mainly consisted in NMPs." Given T expression in the posterior, this result phrased as such appears quite inflated, in fact, looking at cell types in Figures S1, 2a/b/c, this reviewer would state they are all but consistent and indeed it takes sophisticated analyses to find a pattern (of sorts) beyond the coarse domains expected!

      We thank the reviewer for their careful reading of our paper and appreciate that the work would be strengthened by increasing the degree to which quantitative measures are used to justify the statement we make. We have made the following changes to the manuscript to address this criticism:

      (1) We more clearly delineate where we are making qualitative descriptions and have removed summary language (like ‘consistent’, ‘normal’, ‘variable’ etc.) from these sections. For example, the section the reviewer refers to originally read:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we characterized the organization of each by mapping where each cell type was found relative to other types and overall morphology. The approximate location of each cell type relative to gastruloid morphology was quite consistent: the posterior region, although highly variable in size (Figure S1.2a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]....”

      And now reads:

      “Once we had the cell type identity and spatial location of each cell in all the gastruloids, we first qualitatively examined where each cell type was found relative to other types and overall morphology. The posterior region, although variable in size (Figure S1.3a,b,c), mainly consisted of neuromesodermal precursors (NMP, turquoise), a bipotent cell type that contributes to both neural and mesodermal tissues [17–19]...”

      (2) We follow this qualitative description with a quantitative analysis of cell type proportion where we clearly state which variable aspects are statistically significant:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d).”

      (3) We added a summary paragraph at the conclusion of the results from the first two figures which clearly states which aspects of gastruloid organization we find to be variable and which are consistent, with statistical testing:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (3) Figure 6 is one of the most valuable parts of the work, as the authors use the battery of analyses developed to investigate the variable and not-so-robust endothelial clusters in gastruloids. However, this investigation is still very preliminary, and it should be further linked with known biology. It is still unclear what the unique organization of this cell type is (circularity isn't convincing) and whether any signalling cues of adjacent cells could explain it. Is there any evidence that more mature endodermal cell types are generated (like the suggested "liver") to give rise to endothelial cells? It would certainly be interesting to perform IF for this cell type together with mesodermal and endodermal markers to validate seqFISH predictions on a bigger sample.

      We appreciate the reviewer pointing out that the comparisons between different endothelial cell types was interesting, and agree that the clustering methods were insufficiently justified and that a more explicit consideration of the signaling context of the gastruloid could strengthen our findings.

      We have re-evaluated how we calculate differentially expressed genes. We restricted our analysis to only consider genes that are expressed at > 2 counts/cell in at least 50% of the subsets considered. The results are in shown in the revised Figure 6.

      We find that, as the reviewer suggested, some signaling genes are significantly differentially expressed. Specifically, Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in notch signaling in the posterior of the gastruloid. Tek, on the other hand, is more expressed in the somite-associated endothelial cells, and Tek has been annotated to be involved in retinoic acid signaling. These findings align with the reviewer’s observation that signaling from adjacent cells could explain or relate to differentially expressed genes.

      We also did a more thorough review of the literature, and found several papers that reported unique subsets of endothelial precursors, albeit in related systems. In [Rossi 2022] and [Rossi 2021] the authors find a population of endoderm-associated endothelial cells in gastruloids grown with a different protocol that involves Matrigel embedding, treatment with factors that promote blood development, and growth for 168 hours. In [Veenlveit 2020] the authors find a unique somite-associated population of endothelial cells in Trunk-Like Structures, which are similar to gastruloids but model later in development and have more physical organization with discrete somites. To address the reviewer’s request that we further link with known biology we have added the following to the text:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which show more tissue-like organization than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      Finally, we tested several methods of clustering and calculating circularity, and determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and any conclusions drawn from the text.

      (4) Figures 1c and 6b need statistical significance assessments.

      We thank the reviewer for pointing out that without significance testing these plots are difficult to interpret. We have removed plot 6b (see response above about removing the circularity assessments). For plot 1c we appreciate that it is difficult to interpret which cell types vary more than others in their occurrence without significance testing. To address this we did two things: we first calculated the coefficient of variation for the proportion of each cell type across samples:

      Author response image 1.

      To calculate significance, we first considered that since these values are proportions, they must sum to 1 and changes in one cell type will affect at least one other cell type within the same sample. To properly account for this when applying statistical tests, we calculated the CLR-transformed proportion and tested all pairs of cell types for significant variation. The results are shown in the Author response image 2:

      Author response image 2.

      We added a plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We also address said variation in the text:

      “We sought to quantify variability in cell type composition between gastruloids. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centred log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.3a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.4).

      The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.4d).”

      (5) The article should include an analysis of Hox colinearity expression in these gastruloids as a validation of the system.

      We thank the reviewer for pointing out the importance of these genes in validating the gastruloid system and agree that assessing their expression specifically would help readers assess data quality.

      We analyzed the center of mass of expression along the AP axis for the Hox genes included in our panel (Hoxb6, Hoxc10, Hoxd1, Hoxb9, Hoxc8, Hoxc6, Hoxaas3, and Hoxb1). We highlighted these genes in Figure S1.3: their expression along the AP axis is consistent with previously reported expression in the tomoseq dataset from [van den Brink 2020]. A summary of the correlation coefficients for each individual gastruloid for all genes (blue) and the Hox genes (orange) is shown in the Figure S1.2. The Hox genes have similar correlation coefficients overall, although their variation is higher. This is likely due to differences in gastruloid pseudo-age; in future experiments we plan to include more Hox genes and use their expression to further classify gastruloids (see updated Figure S1.2).

      We have updated the text with these new results:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      Reviewer #2 (Public review):

      Summary:

      This manuscript presents an ambitious and technically challenging spatial-transcriptomic atlas of 26 gastruloids using seqFISH. The authors introduce quantitative metrics (mixing score, exposure index, L-metric / scL-metric, spatial L-metric, triplets) to characterize spatial organization at multiple scales. The dataset is valuable, and several analyses are original, particularly the rank-based L-metric family for mutual exclusivity.

      Strengths:

      The authors generate one of the most detailed spatial transcriptomic datasets of gastruloids to date. They propose creative computational metrics (L-metric/scL-metric) to quantify mutual exclusivity of gene expression without predefined thresholds, and they explore organizational principles from single-cell topology to cluster-level structure. Many observations align well with known gastruloid biology, such as posterior robustness and anterior variability. The writing is generally clear, and the figures are rich.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      Several central claims rely on metrics whose computation and justification are insufficiently explained, making it difficult to assess how robust or interpretable the results are. Many choices in the analysis appear arbitrary or are insufficiently motivated (normalization schemes, choice of parameters such as the number of neighbors, the distance cutoffs, hierarchical clustering setup, and so on). The interpretations of spatial consistency, gene-program inference, and endothelial heterogeneity are plausible but might be stronger than the evidence currently supports.

      The manuscript would benefit from stronger benchmarking, quantification of uncertainty, and explicit controls for known artifacts in spatial transcriptomics (e.g., spillover, 2D slicing, cell type assignment entropy). The biological insights are promising, but since several depend on methodological assumptions that have not yet been demonstrated to be stable, they would benefit from clearer methodological explanation.

      We thank the reviewer for spending time to give constructive and actionable comments, and we believe the manuscript is greatly strengthened and more consistent and clear as a result of the changes suggested.

      The work is rich and could become a reference dataset. Then, clarifying and validating the quantitative methods will considerably strengthen the impact and reliability of the conclusions.

      Reviewer #3 (Public review):

      Summary:

      Triandafillou and colleagues report a single-cell resolved spatial atlas of gene expression of 26 gastruloids. While previous work had analyzed either single-cell gene expression or spatially coarse-grained patterns of gene expression (van den Brink et al, 2020), the authors here use multiplexed sequential RNA FISH (seqFISH) to create the first gastruloid atlas, which is simultaneously spatially and cellularly resolved. This atlas adds to a growing list of resources cataloging gastruloid development (see also Suppinger et al 2023).

      To analyze this dataset, the authors also describe a novel analytical framework. Their analysis centers around the 'L-metric', which measures the degree to which pairs of genes are either coexpressed or mutually exclusive. While this metric is similar to calculating correlations in gene expressions, it has important differences (including that it can, in principle, be asymmetric; although the authors symmetrize much of their analysis). In addition to the gene-centric L-metric analysis, the authors also analyze cells in their dataset according to the cell type entropy (an information-theoretical measure of confidence in cell type assignment) and the 'exposure index' (a measure of the similarity of nearest cellular neighbors).

      Using this framework, the authors focus their analysis on two major features of development. The first is the differentiation of the bipotent neuromesodermal progenitor (NMP) cells in the posterior of the gastruloid into either presomitic mesoderm (PSM) or spinal cord SC lineages. They use L-metric analysis to compare overlap in marker genes used to separate NMP, PSM, and SC fates. They highlight that L-metric analysis can recover spatial patterns of gene expression (without explicit spatial information) and discern subtle features of marker genes beyond simple binning of cell types (e.g., that Epha5 expression in anterior NMPs may predict future SC differentiation).

      The second is the formation of endothelial (spatial) clusters within the gastruloid. The authors highlight two subtypes of endothelial clusters: (1) smaller clusters within the somitic anterior region, and (2) larger clusters associated with endoderm. While the authors discern some subtle differences in gene expression between these two clusters, their different spatial patterns suggest a potential physiological difference that would not be captured in traditional droplet microfluidic-based scRNAseq pipelines.

      Overall, this manuscript is a sophisticated and technically sound study that will provide a valuable beachhead for future studies of developmental patterning in gastruloids and organoids.

      Strengths:

      The major strengths of this study are the overall technical sophistication of the data set and analysis, as well as its potential generalizability to other developmental systems (both in vitro and in vivo). The data are extensively analyzed and reasonably interpreted, and this atlas makes good use of the variability in gastruloid development to extract the statistical structure of developmental processes. The L-metric offers a parameter-free tool to analyze transcriptomic datasets that could overcome the pitfalls of other approaches.

      We really appreciate the reviewer’s kind comments about the quality of the dataset and figures, and for pointing out the strengths of the new computational methods we developed in the analysis of this dataset.

      Weaknesses:

      The major limitations of this study are the depth and novelty of the developmental processes studied. The authors provide very convincing proof-of-concept that their dataset can recover known features of gastruloid development, including NMP differentiation and endothelial development. However, further analysis and/or investigation would be required to discover new principles of gastruloid development and patterning.

      We agree that the developmental processes studied here are not inherently novel, and we hope that by showing sufficient overlap with different, less highly resolved methods we have created a convincing document that highlights the potential for this technique to be used to analyze other systems. We appreciate the reviewer’s comments and that the manuscript is improved after making the suggested changes.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) S1.2 plates shown individually, but unclear from which experiment.

      The reviewer was right to point out this oversight — we have updated the figure (now S1.1a) with labels for the individual experiments:

      (2) Figure 2 could include a clearer indication of the types of triples/ doublets to make it even more informative.

      We thank the reviewer for pointing out that the types of triples were not clear — we’ve added a color key and more explanatory text to this figure in order which explicitly explains the type of triplets considered.

      (3) Figures should be presented in order. Figure 3c is before 3b, etc.

      We appreciate the reviewer’s attention to detail and have swapped these two panels so that their order in the figure reflects the order they are referenced in the text.

      (4) Figure 3 is interesting, and the L-metric appears useful to pinpoint crucial genes that, when expressed, indicate a type transition has occurred. It would be great to test this with another set of cell types besides NMPs/presomitic/spinal cord.

      We thank the reviewer for their interest in this biological transition, and agree that testing on another transition would be really interesting. There isn’t another set of cell types expected at this stage of gastruloid development that are predicted to have the same type of bifurcating differentiation. However, in an effort to address the spirit of this comment (that looking at other sets of cell types with the L-score would be interesting), we have used our analytical framework on a non-spatial, single-cell dataset from van den Brink et al 2020.

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (5) Figures 4 and 5 felt exploratory, and I would recommend combining them into a single figure highlighting the usefulness of L-metric and its spatial version.

      We thank the reviewer for the suggestion to merge the content of Figures 4 and 5. While we agree that they are thematically related, given the size of the heatmaps generated in the analyses, we were unable to combine them in a way that preserved the readability of the figure and stayed within the space constraints of the page size; thus, we have chosen to keep these as separate figures.

      Reviewer #2 (Recommendations for the authors):

      (1) Quantification methods require clearer formalization and justification

      A key limitation is that the manuscript relies on several spatial metrics whose definitions are not sufficiently formalized.

      (a) To evaluate biological interpretations, the reader needs a precise description of:

      - How each metric is calculated (mixing score, exposure index, scL-metric, spatial L-metric).

      - Why specific choices were made (normalizations, parameter values, distance thresholds).

      - What the expected ranges and interpretations are.

      We agree that these elements are crucial to interpreting quantitative metrics and thank the reviewer for their close read of the work. The reviewer had many comments on the exposure/mixing values calculated for Figures 1 and 2, and for the L-metric (now called L-score) values calculated for Figures 3, 4, and 5. To address the reviewers concerns we have done the following:

      (1) Created a new, unified framework for calculating cell type exposure and mixing.

      (2) Rewritten the methods section for this section with a particular emphasis on including elements the reviewer suggested, including specifically outlining normalizations and what they account for, distances chosen and the biological rationale behind them, and the range of values expected for each measure:

      “Quantification of Cell Type Spatial Relationships: Exposure index

      To characterize the spatial organization of cell types, we computed two related metrics: a pairwise exposure index matrix capturing type-specific spatial relationships, and a scalar mixing index summarizing overall spatial integration. For each cell, we identified neighbours as all cells whose centroids fell within a specified radius r of the focal cell's centroid. We chose a value of r of 16 μm, which gave an average of 5-6 neighbors per cell. We chose this value as it captures local interactions, which was the primary goal of these analyses. For each ordered pair of cell types (s, t), we calculated the exposure index as the proportion of type s cells' neighbors that are type t:

      where Ns→t denotes the count of neighbour pairs in which the focal cell is type ‘s’ and the neighbour is type ‘t’, and Ns denotes the total number of neighbours across all type ‘s’ cells. Each row of the resulting exposure matrix sums to unity and represents a probability distribution over neighbour types for a given source type. To account for differences in cell type abundance, we normalized exposure indices relative to the expectation under random spatial arrangement:

      Where pt is the proportion of cells that are type ‘t’. Normalized values of zero indicate exposure consistent with random mixing, positive values indicate spatial attraction (co-localization), and negative values indicate spatial avoidance, with a minimum of −1 representing complete exclusion.”

      “Quantification of Cell Type Spatial Relationships: Mixing index

      To summarize overall spatial integration across all cell types, we computed a mixing index defined as the fraction of neighbour pairs involving different cell types:

      where N_cross is the number of neighbour pairs involving cells of different types and N_total is the total number of neighbour pairs. For normalization, we compared the observed mixing to the expectation under random spatial arrangement:

      where

      is the expected cross-type interaction rate given cell type proportions. Normalized values of zero indicate random spatial mixing, positive values indicate greater integration than expected (hyper-mixing), and negative values indicate spatial segregation.

      Exclusion of untyped cells. When computing the mixing index, cells lacking confident type assignments were optionally excluded from both the numerator and denominator, ensuring the metric reflects only spatial relationships among typed cells. These cells were retained in the exposure matrix to quantify how typed cells interact with unclassified cells.

      Statistical analysis of variance. To identify cell type pairs whose spatial relationships varied significantly across samples, we computed the variance in exposure indices across samples for each type pair. To account for the expected relationship between mean exposure and variance, we regressed log-variance against log-absolute-mean across all type pairs and computed residuals. Type pairs with residual variance exceeding the 97.5th percentile (two-tailed α = 0.05) were considered significantly variable, indicating spatial relationships that differ across samples beyond what is expected from sampling variation and composition differences.”

      (3) We have also rewritten the methods for how the L-score is calculated, adding emphasis to where we normalize, what ranges of values are expected, and what the interpretation of these values are:

      “Calculating the single-cell L-score

      The single-cell L-score (scL-score) was computed for each ordered gene pair (gene A, gene B) within a single gastruloid. The cell-by-gene expression matrix was filtered to retain only cells with at least 2 detected transcripts for both genes. Cells were sorted in descending order by gene A's expression values; gene A thus serves as the reference distribution, and the score is asymmetric with respect to gene order.

      Three reference distributions were constructed for gene B: (1) perfect coexpression, in which gene B's values were sorted in the same descending order as gene A; (2) perfect mutual exclusivity, in which gene B's values were sorted in ascending order; and (3) independence, in which every cell was assigned the mean expression value of gene B.

      Cumulative sums of expression values were computed for gene A, for the observed expression of gene B, and for each reference distribution. The cumulative sum of each gene B distribution (observed and reference) was then plotted against the cumulative sum of gene A. This cumulative-sum-versus-cumulative-sum representation captures how gene B's expression accumulates relative to gene A's: if gene B's expression is concentrated in the same high-expressing cells as gene A, gene B's cumulative curve rises steeply at first; if concentrated in opposite cells, the curve rises steeply at the end. The area under each curve was calculated using trapezoidal integration and normalized by the product of gene A's and gene B's total expression, yielding four normalized areas: A_observed (observed relationship), A_positive (perfect coexpression), A_negative (perfect mutual exclusivity), and A_uniform (independence). This normalization ensures that scores are comparable across gene pairs with different overall expression levels (Figure S3.2a,b). The scL-score was then defined as follows:

      If A_observed > A_uniform: scL-score = (A_observed − A_uniform) / (A_positive − A_uniform), yielding values in (0, 1].

      If A_observed = A_uniform: scL-score = 0.

      If A_observed < A_uniform: scL-score = −(A_observed − A_uniform) / (A_negative − A_uniform), yielding values in [−1, 0).

      A score of +1 indicates perfect coexpression, −1 indicates perfect mutual exclusivity, and 0 indicates independence.

      The scL-score was computed for all gene pairs in each gastruloid from the 05/07/2025 dataset (n = 18 gastruloids). The other two datasets (n = 8 gastruloids) were excluded due to lower transcript detection quality. To generate an average scL-score matrix, the analysis was restricted to a common set of 202 genes well-detected across all 18 gastruloids, and per-gastruloid matrices were averaged. Unless otherwise noted, a symmetrized scL-score was used: scL-score_sym(A, B) = [scL-score(A, B) + scL-score(B, A)] / 2.

      Calculating the spatial L-score

      The spatial L-score extends the scL-score to spatial regions. For each gene, a kernel density estimate (KDE) was fitted over all detected transcript spots and evaluated on a regular square grid spanning the gastruloid. Bin side length was set to twice the median nearest-neighbour distance between detected spots, calculated separately for each gastruloid. Spatial bins were ranked by KDE-derived density and processed identically to the scL-score calculation. Low-density bins were not filtered, as KDE smoothing produced non-zero density values throughout the imaging area. The spatial L-score was symmetrized as for the scL-score, except when displaying asymmetric heatmaps.

      Hierarchical clustering of L-score matrices

      Gene-gene distances were defined as Euclidean distances between L-score vectors. Agglomerative hierarchical clustering was performed using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean'). This approach operates on L-score vector differences rather than on pairwise L-score values directly, and therefore does not require the L-score itself to satisfy the properties of a mathematical distance metric; the Euclidean distance between L-score vectors is non-negative and symmetric by construction, satisfying the requirements of Ward's method. Heatmaps display pairwise L-score values, not vector distances.

      We applied this clustering procedure to the following gene sets:

      (1) A subset of NMP, presomitic mesoderm, and spinal cord marker genes in one gastruloid (n = 36 genes; Figure 3g).

      (2) All well-detected genes excluding cell cycle genes, averaged across all gastruloids (n = 166 genes; Figure 4 and Figure S4.3a).

      (3) All well-detected genes common to all gastruloids, averaged across gastruloids (n = 202 genes; Figures S4.1a, S4.3b, S5.1a).

      (4) All well-detected genes excluding cell cycle genes in one example gastruloid (n = 171 genes; Figures 5b, S5.2a).

      (5) All well-detected genes common between our seqFISH panel and those that were detected in > 3 cells in scRNA-seq data from [XXX] (n=207 genes; Figure S4.5a).

      (6) All genes detected in > 3 cells in scRNA-seq data from [XXX] (n=19075 genes, Figure S4.5c-f).

      In Figure S4.2a,b a transformation of the L-metric values was used to cluster genes. The scL-scores were averaged across gastruloids as described above, and then each pairwise scL-score was transformed to a distance-like value with(1 - scL)/2. The matrix was then symmetrized as described previously. Agglomerative hierarchical clustering was performed directly on this transformed gene-gene distance matrix using Ward's linkage criterion (scipy.cluster.hierarchy.linkage, method='ward', metric='euclidean').”

      (4) We have added an illustrative figure about how the L-score is calculated which defines expected behaviour for several cases, gives a visual explanation of the process, and shows several extreme behaviours and their biological interpretation (see Revised Figure S3.2).

      (b) For example:

      - Mixing score: The normalization is unclear. Why only 1-nearest neighbor instead of k-NN? Why not consider existing spatial-autocorrelation metrics such as Moran's I, which would also apply to gene-level mixing?

      - Exposure index: The normalization makes the metric unbounded (e.g., exposure > 1 when local frequency > global frequency). It is unclear whether this behavior is intended. Since exposure to self is meaningful, the same metric could replace the mixing score and simplify the framework. The choice of k = 5 is not justified; parameter-free approaches like Delaunay triangulation could avoid arbitrary cutoffs. If k-nn is preferred, then the robustness of the score to change the k value should be studied.

      These issues make it difficult to interpret the magnitude of reported effects or compare them across studies.

      The reviewer makes an excellent point — we have completely overhauled this analysis in the following way to address the issues raised:

      (1) Created one unified metric (see points 1 and 2 above) which considers for every cell, the identity of its neighbors in a 16 um radius (on average 5 or 6 neighbors for each cell in each gastruloid). This value was chosen so that in most cases, the cells in the immediate vicinity of a cell were considered, but not those further out (i.e. the measure is sensitive to close interactions rather than far ones). We made a matrix of all interaction pairs for a given gastruloid, with diagonal elements representing within-type interactions and off-diagonal elements representing cross-type interactions. We have replaced the previous description with the following:

      “Several studies of gene expression in gastruloids have used pooled measurements to infer the AP axis-location of genes and cell types [3,7,13,24,27] and our data are largely consistent with these lower-resolution findings (Figure S1.3a). Yet it is obvious from individual gene staining [1,4,8,28] and our detailed 2D maps of cell identity and location that gastruloid organization is much more complex than the average order of cells along the AP axis. We thus needed an analytical method for quantifying spatial organization beyond distributions along the AP axis. To further characterize spatial organization, we sought to quantify the degree to which cells were mixed in each gastruloid, and how that mixing might vary between gastruloids. For each cell in each gastruloid, we counted the interactions between that cell and all its neighbours within a 16 μm radius (on average 5-6 neighbors per cell), and summarized all these interactions for all cells in the gastruloid in a matrix, normalizing each element by the frequency of the cell type considered to be the ‘neighbour’ in the interaction.”

      (2) To quantify overall mixing, we calculate the sum across types of the frequency of self interactions (normalized to the total interactions) and then take the inverse (1-M). We then normalize this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of that cell type).

      Because the density of cells is fairly consistent across the gastruloids, w is very close to p (the proportion of that type).

      is the expectation of cross-type rate with random mixing.

      Mixing ranges from -1 (totally segregated) to +1 (more mixed than random, i.e. there is attraction between unlike types). 0 is completely random, and negative values indicate that cell types within that gastruloid tend to cluster. We added the following to the text to explain this:

      “To quantify overall mixing, we calculated the sum (across types) of the frequency of self interactions, normalized to the total interactions) and then took the inverse. We normalized this value to the expected cross-type interactions predicted from random mixing (i.e. the proportion of the neighbouring cell type). This gave us, for each gastruloid, a value that we call the mixing index that ranged from -1 (totally segregated) to +1 (totally mixed with less frequent self-interactions than expected from chance). A mixing index of 0 indicates a random distribution, i.e., neighbour frequency is exactly what would be predicted by that cell type’s frequency alone.”

      We also edited the following description of the overall distribution of mixing indices:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). All gastruloids had a negative mixing index, indicating that they all, on average, had more like-cell type interactions than would be expected given random mixing of types. However, we note that there is a ~14% difference in the mixing index across gastruloids, meaning some variation in overall mixing is present.”

      (3) The exposure index for a given pair can be found from the off-diagonal elements of the interaction matrix, and the normalization means it represents relative overexposure/clustering (positive values) or underexposure/avoidance (negative values). The minimum value is -1 and the maximum is (1-pt)/pt. We changed the description of how the exposure index is calculated to reflect this unified method of quantification:

      We noted that the off-diagonal elements of the matrix we used to calculate the mixing index were informative about cell type-cell type interactions. Specifically they quantify the degree to which each cell type (source) is exposed to another cell type (neighbours). To assess the overall frequency of cell type-cell type interactions, we first pooled the data from all gastruloids together into one interaction matrix (Figure 2b). The measure can range from -1 (no interactions at all), with higher values indicating a greater frequency of being found in close proximity. It is asymmetric in that the exposure of cell type A to B may not be the same as the exposure of cell type B to A.

      We have updated all of the quantification in Figures 1 and 2 with these new measures.

      (2) L-metric: unclear justification and interoperability

      (a) First of all, even if it is not a major issue, the L-metric is not a "metric" at least in the mathematical sense of a metric since a metric is always positive. The L-metric is central to several major conclusions (gene exclusivity, modules, spatial organization), but its conceptual basis and computational steps need more justification.

      We thank the reviewer for pointing this out and have changed “L-metric” to “L-score” throughout. We have also endeavoured to clarify the conceptual basis and have fleshed out the various computational steps as outlined in more detail in our responses below.

      (b) Several steps (ranking, cumulative curves, area under the curve) are difficult to interpret biologically It is unclear why each transformation is required and how it responds to common scenarios (highly expressed genes, correlated vs mutually exclusive patterns).

      Since the metric is rank-based, two genes that are both highly expressed in all cells may show low L-metric despite being truly correlated.

      We appreciate the reviewer’s comments about both the interpretation of the scL-score calculation and how it behaves in common expression scenarios, particularly for genes that are broadly expressed across many cells. To address these points, we generated a set of simulated examples spanning five scenarios: ubiquitously expressed genes with similarly high average expression, ubiquitously expressed genes with differing average expression, ubiquitously expressed genes with similarly low average expression, genes coexpressed across a subset of cells rather than all cells, and genes generally expressed in opposite subsets of cells. For the first three simulations, we independently sampled two genes across 50 cells using Poisson distributions with mean expression set to 100 or 50 (to simulate a gene with high or low average expression, respectively), without any expression bias towards any subsets of cells. For the latter two simulations of dependent expression relationships, we first sampled gene 1 across 50 cells using a Poisson distribution with mean expression set to 4, then generated gene 2 from gene 1 by sampling from cell-specific Poisson distributions fitted to either generally match or oppose the transcript count obtained for gene 1 in that cell. These simulations show that genes can independently appear broadly coexpressed at the population level (regardless of average expression) simply by being ubiquitously expressed, yet still receive low scL-score values. The simulations of dependent coexpression or mutually exclusive expression relationships receive scL-score values near +1 and -1, respectively. These results align with the reviewer’s prediction, but they reflect why we designed the L-metric to follow a rank-based methodology since they preserve the specificity of the metric’s upper bound (+1) for detecting non-random coexpression relationships rather than chance coexpression relationships resulting from independently ubiquitous expression. The results of these simulations are depicted (See Revised Figure 3.3).

      We have made the following edits to the text to specifically address the case the reviewer raised about highly expressed genes:

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3.”

      (c) Interpretation of L-metric values is ambiguous

      What does 0 represent? Randoms? Co-expression? Is 1 the strongest exclusivity? The manuscript currently mixes "co-expression" and "mutual exclusivity" scales.

      We agree with the reviewer that clearly defining what values of the L-score mean is critical to understanding the text. We have added a more explicit discussion of this in the text (excerpted from the response to 2c):

      “A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes.”

      And made a figure representing visually how the L-score is calculated which shows the behaviour and biological interpretation of several extreme values and an example of how the L-score is calculated. See new Figure S3.2:

      We also edited figure captions where we referred to plots as ‘co-expression’ plots, since in some cases the plots showed genes that were mutually exclusive or not expressed together in most cells. Figure S3.1:

      “c. Spatial distribution of the expression of Nkx1-2 and Rfx4 in an example gastruloid.”

      In all other cases we checked, we used the term “co-expression” to mean the opposite of mutually exclusive, as outlined in the definition above.

      (d) Additional issues also limit interpretability - Benchmarking is missing.

      - No tests on synthetic datasets, negative controls, or curated examples.

      - Prior exclusivity methods (EEI, COTAN) routinely benchmark against ground truth; this is now standard.

      - The code for the L-metric seems to be missing in the repository.

      We appreciate this suggestion offered by the reviewer as benchmarking against prior exclusivity-oriented methods provides an important comparison for clarifying both where the scL-score agrees with existing approaches and where it offers distinct advantages. To address this, we explicitly compared the scL-score to the Exclusively Expressed Index (EEI), which is bounded below by 0 and increases with mutual exclusivity, and to the signed coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework, in which positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. We performed this comparison using seven simulated scenarios as well as four representative gene pairs from one gastruloid sample (2025-05-07_roi2). In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (high_high, high_low, low_low), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained low, and COEX was also positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.

      We then applied the same comparison to four gene pairs from one gastruloid sample (2025-05-07_roi2). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types. Pax6-Eogt, Rfx4-Eogt, and Nkx1-2-Rfx4 all showed opposing expression by scL-score, with values of -0.572 and -0.526, -0.970 and -0.955, and -0.537 and -0.615, respectively. EEI detected exclusivity most strongly for Rfx4-Eogt (0.171), but gave values of 0 or approximately 0 for the other two pairs, while COEX was negative for all three pairs (Pax6-Eogt: -0.235, Rfx4-Eogt: -0.350, and Nkx1-2-Rfx4: -0.106), consistent with opposing expression, with strongest signal for Rfx4-Eogt. By contrast, Cdx4-Cdx2 showed moderate levels of coexpression by scL-score (0.402 and 0.424) and EEI (0), but COEX indicated that they were not coexpressed (-0.331).

      These comparisons also clarify the practical advantage of the L-metric over the EEI and COTAN frameworks. EEI is based on binary zero/non-zero quantification and is therefore designed specifically to measure exclusivity rather than coexpression. COEX provides a signed value and, in our simulations, tracked both exclusivity and coexpression; however, on representative gene pairs from one gastruloid sample, scL and COEX diverged in magnitude for highly exclusive expression relationships (Rfx4-Eogt) and sign for a coexpression relationship (Cdx4-Cdx2), motivating our introduction of a signed measure based on the mutual exclusivity of expression with the scL-score (see New Figures S3.3 and S3.4).

      We have updated the text to address the reviewer’s concerns: we benchmark using simulations of commonly occurring scenarios (such as varying expression levels, degree of mutual exclusivity, and amount of noise present in the relationship between the two genes in question) as the reviewer suggested. We also provided a direct comparison to an earlier exclusivity measure (EEI):

      “To this end, we developed a pairwise metric between genes that reported the degree of mutually exclusive expression. It is calculated by rank ordering cells by the expression of one gene and measuring the degree to which the expression of the other gene is anti-rank-ordered (see Methods for details and Figure S3.2a for a visual explanation of how the measure is calculated). We call this measure the “single-cell L-score” (scL-score) because when the per-cell expression of mutually exclusive genes was plotted against one another, the data made an L shape (Figure 3e, right). A value of -1 represents perfectly mutually exclusive expression, which only happens when the genes are never found in the same cell. Higher values indicate more co-expression. Genes that are ubiquitously expressed without a strong correlative relationship between them will have a score of ~0. The maximum possible value is 1, which is obtained when both genes are expressed in a subset of all cells, and are only ever found together in those cells. We refer to this as ‘perfect co-expression’. This scale, which ranges from -1 (mutually exclusive) to 1 (perfect co-expression) captures the range of possible relationships between genes. Our expectation is that ubiquitously expressed genes like cell cycle and housekeeping genes will, due to the rank-ordered nature of the L-score calculation, have L-scores consistently close to zero no matter which genes they are compared with, whereas genes that are specifically associated with a single cell type will have an scL-score value close to -1 when compared with genes specific to other types, but higher values when compared with genes associated with the same cell type. The results of our simulations confirmed these hypotheses (Figure S3.3a).

      To benchmark this measure against existing exclusivity or coexpression measures, we calculated the Exclusively Expressed Index (EEI) [Nakajima 2021] and Coefficient of Expression (COEX) [Galfrè 2021] for the same simulated datasets (Figure S3.3b) and a subset of NMP/presomitic mesoderm/spinal cord genes (Figure S3.4a). All three methods were able, to some extent, to distinguish mutual exclusivity from coexpression, but the scL-score provided clearer separation between these different relationship types; a more detailed description of the analysis is included with Figure S3.3 and Figure 3.4.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We compared the scL-score to two existing measures of exclusivity. The Exclusively Expressed Index (EEI) (Nakajima et al. 2021) is bounded below by 0 and increases with mutual exclusivity. EEI is computed from binary zero/non-zero quantification and is designed specifically to measure exclusivity but not coexpression. The coefficient of coexpression (COEX) from the COexpression Table ANalysis (COTAN) framework (Galfrè et al. 2021) can also be used to quantify relationships between genes: positive values indicate coexpression, negative values indicate mutual exclusivity, and values near 0 indicate little structured relationship. In the two mutually exclusive simulations, all three methods detected exclusivity. In the three simulations of genes independently expressed in all cells (but with varying relative expression levels), EEI and COEX were 0, while the scL-score remained close to 0 (0.102, -0.225, and -0.025, respectively), consistent with little structured relationship. In the weak coexpression simulation, the scL-score was positive (0.770), EEI remained close to 0, and COEX was positive (0.340), indicating detectable but modest coexpression. In the perfect coexpression simulation, the scL-score reached 1.000, EEI was 0, and COEX was strongly positive (1.000). Together, these simulations show that all three methods detect strong mutual exclusivity, and both scL-score and COEX distinguish positive coexpression from exclusivity and from unstructured expression.”

      We added the following explanatory text to Supplemental Figure S3.4:

      “We calculated the scL-score, EEI, and COEX for four gene pairs from one gastruloid sample (2025-05-07_roi2). The scL-score consistently delineated gene pairs possessing opposing expression profiles, while EEI was not always able to measure those exclusivity patterns (Pax6-Eogt: scL-score=-0.572 and -0.526, EEI=0; Rfx4-Eogt: scL-score=-0.970 and -0.955, EEI=0.171; Nkx1-2-Rfx4: scL-score=-0.537 and -0.615, EEI~0). Only the scL-score was able to detect the coexpression pattern present between the positively associated expression profiles of Cdx4 and Cdx2 (scL-score=0.413, EEI=0, COEX -0.331) (shown visually in Figure S3.4a). Thus, while EEI was informative for measuring gene expression relationships characterized by mutual exclusivity, the scL-score more clearly separated positive, random, and mutually exclusive relationships on a single signed bounded scale. The COEX value trended in the opposite direction than expected, but this may be due to the fact that it cannot be calculated on single-gene pairs and necessarily uses information from the entire count table, which here only consisted of 6 genes. These comparisons combined with the simulations in Figure S3.3, clarify a conceptual advantage of the L-metric over the EEI and COTAN frameworks. In contrast, by leveraging the ranked structure of transcript counts across cells, the L-metric framework does not binarize expression and does not require fitting a parametric distribution. It can be calculated on single gene pairs, and is more sensitive to mutual exclusivity.”

      We also amended the Data and Code Availability section to include a specific reference to the L-metric package that was previously missing:

      “All code used to process the raw data and generate figures, as well as the processed data and figures can be found at the following link :

      https://www.dropbox.com/scl/fo/bchkqlbcjb8ub9m606did/AIudcWZaC566toXzb2L-jXc?rlkey=u0wgtkq8j erxoqb5ump599oip&dl=0

      Additional custom scripts used to process the raw seqFISH data can be found on GitHub:

      https://github.com/arjunrajlaboratory/NimbusImage/

      The code for calculating the L-score can be found on GitHub:

      https://github.com/arjunrajlaboratory/l-metric

      Images of all gastruloids generated for this study, as well as single-channel seqFISH images with segmentation and annotations are available here:

      https://app.nimbusimage.com/#/project/69d3fa8f1be4701f5fab6359 Raw seqFISH images are available upon request.”

      (e) Clustering using L-metric vectors

      In this study, the authors use hierarchical clustering to group genes according to the L-metric. This choice is reasonable: hierarchical clustering provides a natural representation of similarity relationships across multiple scales, and the L-metric captures a form of signed dissimilarity between genes. However, this approach raises an important issue. Standard hierarchical clustering methods typically assume a non-negative metric, whereas, as noted earlier, the L-metric can take negative values, meaning it does not strictly satisfy the requirements of a metric in the mathematical sense.

      To the best of our understanding from both the text and the source code, the authors address this issue by defining the distance between two genes A and B as the Euclidean distance between two vectors: the L-metric values from A to all other genes, and from B to all other genes. Although this procedure is mentioned in the manuscript, it is neither justified nor accompanied by any discussion of how such a distance should be interpreted. It is not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes. These two gene distances are not without overlap, but they are not the same.

      This choice magnifies the interpretability issue: readers must understand two layers of transformations. If the [-1,1] range poses problems for hierarchical clustering, simple transformations (e.g., 1 − L) or alternative clustering methods could avoid these issues.

      Given that the L-metric underlies major biological inferences (novel gene modules, spatial subclusters, endothelial states), clearer justification and benchmarking are essential. We think this can lead to more consistency in the spatial metrics.

      We really appreciate that the reviewer took the time to understand our proposed method thoroughly, and apologize for any confusion resulting from a lack of clarity in how it is calculated, and the language used to describe it. The reviewer is absolutely correct that it is not a metric in the mathematical sense; we have replaced the word ‘metric’ with the word ‘score’ throughout the text.

      The reviewer also raised concern about how the hierarchical clustering was performed, and they were absolutely correct about what the vectors represent—they are, as the reviewer states, “not the direct distance between gene A and B, but rather whether A and B have a similar L-metric to all other genes”. The heatmaps in figures 3, 4, and 5 are intended to cluster genes that have similar expression patterns, i.e. similar L-score values with all other genes. The reviewer pointed out that this transformation wasn’t clear, so we have added an explicit explanation of what the vectors represent, as well as an explanation of why we were interested in how these vectors, which represent a ‘fingerprint’ of how the gene interacts with all other genes, clustered (because this is in the section discussing NMP differentiation we focus on a specific subset of genes here, but later apply to the entire panel):

      “For each gene annotated as belonging to any of the three cell types (NMP, PSM, or spinal cord), we calculated a vector of scL-score values with all other genes. Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      We also appreciate the reviewer’s suggestion that alternative transformations of the scL-score may improve clustering interpretability. We tried using the reviewer’s suggestion of doing a simple transform: we averaged the scL-score matrices across the 18 gastruloids using the shared gene panels, transformed each scL-score from the original [-1,1] scale to a [0,1] scale using (1-scL)/2, symmetrized the resulting matrix so that each gene pair was represented by a single value, and then performed hierarchical clustering directly on this gene-by-gene distance matrix using average linkage. We used this analysis to test whether a more direct distance-based approach would change the gene groupings recovered by our original clustering method.

      This alternative approach largely recovered the same cell type-associated groupings, but the separation between branches in the dendrogram was smaller, making fine-scale ordering harder to interpret. We measured this by evaluating the average cell type dispersion, measured in terms of additive branch length, which was 0.682 and 0.671 for the 166-gene and 202-gene panels, respectively. Since the branch separation on our original dendrograms was greater and thus representative of more robust groupings, we chose to continue using hierarchical clustering based on Euclidean distances between scL-score vectors.

      We comment on this alternative transformation we tried in the next section, when considering the clustering of the entire gene panel. We feel this is appropriate as the motivation for looking at the difference vector was derived from expected behaviour of a smaller set of genes, and as the reviewer pointed out it is not clear that that expectation should or would hold for the entire panel.

      “To generate the heatmap shown in Figure 4a and S4.1a, we used the same clustering method as described earlier with the Euclidean distance between scL-score vectors. However, we also tried clustering directly on the scL-scores themselves, by transforming each scL-score from the original [-1,1] scale to a [0,1] distance-like scale using a (1-scL)/2 mapping (Figure S4.3a,b). The results were largely consistent, however the cophenetic distance scale was relatively compressed when the transformed values were used (Figure S4.3a,b). This shallow structure implies that many branches are separated by only modest distances, so fine-scale ordering within the dendrogram should be interpreted more cautiously than the larger-scale cell type block structure. We chose to continue using Euclidean distance-based hierarchical clustering of scL-score profiles, where cell type grouping is observed alongside larger cophenetic separations between clusters” (See New Figure S4.3).

      (3) Claims of "remarkably consistent" spatial organization are stronger than the data currently support

      (a) The manuscript emphasizes reproducible organization across gastruloids, but several factors complicate this interpretation.

      We agree that the distinction between what is reproducible/invariant between gastruloids and what varies was not clear in the original manuscript. To address this, we have updated the text to more explicitly distinguish between the two categories. The other suggestions made by the reviewer to sharpen the quantitative measures to strengthen these claims was very helpful and we appreciate the thought put into them, and have used that framework (emphasizing the statistically significant variations and consistencies) in summarizing our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrate several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior:posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (b) Possible selection bias. Only elongated, QC-passing gastruloids were retained; 18/26 datasets remain. seqFISH runs with uneven housekeeping signals were excluded.

      We agree with the reviewer that our data are elongated, QC-passing gastruloids, although these represent two sources of variation (biological and technical respectively). Our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category, and we have attempted to signal this to readers by consistently including language like “morphologically normal” and “elongated”. We have further updated the language in the manuscript to emphasize this point:

      “To measure the spatial distribution of gene expression, we prepared gastruloids using mouse E14TG2a cells and a standard protocol (see Methods). We harvested mature gastruloids after 120 hours of growth. To ensure consistency we checked that the proportion of the gastruloids that formed correctly was the same or greater than the median of all experiments (Figure S1.1a). Although there was variation in the length, width, and relative amounts of anterior and posterior tissues in the gastruloids considered, they were within the range of what would be qualitatively considered a ‘morphologically normal’ gastruloid [1,10].”

      In regards to the exclusion of datasets, the only time 18 out of 26 were used was when calculating the averaged L scores for all genes in Figures 4 and 5. In this case we used all 18 gastruloids from the seqFISH run performed on 4/4/2025; this dataset had the highest spot counts due to protocol improvement between runs, and integrating the datasets with very different spot counts was problematic because a minimum expression level is needed to calculate L scores. We used all 26 samples for the spatial metrics calculated in Figures 1 and 2 (Figures 3 and 6 focus on specific gastruloids). We have added additional labels in Figures 1, 2, 3, and 6 to make clear when all 26 datasets are used and when only a subset is used.

      (c) Gastruloids are known to be variable; restricting to morphologically "normal" samples could inflate apparent regularity.

      We thank the reviewer for this observation and agree that the degree of variability among gastruloids is an important consideration. As stated in response to b), our goal was to characterize the structures considered to be equivalent and morphologically normal in gastruloid studies, and to characterize gene expression and cell type variation within this category. The rate of occurrence of ‘normal’ gastruloids in our hands is ~80% (Figure S1.1a). We agree that it would be interesting to consider how variations from this baseline affect cell type composition and arrangement, and while we make no claims about it in this paper, we have updated the introduction to highlight this point:

      “To address these gaps, and to create a systematic, high-resolution dataset of gene expression in gastruloids considered to be morphologically normal, we developed a spatially resolved, single-cell molecular map of the location, identity, and gene expression of cells within 26 individual gastruloids with normal morphologies. We found that despite some morphological variability within the qualitative category of elongated and polarized, “normal” gastruloids had largely reproducible cell type composition.”

      (d) Partial lack of statistical validation. The manuscript shows descriptive consistency but no formal tests across runs or batches (e.g., mixed-effects models, ICCs, leave-one-run-out validation).

      We agree with the reviewer that a quantitative comparison between batches is important. We have added the following to the text to address this point:

      “To address potential batch effects due to biological differences between runs, we examined brightfield images of all the gastruloids generated for each experiment (529 total gastruloids across 6 plates on 3 different days), segmented them, and quantified morphological characteristics. When we embedded all 529 gastruloids into PCA space, there was near-complete overlap between all groups, with the exception of one plate from 9/1/2024, which was slightly higher in PC1. Figure S1.1b shows this embedding, and examples of gastruloids at the extreme ends of PCs 1 and 2. We note that the samples collected on 9/1/2024 were on average smaller than the other two experiments, but spanned the same range of elongation (Figure S1.1c). Interestingly, the final size as measured by cross-sectional area of a brightfield image of the gastruloid did not correlate with the initial seeding number (the experiment on 4/4/2025 used 100 starting cells and the other two experiments used 300). Previous studies have demonstrated that the gene expression differences between gastruloids seeded with 100 and 300 cells is extremely small [Bennabi 2025]” (See Revised Figure S1.1a-c).

      (e) 2D sampling limitations. Spatial metrics rely on a single imaging plane chosen as the "midplane," but z-position varies between gastruloids. AP projections, mixing, and triplet analyses could all be sensitive to z-plane choice. Prior work shows that 2D slices can underestimate distances and contacts by large margins (https://pmc.ncbi.nlm.nih.gov/articles/PMC5522766). These limitations should be acknowledged explicitly.

      We agree that sampling in 2D can limit the interpretation of our findings and we thank the reviewer for bringing up this important point. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      “Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.”

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2017].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (4) Uncertainty in cell-type assignment is not incorporated into spatial metrics

      Many spatial measurements depend directly on cell-type calls (exposure, triplets, mixing). However:

      (a) Anterior cell types have higher entropy in their marker-based scores (Figure S1.1a).

      We thank the reviewer for pointing out that several of the cell types in the anterior have high entropy — specifically cardiac mesoderm and paraxial mesoderm. However, we think there is nuance to this point; two of the other prominent anterior cell types (somite and endothelial) have low entropy scores overall, and that spinal cord/neural precursor cells, which are mostly posterior, have somewhat higher entropy; higher entropy scores are not exclusive to the anterior, nor is low entropy exclusive to the posterior. Inspired by the reviewer’s comments, we have re-analyzed our data to include this nuance (see response to point d) below.

      (b) Uncertain labels inflate apparent "mixing" or "disorder," because misclassifications randomly create mixed neighbors and triplets.

      We agree that uncertainty in labels could affect the interpretation of mixing. We appreciate these comments and the reviewer’s suggestions, and we have followed them in our response to point d) below.

      (c) Posterior cell types have low entropy, so comparisons between anterior vs posterior mixing may partly reflect label uncertainty, not biology.

      We thank the reviewer for bringing up this important caveat to our findings. We incorporated discussion of this in our text edits (see point d) below.

      (d) The authors should incorporate confidence measures (e.g., probability-weighted neighbors, entropy filtering, bootstrapping) to confirm that patterns hold independently of classification noise.

      We appreciate these suggestions and have chosen to use entropy filtering to assess whether the spatial organization we observe is highly sensitive to what values are considered ‘low’ entropy. We have updated the text (see below), and added Figure S2.2 to address the reviewer’s comments:

      “The contrast between organized posterior clustering and disorganized anterior mixing matches expectations based on literature that shows that self-organization mechanisms in gastruloids in the anterior vs. posterior are differentially sensitive to culture conditions, with somitic patterning requiring external matrix support [1,4,8,24], distinguishing it from the seemingly more autonomous organization observed in posterior cell types.”

      “One potential caveat to this finding is that differences in uncertainty in cell typing could be the primary driver of mixing and cell type interaction differences, both between the anterior and the posterior within an individual gastruloid, or overall between gastruloid. To control for this, we applied an entropy filter to our dataset. We filtered out cells that had entropy > 1.5 (see plot below for cutoff), which was chosen based on the distribution of entropy values for ‘none’ type cells, which effectively describe the upper limit of random transcript assignment (99.7% of ‘none’ type cells are removed with this filter, and about 50% of cardiac mesoderm cells and paraxial mesoderm cells, see Figure S2.2a). We first examined overall mixing; there was strong correlation between the per-gastruloid mixing indices before and after entropy filtering (Pearson r = 0.809, Figure S2.2b). Globally, mixing indices decreased with filtering, meaning that overall the cell types were more clustered. When we compared the absolute value of the change in mixing index pre and post-filtering to the proportion of each cell type, the only significant correlation was with cardiac mesoderm (Figure S2.2c). Exposure indices were overall quite similar after filtering, although the strength of somite-somite and somite-paraxial mesoderm interactions increased (Figure S2.2d).”

      “We also examined how entropy filtering might affect the exposure index, given that the mixing index is calculated from the exposure index of across cell types. In general, the magnitude of the exposure index values increased when more uncertain cells were excluded, but the directionality and relative ordering was not affected. Although the magnitude of change in the posterior cells types was less than the anterior cell types, the cross-cell type exposure values, particularly between paraxial mesoderm/endothelium and somites, doubled. From these results we conclude that mixing in the posterior is driven mainly by NMP/presomitic mesoderm interactions, and is overall lower than mixing in the anterior, which is driven by rarer cell types like cardiac mesoderm, paraxial mesoderm, and endothelium, being interspersed within somite cells” (See Figure S2.2).

      (e) This leads to reviewing the claim on endothelial heterogeneity, which strongly depend on spatial adjacency and gene exclusivity metrics.

      We have extensively considered claims of endothelial cell heterogeneity, and these are discussed in detail in response to the reviewer’s next point. We have also copied them here for the reviewer’s convenience:

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in the Author response image 3:

      Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although many of the somite-associated genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0). See Author response image 4.

      Author response image 4.

      (5) Endothelial "spatially dependent" gene expression may reflect spillover rather than intrinsic state

      The comparison between anterior-associated and posterior-associated endothelial nuclei suggests two transcriptional states. However, spatial adjacency confounds the interpretation:

      (a) seqFISH assigns transcripts to nuclei in dense tissue; partial-volume effects can mix RNA from neighboring endodermal or somitic cells.

      We thank the reviewer for their attention to detail and agree that a careful consideration of these points is important. We also note that given that some of our differentially expressed genes in endothelial cells are endoderm or somite genes, there indeed may be some transcript misassignment.

      We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript misassignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for none-typed cells at any dilation considered.

      (b) Endoderm and endothelium are closely intermixed (Figure S6.1d), and their gene expression co-localizes in KDE maps (Figure 5b-c).

      We agree with the reviewer and thank them for their close reading of the manuscript. We re-assigned transcripts to nuclei at varying levels of nuclear dilation. If, as the reviewer suggests, the differences in gene expression are due to transcript mis-assignment, then reducing the nuclear dilation should reduce the entropy in cell type score. We re-analyzed the gastruloid shown in Figure 6, and assigned spots at various levels of nuclear dilation. Without dilation, all nuclei get 132 transcripts on average, and with dilation of 12 pixels (the maximum we tested) each got 181. We reassigned cell types and calculated the cell type score entropy. The results for endothelial and endodermal cells are shown in Author response image 3.

      While we do see a small increase in entropy score with dilation for endothelial cells, neither cell type comes anywhere near approaching the cell type entropy for non-typed cells at any dilation considered.

      (c) Without explicitly quantifying spillover, differential expression between these two endothelial subsets cannot be confidently attributed to cell-intrinsic differences.

      (d) You may control for spatial proximity with any of the following:

      - Include adjacency index as a covariate in DE models.

      - Use scL-metric to test the mutual exclusivity of endothelial vs endoderm genes within the same nucleus.

      - Apply local permutation nulls: shuffle transcripts within local windows and recompute DE.

      - Restrict analysis to gastruloids that contain both endothelial subsets.

      We thank the reviewer for bringing up this important point, which we were eager to address. We agree with the reviewer’s point that by only considering gastruloids that contain both subsets of endothelial cells is the correct way to do the analysis. We were already only considering this case (specifically the values calculated in Figure 6 and for 1 gastruloid pictured in Figure 6a). We added an n=1 label to revise Figure 6d to emphasize this point.

      Additionally, we took several steps to verify that the cell states we found were a true reflection of endothelial cell biology. First, we pre-filtered genes on expression, so we only considered genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level. This was to ensure that the genes we detected were unique to that location spatially and that no effects were driven by expression from nearby tissue that could affect some cells more than others. Our list of differentially expressed genes changed — although some of the genes we had originally highlighted were still present, Gadd45g specifically was no longer present. The updated plot is shown in Figure 6.

      If we do the same analysis with the nuclear dilation equal to 0, we find similar results, although some genes are no longer present (likely due to the filtering, since overall counts are lower when the nuclear dilation is 0) (See Author response image 4).

      We have also included in the supplement larger images of some of the top differentially expressed genes, which more intuitively show the differential expression results (Endoderm enriched and Somite enriched).

      Finally, we have referenced several previously-reported instances in the literature where distinct subsets of endothelial precursors with unique gene expression programs were identified. Although in these cases 1) the embryo models were different (in [Rossi 2021, Rossi 2022] gastruloids made with a different protocol and treated with factors designed to promote blood development, and in [Veenlveit 2020] trunk-like structures) and 2) the methods were different (IF and 10x single-cell sequencing) this at least establishes a precedent for the observation of multiple types of endothelial precursors. In the case of [Veenlveit 2020] the authors specifically note that one subset is associated with somites, and we have updated the text to reflect these new results:

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b,c). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left-hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Pecam1 also clustered uniquely in our L-metric analysis (Figure 4c), suggesting this differential expression is conserved across gastruloids. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behaviour is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6)

      (6) Interpretation of gene-program modules may be overstated

      Claims that the L-metric reveals "novel gene programs" should be softened:

      (a) The seqFISH panel is an approx. 200-gene marker-enriched panel, already biased toward known cell-type markers.

      This is true and we appreciate that this came through in the text since it’s important for the reader to understand the approach we took in this study.

      (b) Strong blocks in Figure 4a and S4.1a may reflect panel design rather than newly discovered programs.

      We agree with the reviewer that the panel design was not sufficiently highlighted, so we have made the following changes to the text to emphasize which patterns would be expected due to the genes we are probing for, and which findings were surprising given the known functional role of the gene.

      At the end of the section titled ‘The L-metric captures the spatial distribution of gene expression despite being calculated without spatial information’:

      “These analyses demonstrate that information contained within the hierarchical relationships between genes, determined by scL-score can reveal novel information about cell states within cell types, although we acknowledge that since cell type is determined by a limited panel of marker genes, results should be further functionally verified. scL-score analysis can also identify distinct spatial locations of cells in this cell state, all without explicit encoding of spatial information, but rather quantifying and clustering the degree to which genes are mutually exclusively expressed with one another.”

      At the end of the section titled ‘Clustering scL-metric vectors clearly resolves cell types and reveals novel genetic interactions’:

      “Finally, although the strong blocks we find in the heatmaps in Figures 4a and S4.1a largely reflect cell types, as is consistent with our panel design, we discovered some novel functions of genes in the panel through their location in the scL-score tree: although Tgfβ was initially included in our panel to generally detect inflammatory and growth signaling, clustering by expression patterns revealed its unique association with endothelial precursors.”

      To further address the concern that the generality of clustering is due to gene selection, we performed random gene drop-out and assessed how well cell types clustered as a function of the number of genes removed:

      Author response image 5.

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      Finally, we performed scL-score analysis on an unbiased, scRNA-seq dataset without any panel selection, and were able to show similar groupings of cell type markers:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (c) Cluster robustness is not assessed (bootstrap, stability).

      We appreciate the reviewer’s suggestion that cluster robustness should be assessed. To address this point, we performed a clustering-stability analysis on the common 202-gene panel that included cell cycle genes by asking to what extent hierarchical clustering could recapitulate cell type-based groupings of genes as the number of genes used for clustering was varied across progressively larger, randomly sampled panel subsets. For each resulting tree, we averaged cell types’ dispersion of genes across the tree using the framework described in our response to suggestion 6m and compared the observed value to a permutation-based distribution generated on the same tree.

      This analysis showed that the observed cell-type dispersion remained consistently lower than the corresponding permutation distribution across all subset sizes examined. In other words, genes assigned to the same annotated cell type remained closer together in the dendrogram than expected by chance even when clustering was performed on reduced random subsets of the panel. We also observed that dispersion values increased as larger gene subsets were included, which is expected as the clustering problem becomes more complex with increasing panel size; however, the separation between the permuted distribution of average cell type dispersion and the observed dispersion value increased as the panel subset size increased. Taken together, these results indicate that the cell type-resolved organization captured by the scL-score derived hierarchy is not dependent on one particular subset of genes, but is instead a stable property of the broader gene panel. New Figure S4.2 addressing cluster stability:

      We have added the following explanation in the text:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b). To assess cluster stability, we randomly selected subsets of the panel and repeated the clustering. Regardless of panel size, the tree produced by clustering on scL-score vectors was always significantly less dispersed than permuted nulls (Figure S4.2c). Although our method of calculating dispersion can only be compared between trees clustered on the same gene set, we noted that as we increased the number of genes, the gap between the dispersion of the real tree and the dispersion of the permuted trees increased (Figure S4.2d), indicating that, as would be expected, better clustering was achieved when more genes were considered.”

      (d) Agreement with cNMF (claimed in text) is not quantified (ARI, Jaccard, hypergeometric overlap).

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See Revised Figure S4.2 (now S4.4)).

      We have updated the text to reflect these quantitative comparisons:

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (e) Testing the scL-metric on larger, unbiased scRNA-seq datasets would help demonstrate generality.

      We agree with the reviewer that this would demonstrate generality, so we applied scL-score analysis to the scRNA-seq dataset from van den Brink 2020 — several gastruloids at the same stage of development as those used in this paper were pooled and sequenced. We added a figure, new Figure S4.5 with the results of this analysis, and a new section in the paper describing them:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.”

      (7) Minor Comments

      (a) We suggest including representative raw seqFISH images. The manuscript does not show raw images, which makes it difficult to evaluate the quality of the underlying data that all spatial analyses depend on. A figure showing raw fluorescence channels, detected spots, and nuclei segmentation masks for at least one anterior region, one posterior region, and one dense interface (e.g., endoderm-endothelial) would allow readers to assess spot intensity and background levels, signal-to-noise ratio, segmentation accuracy, and potential over/under-segmentation, channel cross-talk, and transcript crowding or dropouts in dense tissues. A small panel of raw images would improve the transferability to the spatial metrics.

      This is an excellent suggestion and we have included examples of the raw images for a representative gene for all hybridizations, the spots as determined by the spot-finding algorithm distributed with the seqFISH instrument, deconvolved spots, and nuclear segmentation. These images can be found in Revised Figure S1.1d.

      (b) The Introduction mentions 3D organization, which can confuse readers into thinking the seqFISH dataset is volumetric. The data shown and analyzed come from a single 2D plane per gastruloid, not from full 3D z-stacks. Since all spatial metrics rely on true adjacency, the manuscript should explicitly state early on that the dataset is 2D and briefly justify why a single plane is sufficient for the analyses.

      We appreciate the reviewer bringing up this point and we have removed references to 3D so as not to confuse readers.

      We also have included an extensive discussion of the 2D nature of the data. The current version of the manuscript addresses the limitations of 2D sampling in the following paragraph at the end of the section titled “Cell types’ locations and relative proportions are consistent across morphologically normal gastruloids”:

      Our spatial transcriptomics is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that, in any individual gastruloid, we may collect data from a different part of the gastruloid. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. The relatively wide distribution of mixing coefficients demonstrates that even within gastruloids with broadly similar morphologies and cell type proportions, the underlying organization of cell types can vary substantially.

      To further emphasize the specific issues raised we have amended this paragraph to the following:

      “The seqFISH technique is imaging-based, and the fact that we image transcripts in a single plane admits the possibility that due to rotational differences, different parts of the gastruloid are imaged in each sample. We controlled for this to the extent possible within experimental limitations by imaging multiple gastruloids across several experiments and keeping our imaging parameters, particularly the instrument z-depth relative to the coverslip, nearly identical across experiments. We also note that previous analysis of 2D and 3D distances has indicated that in many cases, 2D distances (as we use in this work) are preferable for making comparisons between cells in a sample [Finn 2018].”

      The final sentence is derived from the abstract of the paper referenced by the reviewer, which states “We conclude that 2D distances are preferred for comparative analyses between cells, but 3D distances are preferred when comparing to theoretical models in large samples of cells. In general, 2D distance measurements remain preferable for many applications of analysis of spatial genome organization.” We thank the reviewer for bringing this paper to our attention.

      (c) Results, first paragraph: "good agreement" along the AP axis should be quantified or defined.

      In the first paragraph we compare the peak in gene expression along the (length-normalized) AP axis of each gene with a similar but orthogonally measured dataset from another group (van den Brink 2020). The text specifically reads:

      “When we compared how gene expression varies along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section.”

      We have revised Figure S1.2 with a summary plot showing the distribution of correlation coefficients for all genes and for the Hox genes in our panel (which are known to be expressed sequentially along the AP axis).

      We have updated the text as follows:

      “To assess the quality of our data, we first assigned an AP axis to each gastruloid using the expression of T, a canonical marker for the posterior (Figure 1a). When we compared how gene expression varied along the AP axis, we saw good agreement at a coarse-grained level with a previous study that sectioned gastruloids along the axis and analyzed gene expression in each section [2] (Figure S1.2a). The colinearity of the peak expression of Hox genes in our panel was also consistent with this dataset, with a median Pearson correlation of 0.695 (compared to 0.663 for all genes (Figure S1.2b).”

      (d) Clarify what "greater cell type distinction" means when using marker panels, and what metric demonstrates improvement?

      By “greater cell type distinction” we meant that when using traditional clustering methods we were not able to individually resolve some cell types: there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together. Cluster labeling was performed by considering which genes showed up as being differentially expressed in each cluster using the same associations as were used to perform the cell type scoring with the marker gene panel. While we acknowledge that these results are subjective to clustering parameters, this is a general problem with clustering and not specific to this study. We have added additional clarification in the text and removed the phrase “greater cell type resolution” since it wasn’t clear what we were comparing to:

      “To profile the spatial organization of individual gastruloids, we assigned a cell type to each nucleus using a cell type scoring method with known marker genes. We compared these results to those obtained with unsupervised clustering. We found that although clustering did produce clusters, they were not strongly separated and we were not able to individually resolve some cell types we expected to find: specifically, there was a mixed differentiation front and presomitic mesoderm cluster, and NMP and spinal cord cells were also clustered together (Figure S3.6a). Therefore, we proceeded with the scoring-based method; see Methods for additional details. A representative gallery of typed gastruloids is shown in Figure 1b; the full dataset is in Figure S1.3. Because each cell received a cell type score for each type, we could use the entropy of the cell type score probability distribution to assess confidence of our cell type assignment: a cell that received a similar score for multiple cell types would have a high entropy distribution, while one which scored highly for one type and low for the rest would have low entropy. On average, the cell type entropies for most cells in a given cell type were low, with the exception of paraxial mesoderm and cardiac mesoderm cells, which had intermediate values (see additional discussion below). The generally low entropies indicate that most of our cell type assignments were high-confidence (Figure S1.4a).”

      (e) The variation in "none-typed" cells across gastruloids should be expanded and shown quantitatively.

      We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      (f) Figure S1.1b: explain the criteria used to decide which tissues "varied significantly." Showing the proportion of "None" per gastruloid would help.

      We thank the reviewer for pointing this out. We agree that this is important and we have added the proportion of none-typed cells to Revised Figure S1.4c.

      To justify the use of the phrase ‘varied significantly’ we have calculated the degree to which cell type pairs significantly co-vary (adjusted p value < 0.05) in their proportions and added the plot to Supplemental Figure 1.4, highlighting the significantly varying pairs.

      We address said variation in the text:

      “We sought to quantify variability in cell type composition between the 26 morphologically normal gastruloids profiled. Previous single-cell datasets relied on pooling multiple gastruloids, thus obscuring the degree to which the overall cell type distribution was reflected in each individual gastruloid. However, recent single-cell measurements of individual gastruloids have suggested substantial gastruloid-to-gastruloid variation in cell type proportions [13]. Figure 1c shows distributions of cell type proportions across samples, and Figure 1d shows the coefficient of variation of these proportions. Individual gastruloid cell type distributions, including the proportion of cells that had insufficient reads to be confidently assigned a type, are shown in Figures S1.4b and c. We found that cardiac mesoderm, endoderm, and spinal cord cells had the greatest coefficient of variation in proportion between gastruloids (Figure 1d). To calculate statistical significance, we first performed a centered log-ratio (CLR) transform on the proportions, then looked for covariation between cell types across gastruloids. We found there was a statistically significant inverse correlation between the proportion of endoderm and NMP, presomitic mesoderm, and differentiation front (Figure S1.4d). We did not observe gastruloids that were as strongly neurally-biased as those reported in [13], but we did see some gastruloids with a relatively high proportion of spinal cord precursor cells (Figure S1.32a ii., xv., b vii.), and overall the proportion of spinal cord had a negative covariation with the mesodermally-derived cell types, consistent with the anticorrelation also reported in [13] (Figure S1.41dc).”

      “The proportion of somite cells was significantly positively correlated with the proportion of presomitic mesoderm cells (covariation = 0.63, Figure S1.41dc).”

      (g) In Figure 1c, showing the variability of "None" cells is important.

      We made this adjustment and added the plot to Figure S1.3c

      (h) For Figure 1e, consider normalizing to T-expression proportion, not only AP length. This could clarify multimodality in the endoderm and spinal cord.

      We thank the reviewer for this suggestion to normalize by molecular as well as physical features. We agree this could be a useful projection of the data, since expression of T is used in many contexts to define the posterior of gastruloids.

      We incorporated T expression into the length normalization in the following way: we generated a cumulative distribution of all T spots along the AP axis, and when this value exceeded a threshold (specifically 90% of all spots) we defined this as the midpoint of the gastruloid, and linearly normalized space before and after it from 0-0.5 and 0.5-1 respectively. We then re-projected nuclei for each gastruloid onto this new coordinate system, and visualized in the same way as the main text figure (See Figure 1f).

      To assess whether this increased or decreased variability in position, we calculated how much the location of the peak of each cell type in each gastruloid differed from the peak position of that cell type in all samples pooled together. We compared this difference to a bootstrapped null drawn from the pooled distribution. Interestingly, we found that nearly all cell types had statistically significant variation (meaning the average distance from the mean fell outside the bootstrapped distribution in the positive direction), however the effect size for most cell types was very small.

      Notably, as the reviewer suggested, normalizing to T expression decreased the effect size of this variability in all cases but one, and particularly decreased spinal cord variability (despite not being a marker for spinal cord):

      We interpret these findings in the following way — that most of the variation observed in cell type arrangement along the AP axis is due to morphological and molecular variability (likely due to stochasticity in initial cell number and differences in developmental timing between gastruloids), and that once these factors are taken into account, for most cell types the effect size of variability is very small. However for some cell types, notably the location of the differentiation front, the position varies among gastruloids. We hypothesize that this may be due to the rhythmic nature of somite development. We have updated the text to reflect these changes.

      “Given that the proportions of cell types within each gastruloid were fairly consistent, to what degree did their spatial organization vary? We first projected each cell’s expression onto the AP axis and looked at the distribution of AP axis locations across gastruloids (Figure 1f, left-hand side). The most posterior cell types (NMP and presomitic mesoderm) showed wide distributions from 0-30% of the axis. Centred around 30%, spinal cord precursors and endoderm had distinctive peaks (clusters), the exact location of which varied between gastruloids. Using a threshold of T expression to define the midpoint of each gastruloid and uniformly length-normalizing each half caused these cell types to collapse into a single peak (Figure 1f, right-hand side). The differentiation front was similarly located in one peak (between 30-40% of the AP axis), and the variation in the location of this peak decreased with T-expression normalization. To quantify this variability, we bootstrapped a null distribution of cell type locations by pooling each cell type together across samples and using the resulting distribution to define a reference mean. When we compared how much peak variation there was among samples randomly drawn from this null to our observed data, we found that although nearly every cell type (excepting paraxial mesoderm) had more variability than expected by chance, the effect size of this variation was small (Figure S1.5a). Normalizing location to T-expression decreased variability in most cell types, most notably for differentiation front and spinal cord (Figure S1.5b).

      We also asked how the order of cell types along the AP axis varied between gastruloids. When we ranked the peaks shown in Figure 1f per gastruloid, we found that Kendall’s W, an overall measure of rank coherence across independent samples that spans from 0 (no agreement) to 1 (complete agreement), was 0.834 (Figure S1.5c). We found that the cell types most likely to swap rank order were spinal cord and endoderm, and paraxial mesoderm and endothelium (Figure S1.5d).”

      (i) Clarify "mixing coefficient values range from 0.29-0.58" (incomplete sentence). Section "Cell types' arrangement is consistent across gastruloids, but varies by type" second paragraph.

      We appreciate that there was some confusion here and we thank the reviewer for pointing it out. We have updated the text with the new values and a more specific interpretation to aid the reader and address the reviewer’s comment:

      “The mixing index values range from -0.50 to -0.22 (Figure 2a). While all values are negative, the range was large. This observation led us to conclude that while in all the gastruloids profiled cell types tended to cluster together, there was variation between gastruloids in the degree of coherent clustering between types.”

      (j) The paragraph claiming organization consistency may be too strong, given the lack of statistical validation. "Overall, these data speak to the consistency of gastruloid organization...".

      We agree that quantification of variability, which we only assessed qualitatively, would help readers better evaluate claims of organizational consistency, and we appreciate the reviewer pointing this out.

      We have taken several steps to add quantification, including:

      (1) Calculating the coefficient of variation for the proportion of each cell type

      (2) Quantifying co-variation and assessing statistical significance

      (3) Quantifying the variation in AP-axis location, including comparison to a bootstrapped null, significance testing, and a new form of normalization which decreases some of the variability.

      (4) Quantifying order along the AP axis and calculating Kendall’s W to quantify how concordant this ordering is across gastruloids

      (5) Overhauling the methods used to quantify local patterning and mixing into one unified metric

      (6) Calculating significance for variation in exposure index across samples

      To address the question of whether claims of organization consistency are appropriate, we have added the following summary paragraph at the end of the results for the first two figures, and have taken the reviewer’s suggestion of only discussing statistically significant or quantified results in listing both consistent and variable features. Now rather than making an argument about whether gastruloids are consistent or not, we merely provide the readers with our findings:

      “Variation in cell type abundance and organization is structured and concentrated in specific cell types

      We have demonstrated that some aspects of gastruloid composition and spatial organization are consistent across gastruloids, while others are more variable. Consistent features include proportions for NMP, presomitic mesoderm, somite, and paraxial mesoderm, whose coefficients of variation were lower than other cell types (Figure 1d). Organizationally, all cell types across gastruloids are more physically clustered than random (Figure 2a), and the order in which cell types are found along the AP axis has statistically significant high agreement between gastruloids as measured by Kendall’s W (Figure S1.5c). At the local neighbourhood scale, we found that most cell type interactions were conserved across gastruloids (Figure S2.1c). At the local scale, across individual gastruloids, we found many motifs of three cells that were statistically enriched over random, suggesting a conserved local order (Figure 2c). While the normalized distance along the AP-axis of all cell types significantly varied compared to a bootstrapped null (Figure S1.5a), the effect size was small, and decreased in almost all cases when normalized to gene expression (of T) in addition to morphology (Figure S1.5b).

      However, there were also variable features. The proportion of cardiac mesoderm, endoderm, and spinal cord had the highest coefficient of variation between gastruloids (Figure 1d). Because proportions must sum to one, a change in the proportion of one cell type is necessarily linked to changes in others; we performed centred log transformation and looked for statistically significant covariation. Of all possible pairings, the following proportions had a significantly negative correlation across samples: endoderm/differentiation front, NMP/endoderm, presomitic mesoderm/endoderm, none/endothelial, and spinal cord/endothelium. This result shows that the proportions of these cell types predictably co-vary between samples, potentially suggesting some kind of biological trade-off in cell type specification or organization (Figure S1.4d).

      Across gastruloids, intra-cell type interactions (degree of clustering) of spinal cord, endoderm, and differentiation front vary (Figure S2.1b). This variation suggests that these cell types may be patterned differently between gastruloids. For example, the local motif of 3 endoderm cells found next to one another was statistically enriched within some but not all individual gastruloids, and by definition is completely absent from gastruloids lacking endoderm (Figure 2c). We interpret this contrast to mean that when endoderm is found in a gastruloid, it is consistently patterned at a local level, but may vary more at a global level. This interpretation is concordant with the findings from [Farag 2024], which demonstrates several distinct classes of endoderm organization in gastruloids.

      To summarize, while changes in the amount of individual cell types can vary, these changes are in most cases explained by variations in morphology and molecular characteristics (such as anterior: posterior ratio and the expression of morphogens like T). For patterning, we found that, in most cases, global patterns were conserved, but there were small variations in local patterning that may lead to variable meso-scale organization of specific cell types, particularly those found in the middle of the anterior-posterior axis.”

      (k) Figure 3c: gene set sizes differ substantially; proportion-based normalization may be more appropriate.

      We thank the reviewer for carefully noting the gene set differences. While the NMP-only and spinal cord-only gene sets each have 10 genes, PSM-only and NMP+PSM have 6 genes, and NMP+spinal cord has 5 genes. Given the relatively small N, we feel that normalization would likely introduce a layer of abstraction that would be more confusing for the reader, especially given the qualitative nature of the claims made about the shape of the plots in question. However, we agree this point is important, so we have now indicated the size of the gene sets on the plot (revised Figure 3b, previously c).

      (l) The purpose of the first two graphs in Figure 3c is unclear.

      We thank the reviewer for pointing out that this is unclear. The purpose of this visualization is to demonstrate correlation between genes shared between cell types and genes exclusive to only one of the cell types. The text reads:

      “Figure 3b shows the total expression of each gene group versus the NMP genes for all cells in all gastruloids that we typed as NMP, presomitic mesoderm, or spinal cord. As expected, there was a clear correlation between the mixed categories and NMP genes, supporting the notion of a continuous differentiation process.”

      We have added a correlation line to revised Figure 3b (former Figure 3c) to emphasize this point.

      (m) Clustering of all genes: quantify whether clusters align with cell types.

      We appreciate this suggestion offered by the reviewer as this analysis will allow readers to quantitatively assess the overlap between cell type-based groupings of genes and our scL-score-determined clusters for this dataset. To address this, we have used a bootstrapping approach to evaluate how well our scL-score-based hierarchies capture cell type-based groupings. Specifically, following hierarchical clustering of genes using the scL-score, for each cell type, we computed the average of the minimum cophenetic distance between each pair of genes associated with that cell type. After obtaining each cell type’s average, we took the mean of these averages which we refer to as the cell type dispersion for that tree. We note that the absolute value depends on the topology of the tree. To establish a reference for this measure, we bootstrapped a null distribution by preserving the same tree topology and randomly assigning genes to leaves and calculating the resulting dispersion. We performed this 10,000 times to create a reference null distribution. We compared the true dispersion value to this null, and computed a bootstrapped p-value. In both cases (with and without cell cycle genes), the observed average cell type dispersion is less than that for all permutations (p-value=0.0001, bootstrapping), indicating that genes associated with the same cell type were significantly more likely to cluster together on the observed scL-score hierarchy than would be expected under random clustering.

      We have updated the text to reflect these quantitative comparisons:

      “Given the amount of spatial and state information that was encoded in the scL-score heatmap for a subset of our gene panel, we expanded our analyses to all genes, hoping to discover new genetic interactions or refine existing ones. We first calculated the scL-score for all genes in all gastruloids, then averaged across gastruloids and clustered the resulting interaction vectors (see Methods for details). The heatmap is shown in Figure 4a (heatmap including cell cycle genes is shown in Figure S4.1a). We noted that just as when we clustered genes associated with NMPs and their direct descendants, genes associated with cell types tended to cluster together. Specifically, NMP, spinal cord, endoderm, and endothelial genes clustered very strongly together, while presomitic mesoderm genes again were split into two groups, one of which was more closely associated with genes involved in early somitogenesis. We quantified how well cell type-specific genes clustered compared to a random null by first calculating the dispersion of cell types within the tree topology using cophenetic distance (see Methods), and then permuting the leaves of the tree to create a null distribution of the dispersion expected by random. The results produced by hierarchical clustering on scL-score vectors were significantly (p=0.0001) more clustered than would be expected by chance (Figure S4.2a,b).

      Author response image 6.

      (n) Clarify how the distance between genes is computed for clustering; the current method is hard to interpret.

      We appreciate the reviewer’s suggestion to clarify how distances between genes were computed for hierarchical clustering. To clarify this point, we have added the following to the text:

      “Two genes that play similar regulatory or functional roles would be expected to have similar patterns of coexpression and exclusivity across the full gene panel and thus similar L-score vectors. We reasoned that the Euclidean distance between these vectors could be used instead, as it represents the degree to which A and B have a similar scL-score to all other genes considered and satisfies the requirements of a distance measure for the purposes of clustering. We performed hierarchical clustering using the distance between these vectors; the clustering therefore groups genes by the overall similarity of their coexpression profiles rather than by any single pairwise relationship. A heatmap of this clustering (with the pairwise scL-score values displayed between individual genes displayed for clarity) is shown in Figure 3g.”

      (o) Quantify similarity between cNMF and L-metric clusters.

      We appreciate this suggestion offered by the reviewer as quantifying the agreement between cNMF-derived gene programs and our scL-score-determined clusters will allow readers to more rigorously assess the extent to which these two approaches recover similar groupings of genes. To address this, we compared the top 24 genes of K=7 clusters identified using cNMF to 7 clusters (average 24 genes) obtained from scL-score-based hierarchical clustering at the appropriate cophenetic distance threshold (as originally depicted in Figure S4.2, now in updated Figure S4.4c). We computed the pairwise overlap between every scL cluster and every cNMF cluster and quantified each comparison using the Jaccard Index and Adjusted Rand Index. For each scL cluster, we plotted only the maximum value observed across its 7 possible cNMF cluster comparisons, thereby capturing the strongest correspondence between each scL cluster and the cNMF-defined programs for a given metric.

      To establish a baseline for these overlap measures, we designed a reference simulation by preserving the same cNMF clusters while defining a “permuted” set of scL clusters obtained by randomly assigning genes to clusters of the same number (7 clusters) and set of sizes (average 24 genes) as the scL clusters. As above, for each permuted scL cluster and each metric, we retained only the maximum overlap value across 7 possible cNMF cluster comparisons. We note that under this framework, the same cNMF cluster can serve as the highest-overlap comparison for more than one scL cluster.

      The following plots summarize the results of applying this approach. Higher values (closer to +1) for the Jaccard Index and Adjusted Rand Index correspond to greater overlap between observed or permuted scL clusters and cNMF clusters. Across both metrics, the observed scL clusters consistently exhibited substantially higher overlap with cNMF clusters compared to permuted scL clusters. For the Jaccard Index, the observed clusters showed markedly elevated values relative to the narrow distribution centered near 0 obtained under permutation, demonstrating that gene overlap between scL clusters and cNMF programs is greater than expected by chance. This similarly holds when gene overlap is assessed using the Adjusted Rand Index. Together, these results quantify how the scL-score can hierarchically derive clusters of genes that recapitulate major gene programs identified by cNMF to an extent beyond that expected under random clustering (See updated Figure S4.4 (formerly Figure S4.2)).

      “To validate the clustering produced by the scL-score, we compared our results to a state-of-the-art method for identifying gene programs in an unbiased fashion from single-cell data: consensus non-negative matrix factorization (cNMF) [39]. We pooled nuclei from all individual gastruloids and ran cNMF. We found that many of the resulting clusters (Figure S4.4a,b) corresponded to the clusters identified when the scL-score tree was truncated to produce exactly the same number of clusters (Figure S4.4c). The similarities were even greater when the scL-score clusters were hand-selected based on visual inspection of the tree and density of marker genes (Figure S4.4d). To quantify the overlap between clusters, we calculated both the Jaccard Index and the Adjusted Rand Index (ARI) between each scL-score cluster (Figure S4.4c) and the most similar cNMF cluster. These distributions are shown in Figure S4.4e (blue). We compared to a bootstrapped null where we permuted the genes found in the scL-score clusters, and found that permuted clusters were far less similar to the cNMF clusters than those derived from the real scL-score tree (Figure S4.4e). From these observations, we conclude that the two methods are capable of producing similar results at a high-level, but are different in their application. Individual cells receive component scores for cNMF gene programs, yielding more per-cell information, while the tree produced by L-score clustering reveals hierarchical information about gene programs, which quantifies their similarity in expression on a more global scale.”

      (p) In Figure 3e, clarify what the orange lines represent.

      We have updated the text:

      “Per-cell expression scatterplots of the two pairs of genes shown in b). The y-axis of each is the per-cell expression of Eogt. The x-axis is the per-cell expression of Pax6 (left) or Rfx4 (right). R is Pearson’s r, scL is scL-score. Count data is shown in black; smoothed 2D densities are shown in orange.”

      (q) Ensure figures and panels follow the text order. For example, Figure S1.3a is referenced earlier than 1.1 and 1.2.

      We appreciate the reviewer’s attention to detail. We have changed the order of these figures so that their reference in the text follows their numeric order.

      (r) Typo in "by covariation with any other cell type (Figure S1.1e)" did you mean S1.1.c?

      We have updated the text with this change.

      (s) For circularity (Figure 6), justify the convex-hull-based measure; thin protrusions can distort interpretation. Consider the volume difference between the convex hull and the original shape.

      We thank the reviewer for this helpful suggestion. We tested the difference method suggested by the reviewer, as well as several other methods of clustering and calculating circularity. We determined that the difference in spatial organization of endothelial cells was not robust to changes in method and parameters, so we have chosen to remove that section of the figure and text.

      Summary

      This work delivers a rich spatial dataset and introduces creative computational tools. The main limitations lie not in the data but in the clarity, justification, and validation of the quantitative methods. We suggest strengthening these aspects by adding formal definitions, parameter justification, benchmarking, robustness tests, and controlled interpretations. This will improve the manuscript's impact and reproducibility of the methods described, besides making it easier to understand for the readers.

      We thank the reviewer for their kind assessment and also for their many insightful comments for improvement. We feel the revised manuscript is greatly improved because of them.

      Reviewer #3 (Recommendations for the authors):

      In my view, this manuscript is well-designed and clearly written, and supports all of the claims made. I have no suggestions for major revisions for this manuscript; rather, I would suggest the following as outstanding questions for future investigation:

      We thank the reviewer for a careful reading of our manuscript and the several interesting suggestions and useful references in the literature. Including the discussion of these ideas in the text (see below) has improved the flow and scope of the manuscript.

      (1) On the NMP fate, bifurcation has been studied extensively in gastruloids and related structures (see, for example, Underhill et al 2023; Bolondi et al 2024). Can this dataset from Triandafillou and colleagues reveal new regulatory hierarchies in this process? This seems possible in principle, but I was not able to reach this interpretation (for example, it seems the authors interpret the clustering in Figure 3H as reflecting spatial patterns rather than a regulatory hierarchy).

      We agree that this is an exciting implication of the work, but we feel that with the current panel (which was originally chosen primarily to type cells and not to infer regulatory structure) we would not be able to comment on this. However we feel that with a larger set of genes such inference may be possible, so we took your suggestion in point 3) and applied the L-metric to a previously existing single-cell dataset. We looked for possible regulatory structure, and found at least one interesting case where a transcription factor showed strong co-expression with two previously unconnected genes. While more specific analyses and experimental validation would be required to establish a direct relationship, we feel that this suggestion by the reviewer represents an important potential future application of this methodology, so we’ve included a new figure and explanatory text to address this possibility:

      “scL-score analysis reveals cell type groupings and new transcription factor associations in a single-cell RNA-seq dataset

      To test the generality of scL-score analysis, we analyzed a previously-published dataset from [van den Brink 2020] where individual gastruloids (at the same stage as those used in this study) were pooled and subjected to single-cell RNA-seq analysis. After filtering for cell quality and common gene detection, we calculated the scL-score values using a cell-by-gene table of 14304 cells x 19075 genes.

      To first test whether we could reproduce the results from this study, we performed hierarchical clustering on the scL-score difference vector (as previously described) on the set of 207 well-detected genes that were also present in our seqFISH gene panel. The resulting tree showed clustered cell types, similar to the tree produced with the expression data in this study (compare Figure S4.5a to Figure S4.4d). The cell types were clustered significantly more than expected by chance (Figure S4.5b).

      Then, to test whether scL-score analysis would be effective for analyzing the entire dataset, we performed hierarchical clustering on all 19075 genes. We then examined the resulting heatmap (Figure S4.5c) for clusters of interest. We observed a cluster enriched for endothelial genes (Figure S4.5d), which contained some genes in our panel but many others which were not; this finding demonstrates that the clustering in Figure 4a is not solely due to the selection of genes in our seqFISH panel. We also observed a large cluster that contained genes associated with pluripotency or primordial germ cell fate (Figure S4.5e). Although some of the genes in this cluster were in our seqFISH panel, when we performed scL-score analysis we did not see them cluster with each other or with any other cell type genes. This lack of clustering implies that in our dataset cells that co-express these genes may be rare or too poorly detected to cluster strongly; however, the same analysis performed with more cells and genes showed association. This result demonstrates that clustering scL-score difference vectors can identify known cell-type-associated genes, even within transcriptome-scale data. Finally, we also found a small cluster showing strong co-expression of the transcription factor Gata4, a crucial regulator of the development of visceral and parietal endoderm, and two other genes: a predicted gene of unknown function (Gm43715) and Troponin C (Tnnc1) (Figure S4.5f). Intriguingly, Gata4 has been implicated in heart development (albeit in an indirect manner) [Watt 2004], and troponin C is important for cardiac muscle cell contraction and has been implicated in cardiomyopathy, although at a much later stage of development than that modeled by gastruloids [Li 2015].

      Together, these results demonstrate that scL-score analysis is reproducible across datasets, even when different numbers of genes are compared. It effectively clusters genes associated with cell types, and can reveal developmental transitions. Moreover, increasing the number of cells and genes can reveal new clusters, some of which may predict novel regulatory interactions or spatial co-occurrence not previously observed.

      (2) On the formation of endothelial clusters ('blood islands'), this process has been observed in gastruloids (Rossi et al 2022). The observation of endoderm-associated endothelium in gastruloids is interesting, but it was also not clear how the authors interpret this finding. Are there two separate endothelial differentiation paths captured here? Or is there one path, and only some migrate towards the endoderm? The authors seem to raise possibilities, and it was slightly unclear on my reading how they interpreted their findings.

      We thank the reviewer for pointing us to this paper. We think it’s especially interesting that we also see close association between endoderm and endothelial precursors, especially given that the protocol used in the referenced paper was designed to generate blood precursors (treatment with VEGF, bFGF and ascorbic acid). In light of the comments of other we re-did this analysis — Author response image 7 shows the updated list of differentially expressed genes.

      Author response image 7.

      Some of the spatially differentially expressed genes are linked to signalling, and likely reflect overall signalling differences between the anterior (where the somite-associated endothelial cells are) and the posterior (where the endoderm-associated endothelial cells are). For example, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA is higher in the anterior). Pecam1 is involved in adhesion and cell morphology, and perhaps is higher in cells interacting with endoderm due to tighter packing/association; the same could also be true of Cdh5.

      While none of these answer the reviewer’s questions about the origin of the cells (and whether it is common), we found another instance of differential endothelial populations in embryo models: in Veenvliet 2020, they find that some endothelial cells have a mesodermal (somitic) origin. Thus we may be seeing a similar phenomenon in our samples. We have updated the text to reflect these additional lines of evidence and to clarify how we think the two populations may differ (while acknowledging that we lack the tools to confidently assign cell of origin or functional differences with this technique):

      “We observed that in 5 out of the 26 gastruloids, there was a large central patch of endoderm cells intermixed with endothelial precursors; these samples also had unique spatial L-score clustering of endothelial and endoderm genes (Figure 5b). An example of one such gastruloid is shown in Figure 6a. Migration to and association with the endoderm is also a hallmark of endothelial development [47,48], and we were curious whether there were differences between these cells and the cells we observed forming anterior, somite-associated clusters. When we computed the cell type exposure index for just this gastruloid, we found that, consistent with our visual observations, in this particular sample, endothelial and endoderm cells were much more frequently found next to one another than on average (Figure 6b). To determine whether these spatial and organizational differences reflected gene expression differences, we divided the gastruloid normal to the anterior-posterior axis to separate the endothelial cells into endoderm-associated and somite-associated and looked for differentially expressed genes between the two groups in this gastruloid. To ensure we were focused on genes that truly varied in expression in endothelial cells and were not merely a reflection of spillover from surrounding cells, we pre-filtered genes on expression, so only genes that were present in at least 50% of the cells in either group at a greater than 2 count per cell level were considered. The significantly differentially expressed genes after filtering are shown in Figure 6d. As an additional check on the degree to which transcript mis-assignment affected our analysis of gene expression in these cells in particular, we varied the nuclear dilation in this gastruloid specifically, and calculated cell type score entropy as a function of nuclear dilation (Figure S6.1a). Because cell type score entropy of a cell reflects the degree to which that cell specificity expresses genes associated with a single cell type, our expectation was that if spillover between endoderm and endothelial cells was a significant issue, then decreasing the nuclear dilation should greatly decrease the entropy scores for both groups. Although we saw a slight increase in the spread of the distribution as nuclear dilation increased, the median cell type entropy stayed extremely low for both groups (Figure S6.1a). From this analysis we conclude that the genes we identify as differentially expressed are not due to spillover from surrounding cells, but instead are due to spatially-dependent differences in endothelial cell biology.”

      The genes with the highest fold-change in expression in endoderm-associated endothelial genes are shown on the left hand side of Figure 6d. Two are endothelial genes: Pecam1 and Cdh5, both of which are associated with angiogenesis. Spatial expression of these genes is shown in the top row of Figure 6e (larger version in Figure S6.1b). Notch1 is more expressed in endoderm-associated endothelial cells, and this could reflect an increase in Notch signaling in the posterior of the gastruloid. [Chan et al 2017] demonstrated that Notch signalling can be sensitive to shear stress, raising the possibility that the differences in cell state we observe may be driven by differences in mechanical forces in the anterior and posterior. Although most endothelial cells are thought to be of mesodermal origin, some evidence suggests that, in the organogenesis of specific tissues like the liver, the endoderm can give rise to endothelial cells [49]. Furthermore, in [Rossi 2022] the authors show that in a gastruloid-like model specifically designed to model blood development, there is strong spatial adjacency between endothelial and endoderm cells. They hypothesize that these may be a subset of endothelial cells, specifically hemogenic endothelial cells (which have the potential to become blood progenitors). Our data demonstrate a molecularly driven organization distinct from the clustering we observed in the anterior and suggest that multiple mechanisms of endothelial specification could be modeled in gastruloids, even simultaneously within the same structure, although further characterization is needed to determine exactly what processes these unique endodermal/endothelial structures model.

      Several other endothelial genes are instead differentially expressed in somite-associated endothelial cells: Nrp2, Tek, Apoe, and Cldn5. Although these genes have less obvious functional distinctions than the endoderm-associated genes, Nrp2 enables semaphorin receptor activity, including nervous system development and ventral trunk neural crest cell migration and Tek negatively regulates endothelial cell apoptotic process and response to retinoic acid (RA), which is known to be higher in the gastruloid anterior. Furthermore, a specialized population of endothelial precursors associated with somites was also observed in trunk-like structures, which are more organized organoids than gastruloids [Veenvliet et al. 2020].

      Although endothelial cells have consistently been observed in single-cell measurements of gastruloids, their relative rarity has precluded in-depth analysis of subtypes or inference of spatial location. Our results strongly suggest that endothelial precursor formation, migration, and organization may all be modeled in 3D gastruloids, even without treatment with additional factors as in [Rossi 2021, 2022]; recent advances in 2D gastruloids have allowed modeling of cardiac and hepatic vascularization [45], and our data suggest that 3D gastruloids may similarly be adapted to model more specific aspects of hematopoiesis and vascularization. Early specification from a pool of mesodermal precursors is a hallmark of the endothelial lineage [47]; given the consistency with which we observe endothelial precursors, we speculate that this behavior is recapitulated in gastruloids, but further epigenetic measurements are required to validate this hypothesis” (See Revised Figure 6).

      (3) Finally, a broader question about the analytical framework. The authors emphasize that the L-metric is parameter-free; however, much of their analysis still appears to rely on baked-in priors about known marker genes for cell type assignment. Is there a way to extend their analysis to infer cell types directly from the structure of the L-metric? The hierarchical clustering in e.g., Figure 3G suggests something like this: the hierarchy mostly (but not exactly) follows the marker gene annotation. Does this suggest that the cell type labeling should be revisited?

      Once of the initial motivations behind the creation of the L-metric (now L-score) was to have a more reliable and specific way of quantifying the interaction between genes that are known in the literature to be cell type markers — with more sophisticated (and noisy) methods of analysis like single-cell RNA sequencing, we found that there was substantial variation in the specificity and ubiquity of so-called ‘marker genes’, and that their usefulness often depended on context. We appreciate that the reviewer raises this point as well, and we think that scL-score analysis, such as that exemplified in Figure S4.4, can help identify new marker genes, or at the very least distinguish the biological context in which a marker gene is useful.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Comments on revised version:

      I thank the authors for their extensive efforts to revise the manuscript. I have no further concerns.

      Reviewer #2 (Public review):

      The authors have made several corrections to the original manuscript. For example, they revised the bootstrapping analysis to avoid arbitrarily inflating the degrees of freedom. However, most substantive concerns remain inadequately addressed.

      (1) The primary issue is still the lack of baseline models against which to benchmark the predictive performance of the proposed DenseNet model. This concern was raised independently by two reviewers. Without such benchmarks, it is difficult to interpret the reported results in the context of prior work on MRI-based cognition prediction.

      Notably, the authors state: "While we compared our model with the connectome predictive modeling (CPM) approach and observed better performance with our deep learning framework, we did not conduct a comprehensive benchmark across all available machine learning methods, nor was this the aim of the present study."

      However, I could NOT find any discussion or results related to the CPM model in the manuscript. It is therefore unclear whether the DenseNet model was actually statistically compared with CPM, and, if so, how the comparison was conducted.

      Note that the statement, "While Vieira et al. show that the majority (76%) of prior studies used linear modeling approaches, including CPM and penalized regressions, these models are often vulnerable to overfitting, especially when applied to high-dimensional fMRI data," is not entirely accurate. Linear models typically have far fewer parameters than deep-learning models and are therefore often less prone to overfitting. In fact, it is well established that deep-learning models are particularly susceptible to overfitting and usually require substantially larger sample sizes to achieve stable and reliable performance. Although deep-learning models may outperform shallower models once sufficient data are available and training is well controlled, this does not justify the authors' claim as stated. I therefore disagree with the argument put forward by the authors.

      The authors further justify the absence of benchmarking by stating: "In this context, deep learning was employed as a flexible framework capable of modelling high-dimensional functional connectivity patterns across cognitive states, rather than as a claim of inherent methodological superiority. Thus, our goal was not to propose a universally superior prediction model, but rather to test how brain state influences predictive utility for WM and EM using a deep learning approach." However, most shallow models can likewise be applied across different brain states and cognitive targets. This rationale does not establish deep learning as a uniquely appropriate or necessary choice. If deep learning is indeed a better approach in this context, the authors should demonstrate this empirically through appropriate benchmarking against established baseline models.

      We thank the reviewer for reiterating this point. As noted in both the manuscript and our previous response, the primary goal of the present study was not to benchmark predictive algorithms, but rather to compare the predictive utility of different brain states using a consistent modeling framework.

      We did perform an exploratory CPM analysis using a conventional implementation that included correlation-based feature selection (p < 0.01), summarization of positive and negative networks, and robust regression for prediction using 3-fold validation. However, CPM performance can depend substantially on analytical choices, including feature-selection thresholds, treatment of positive and negative networks, cross-validation strategies, and model specification. Although we obtained CPM results (see below), we did not systematically evaluate how these choices influenced performance, nor did we optimize CPM to the same extent as would be required for a rigorous methodological comparison.

      For this reason, we chose not to include a CPM-versus-DenseNet comparison in the manuscript. Any direct comparison could easily be overinterpreted as evidence for the superiority of one approach over another, despite the absence of a comprehensive benchmarking framework. We therefore deliberately avoided such claims and instead focused on the scientific question motivating the study: whether predictive performance differs across brain states when the same predictive framework is applied consistently. We agree that comparisons with CPM and other predictive approaches would be valuable, but we believe such analyses are better suited to a dedicated methodological benchmarking study.

      (2) Additional analysis shows that "BCG is not significantly associated with cognition itself". This is the most perplexing result. This is like saying Brain Age Gap is not related to chronological Age. It is counterintuitive since the Brain Age Gap is calculated by chronological age minus actual age, and most research has shown a strong relationship between the Brain Age Gap and age.

      If the brain cognition gap is not related to cognition, is it possible that the results found are mainly due to the predictive model not fitting well with another dataset? Regardless, the lack of association between BCG and cognition deserves a discussion.

      We thank the reviewer for this comment. The absence of a significant BCG-cognition association might be unexpected. We agree it warrants careful interpretation. Theoretically, when a predictive model is trained and evaluated across samples with differing age distributions, the regression-to-the-mean dynamics that typically create BAG–age dependence may not transfer in the same way to BCG–cognition relationships, particularly in age-homogeneous cohorts such as COBRA. However, we acknowledge that the present findings alone are insufficient to resolve this question fully.

      We have added additional findings as supplementary material (Figure S4).

      (3) I still do not fully understand the rationale of the mediation analysis. The analysis and findings are still not related to aims 1 and 2, since DA and entropy are not part of the prediction models. But I appreciate the explanation that this part is related to the authors' previous work, and that the authors attempted to link to them somehow.

      We appreciate the reviewer's comment. The mediation analysis was not intended as a replication of our previous work, but rather as a mechanistic analysis to examine whether the association between BCG and cognition operates indirectly through brain age. Given the established links between BCG, brain age, and cognitive function, mediation analysis provides a principled framework for testing this hypothesis. Essentially, this mediation analysis supports previous theoretical model and our own empirical data where we showed lower DA contributes to more noise in FC metric, which in turn results in less accurate prediction (i.e., larger gap).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I hope the authors report CPM analysis as they claimed in the response letter, with actual statistical tests to compare the performance of DenseNet vs. CPM. It will be better to compare DenseNet with other models as well.

      We thank the reviewer for reiterating this point. As noted in both the manuscript and our previous response, the goal of the present study was not to benchmark DenseNet against alternative predictive frameworks, but rather to compare the predictive utility of different brain states using a consistent modeling approach. Although we explored CPM as a potential reference model, we found that its performance was sensitive to analytical choices, including feature-selection thresholds, summarization of positive and negative networks, and model specification. A rigorous CPM comparison would therefore require systematic optimization and validation of these choices, which would constitute a separate benchmarking analysis beyond the scope of the present study. We therefore deliberately avoided claims regarding methodological superiority and revised the manuscript accordingly. While comparisons with CPM and other machine-learning approaches would be valuable, we believe such analyses would constitute a separate benchmarking study beyond the scope of the present work.

      (2) "n-back did not significantly predict EM in DyNAMiC, and rest did not significantly predict WM. For this reason, we highlighted only the conditions that showed meaningful predictive power in the original analyses."

      These null results should also be reported; otherwise, this suggests the authors cherry-picked only the "meaningful" results, making the claims overly optimistic.

      We thank the reviewer for this comment. We agree that reporting only significant findings could create the impression of selective reporting. However, this was not our intention. All within-dataset prediction results, including both significant and non-significant findings, are reported in Tables 1 and 2. For the cross-dataset analyses, we chose to evaluate only the best-performing models identified in the DyNAMiC dataset. This decision was made a priori to test whether the most reliable predictive models generalized to an independent dataset, rather than to maximize the number of significant findings. For example, resting-state FC emerged as the strongest predictor of episodic memory in DyNAMiC and was therefore selected for external validation in COBRA. Similarly, the movie-watching model showed the strongest performance for working memory and was consequently carried forward to the cross-dataset validation. We have clarified this rationale in the manuscript.

      (3) I appreciate the correlation plots between BCG and physical activity and cardiovascular risk. The results are much weaker in COBRA (r = .17 and -.10 vs. .40 and -.27 in DyNAMIC). Perhaps this warrants discussion. Note that there are potential outliers in DyNAMIC. Perhaps the authors might like to include Spearman's rank.

      We thank the reviewer for this helpful observation. We agree that although the associations between BCG and both physical activity and cardiovascular risk were statistically significant in COBRA, the effect sizes were smaller than in DyNAMiC. To assess whether the observed association was influenced by potential outliers, we repeated the analysis using Spearman’s rank correlation. The association remained the same after controlling for age using partial Spearman rank correlation - between GAP and physical activity (DyNAMiC: r =0.40, p =0.001; COBRA: r =0.17, p=0.03) and for GAP vs. CVD risk score (DyNAMiC: r =–0.27, p =0.03; COBRA: r = –0.10, p =0.40). Nevertheless, we have now tempered the interpretation in the Discussion (P.12) to clarify that the direction and significance of the associations were consistent across datasets, but that the magnitude of the effects was weaker in COBRA. This difference may reflect cohort differences, including age range, sample composition, and differences in how physical activity was assessed.

      (4) Yes, adding figures comparing BCG and BAG in the main text would be helpful, given BAG's popularity.

      We have added additional findings as a supplementary figure.

      (5) The authors should provide this reason as a justification in the method: "We initially attempted to predict both episodic memory (EM) and working memory (WM). However, EM prediction was only reliable within and across samples for the resting state, whereas WM prediction generalized most strongly from the movie-watching condition. Because COBRA does not include a movie-watching paradigm, we could not evaluate WM prediction across datasets. For this reason, we focused on EM when examining the brain-cognition."

      We thank the reviewer’s suggestion. This clarification is included in the method section of the revised manuscript [P 21].

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study is built on the emerging knowledge of trained immunity, where innate immune cells exhibit enhanced inflammatory responses upon being challenged by a prior insult. Trained immunity is now a very fast-evolving field and has been explored in diverse disease conditions and immune cell types. Earhart and the team approached the topic from a novel angle and were the first to explore a potential link to the complement system.

      The study focused on the central complement protein C3 and investigated how its signalling may modulate immune training in alveolar macrophages. The authors first performed in vivo experiments in C57BL mouse models to observe the presence of enhanced inflammation and C3a in BAL fluid following immune training. These changes were then compared with those from C3-deficient mice, which confirmed the involvement of C3a. This trained immunity was further validated in ex vivo experiments using primary alveolar macrophage, which was blunted in C3-deficiency, and, intriguingly, rescued by adding exogenous C3 protein, but not C3a. The genetic-based findings were supported by pharmacological experiments using the C3aR antagonist SB290157. Mechanistically, transcriptomic analyses suggested the involvement of metabolism-linked, particularly glycolytic, genes, which was in agreement with an upregulation of glycolytic flux in WT but not C3-deficient macrophages.

      Collectively, these data suggest that C3, possibly through engaging with C3aR, contributes to trained immunity in alveolar macrophages.

      Strengths:

      The conclusions reached were well supported by in vivo and ex vivo experiments, encompassing both genetic-knockout animal models and pharmacological tools.

      The transcriptomic and cell metabolism studies provided valuable mechanistic insights.

      We thank the reviewers for acknowledging the importance of the work.

      Weaknesses:

      For the in vivo experiments, the histopathological and other inflammatory markers (Figure 1) were not directly linked to alveolar macrophages by experimental evidence. Other innate immune cells (eg. dendritic cells, neutrophils) and endothelial cells could also be involved in immune training and contribute to the pathological outcomes. These cells were not examined or mentioned in the study.

      We agree with the suggestions from Reviewer 1 that other cell types such as dendritic cells, neutrophils, and endothelial cells can also be involved in immune training. As the focus of this study was on alveolar macrophages, we specifically focused on these cell types. However,

      (1) We have re-analyzed a recently published dataset of human volunteers who received aerosolized BCG exposure compared to saline. We observe that by Day 7, aerosolized BCG exposure alters the expression of C3 and C3aR1 in alveolar macrophages in the human bronchoalveolar lavage (BAL) fluid, compared to saline. We have included this new analysis in a revised Figure 1 to clarify why our focus is on investigating the C3-C3aR1 axis in alveolar macrophages.

      (2) We have conducted new experiments where we train the mice in vivo, collect the alveolar macrophages, and then provide the second stimulus ex vivo. We observe a similar phenotype in the in vivo trained, ex vivo stimulated alveolar macrophages. We have included this new data in a new Figure S2.

      (3) We have updated our Discussion to state “However, we also acknowledge that other immune cells such as dendritic cells and neutrophils, and non-immune cells such as epithelial cells, endothelial cells and fibroblasts can also be involved in immune training (Bigot et al., 2025; Friščić et al., 2021; Moorlag et al., 2020).”

      (2) For the ex vivo experiments assessing immune training in alveolar macrophages, only the release of selected inflammatory factors were measured. Macrophage activities constitute multiple aspects (e.g. phagocytosis, ROS production, microbe killing), which should also be considered to better depict the effect of trained immunity.

      We agree with the reviewer and have conducted additional experiments to assess immune responses influenced by training in alveolar macrophages. Specifically, we show that in addition to impairing the release of proinflammatory cytokines such as TNFα and IL-6, C3-deficient alveolar macrophages exhibit significantly lower phagocytosis and ROS production compared to WT alveolar macrophages post-training with heat-killed Pseudomonas aeruginosa. Results from these additional experiments have been included in new Figure S2.

      (3) The proposed mechanism of C3 getting cleaved intracellularly and then binding to lysosomal C3aR needs to be further supported by experimental evidence. 

      The mechanism of C3 being cleaved intracellularly involves serine protease-dependent cleavage of C3 to C3a and has been experimentally demonstrated previously (Liszewski et al. Immunity 2013; Elvington et al. J Clin Invest 2017). A prior report demonstrated that intracellular C3a interacted with a lysosomal C3aR to promote CD4<sup>+</sup> T cell survival (Liszewski et al. Immunity 2013). Based on the reviewer’s suggestions, we performed confocal microscopy on alveolar macrophages. Although we clearly observed intracellular colocalization of C3a (using a monoclonal antibody to the neo-epitope) with C3aR, we observed only some colocalization with LAMP1, a lysosomal marker (see Author response image 1). Hence, we will refrain from making comments on how C3 binds to lysosomal C3aR intracellularly in alveolar macrophages, as this may be cell type-specific or stimulation-specific. We have now revised the sentence in the manuscript to remove any references to lysosomal C3aR and now state – “Upon internalization, C3 is cleaved to C3a (Elvington et al., 2017), binds to C3aR, and affects cytokine production in CD4<sup>+</sup> T cells (Liszewski et al., 2013)”. We have not incorporated the Author response image 1 in the main manuscript as we would like to explore this further to precisely define the subcellular localization of C3a-C3aR in alveolar macrophages, but have provided it for the reviewer to explain the basis of the rewording in the revision.

      Author response image 1.

      C3a-C3aR colocalization in mouse ex vivo cultured alveolar macrophages (mexAM). mexAMs were harvested and cultured as per the protocol from Gorki et al. (2022). Cells were incubated in a Millicell EZ Slide 8-well glass chamber slide overnight to allow for adherence, then fixed, permeabilized, and incubated with anti-C3a conjugated to AF555 (blue, Hycult HM1072), anti-C3aR conjugated to AF647 (red, Hycult HM1123), and anti-LAMP1 (green, Cell Signaling 99437) overnight at 4°C. Slides were washed 3X in PBS (5 min each) and mounted overnight at 4°C in ProLong Diamond Antifade Mountant with DAPI (white). Images were acquired on a Zeiss LSM 880 confocal microscope at 63X. At least 6 cells per condition imaged. Experiments were conducted in duplicate (technical replicates) and repeated (for biological replicates). Scale bar, 2 μm.

      (4) There was an absence of any validation in human-based models.

      We acknowledge that the observations need to be validated in human-based models. The focus of our manuscript is on training in alveolar macrophages. Unfortunately, we do not have access to an adequate representation of human alveolar macrophages for our ex vivo testing to account for individual-level variation in immune responses. We anticipate this work will form the basis of these future studies. In the interim, we re-analyzed a recently published publicly available dataset of human BAL specimens from human volunteers who underwent aerosolized BCG administration (Marshall et al. Nat Comm 2025). We observe an increase in C3 and C3aR1 expression at Day 2, which persists through Day 7 post-training with aerosolized BCG compared to aerosolized saline specifically in human alveolar macrophages. We have included this data in Revised Figure 1. We also validated C3 uptake in alveolar macrophages using precision-cut lung slices from human donors. We have included this additional data in new Supplementary Figure 3.

      Reviewer #2 (Public review):

      Earhart et al. investigated the role of the complement system in trained innate immunity (TII) in alveolar macrophages (AM). They used a WT and C3 knockout murine model primed with locally administered heat-killed P. aeruginosa (HKPA). Additionally, they employed ex vivo AM training models using C3 knockout mice, where reconstitution of C3 and blockade of C3R were performed. The study concluded that the C3-C3R axis is essential for inducing TII in macrophages in the ex vivo model. The manuscript is well-written and easy to follow. However, I have the following major concerns.

      (1) The secondary challenge to assess the reprogramming of innate cells in the BAL was conducted 14 days after the initial exposure to HKPA. However, no evidence is provided to confirm that homeostasis was re-established following the primary exposure. Demonstrating the resolution of acute inflammation is essential to ensure that the observed responses to the secondary challenge are not confounded by persistent inflammation from the initial exposure.

      We thank the reviewer for giving us an opportunity to clarify this point. We have now included additional data from the bronchoalveolar lavage fluid of these mice to show that the levels of protein leaked into the BAL, levels of proinflammatory cytokines (e.g., TNFα, CXCL1) and the neutrophils (all relevant to the acute phase of inflammation) were similar between the untreated and treated wildtype mice. This new data has been included in Figure S1.

      (2) In Figure 1D, cytokine production by BAL cells from WT and C3KO mice after HKPA exposure and LPS challenge is shown. However, it is unclear whether the reduced response in trained C3KO mice is due to a defect in trained immunity or an intrinsic inability of C3KO cells to respond to LPS. To clarify this, the response of trained C3KO cells should also be compared to untrained C3KO controls after the LPS challenge. This comparison is necessary to determine if the reduction is specifically related to innate immune memory or a broader impairment in LPS responsiveness. Such control should be included in all ex vivo training and LPS stimulation experiments as well.

      We thank the reviewers for their suggestions. We have conducted additional experiments and we observe no significant differences in the BAL cytokine levels between the wildtype and C3-deficient mice post-training in the absence of infection. This new data has been included in Supplementary Figure S1.

      Additionally, we came across several manuscripts, including a recent one in eLife as a part of this Series (Gu et al. Elife 2021; Zahalka et al. Mucosal Immunol 2022; Prevel et al Elife 2025) that have done in vivo training followed by an ex vivo challenge. Hence, we have conducted new experiments to compare the response of in vivo HKPA-trained wildtype (WT) and C3-deficient (C3KO) alveolar macrophages compared to untrained AMs after an ex vivo LPS challenge. This new data has been included in Figure S2.

      (3) The data presented provide evidence of alterations in the functional and metabolic activities of innate cells in the lung, indicating the induction of innate immune memory in a C3-C3R axis-dependent pathway. However, it remains to be established whether such changes can lead to altered disease outcomes. Therefore, the impact of these changes should be demonstrated, for instance, through an infection model to support the claim made in the study that C3 modulates trained immunity in AMs through C3aR signalling.

      We acknowledge this is a Limitation of our manuscript. As this is a Short Report, we focused on how C3, via the C3aR, affects the reprogramming of alveolar macrophages. Recent work demonstrated that systemically administered β-glucan induces peripheral trained immunity and aggravates lung injury (Prével et al., 2025), similar to disease in models of periodontitis and arthritis (Haacke et al., 2025). However, training with β-glucan also reduces bleomycin-induced lung fibrosis (Kang et al., 2024). Hence, our ongoing work involves optimizing relevant intrapulmonary exposures to assess how trained immune responses are modulated by the C3a-C3aR axis. We have included the Reviewer’s critique in our revised Discussion as a limitation, while referencing the abovementioned manuscripts.

      (4) Figure 3, panels B and C - stats should be shown for comparing WT-HKPA-trained and C3KO HKPA-trained.

      These suggestions have been incorporated into Revised Fig 3B and 3C (now Figure 4).

      (5) In Figure 4, where the proper untrained C3KO is included, the data presented in Figure 4C show an increase in basal and maximum glycolysis in trained C3KO compared to their untrained control counterparts. Statistical analysis should be provided for this comparison. Based on these data, it appears that metabolic reprogramming occurs even in the absence of C3. Furthermore, C3KO cells intrinsically exhibit reduced glycolytic capacity compared to WT. These observations challenge the conclusions made in the manuscript. Therefore, without the proper control (untrained C3KO) included in all experimental approaches, it is impossible to draw an evidence-based conclusion that the C3-C3R axis plays a role in the induction of innate immune memory.

      We have included the statistical comparisons for all the groups in Figure 4C (now Figure 5), as suggested by the reviewer. The data suggests that C3-deficient (C3KO) alveolar macrophages have a blunted metabolic response to training, as compared to C3-sufficient (WT) alveolar macrophages. However, the C3KO cells do not have reduced glycolytic capacity compared to WT in the absence of training. The blunted response in trained C3KO AMs is rescued by exogenous C3, but is then reversed by C3aR antagonism. We have also provided new data/analyses with proper controls (untrained C3KO) in the other Figures (for example, in Figures S1, S2 and 3B&C (now Figure 4)). Taken together, the data would suggest that the effects of C3 in AM reprogramming are C3aR-dependent.

      (6) The Results and Discussion sections should be separated, and the results should be thoroughly analyzed in the context of published literature. Separating these sections will allow for a clearer presentation of findings and ensure that the discussion provides a comprehensive interpretation of the data.

      We thank the reviewer for this suggestion. The manuscript has been submitted as a Brief Report, and hence, we adhered to the instructions to authors for this format. However, we have added an additional section towards the end of the manuscript based on the Reviewer’s suggestion.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      It is intriguing that whilst there was a significant elevation of C3a in the cell culture medium of alveolar macrophages, the addition of exogenous C3a failed to rescue the phenotypes of C3-deficient alveolar macrophages. Please help discuss this.

      Alveolar macrophages secrete both C3 and proteases, which can cleave C3 to C3a in the supernatant. However, we used the addition of exogenous C3a to compare it to the addition of full-length C3. C3 is internalized by multiple cell types, including alveolar macrophages (as demonstrated in new Figure 3 of our Revised Manuscript) as C3(H<sub>2</sub>O) in comparison to C3a. Hence, we propose that the internalization of C3(H<sub>2</sub>O) provides an intracellular source of C3a (previously reported in Elvington et al. J Clin Invest 2017), which engages with the C3aR to result in alveolar macrophage reprogramming. In comparison, incubating cells with C3a does not exert similar effects. We have included these comments in a separate section towards the end of the manuscript.

      Please provide details for the statement "cell-permeable C3aR antagonist (SB290157)" (Figure 3E). Could a paracrine-based mechanism also be at play?

      SB290157 does not act selectively on the cell surface, but rather, can also enter cells. Our data, along with previously published reports (Quell et al. J Immunol 2017; Zha et al. Cancer Immunol Res 2019), suggest that C3aR may be intracellular in AMs. However, SB290157 can also block any receptor that may be present on the surface. For this reason, we used exogenous C3a as a way to interrogate surface C3aR signaling, and did not observe significant changes in AM reprogramming with exogenous C3a. However, as this is an indirect approach, we cannot completely rule out a paracrine-based mechanism and have included this limitation in the Discussion section of the revised manuscript.

      For Figure 3, please also provide the statistical analysis results for WT versus C3KO HKCA-trained cells. The statistical tests described in the legend for Figure 3D seem to apply to Figure 3E. Please check the labels.

      These suggestions have been incorporated into Revised Figure 3 (now Figure 4).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      We thank both reviewers for their thoughtful and constructive evaluations of our manuscript. We are pleased that the reviewers recognize the value of our multimodal single-cell approach and the diversity of model systems used to define gene regulatory networks underlying heterogeneity in Ewing sarcoma.

      Reviewer #1 (Public review):

      The investigators elegantly utilized a single-cell co-assay of RNA and ATAC seq to unveil the heterogeneous gene regulatory networks in Ewing sarcoma. The authors should be commended on their ability to identify multiple unique modules of gene regulation of Ewing sarcoma utilizing complex computational methods between numerous Ewing sarcoma cell lines. Additionally, they complemented their single-cell findings with xenografts as well as primary Ewing sarcoma patient tumors - validating the intratumoral heterogeneous gene regulatory networks of Ewing sarcoma. More importantly, they have revealed that exogenous TGF-β may modify these distinct epigenetic and transcriptional signatures within Ewing sarcoma tumors. Overall, the manuscript highlights an important discovery of the heterogenous gene regulatory programming of Ewing sarcoma and further highlights the role that TGFB plays within the tumor microenvironment of Ewing sarcoma. There are some areas of ambiguity that require clarification to increase the impact of the manuscript.

      We appreciate Reviewer 1's positive assessment of our work and their recognition of the importance of identifying heterogeneous gene regulatory programs in Ewing sarcoma, including the role of TGF-β in the tumour microenvironment. We have addressed the areas of ambiguity noted by the reviewer, including clarifying cluster assignments and the selection of k=3 for module identification, adding statistical comparisons to relevant figures, correcting figure cross-references, and improving figure labeling for clarity. We have also added higher-resolution images of the spatial profiling data and highlighted relevant correlations between CHLA9 and CHLA10 clusters. We believe these revisions improve the clarity and rigour of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      This work by Waltner et. al. provides a comprehensive single-cell multiomics analysis of plasticity in gene regulatory networks present in Ewing sarcoma using single-cell RNA-sequencing (scRNA-seq) and single-cell assay for transposase accessible chromatin with sequencing (scATAC-seq). They find that Ewing sarcoma cell line models have distinct patterns of chromatin accessibility compared to non-Ewing sarcoma models, and that there is significant variability across Ewing sarcoma cell lines, and sometimes within a single cell line. These differences across models are linked to 3 distinct gene regulatory modules, 2 of which are present across the range of model systems studied here. The first modules present across models are activated when the fusion is expressed and include genes enriched for the known EWSR1::FLI1 response element, GGAA microsatellites, along with other neural crest transcription factors. The other module primarily consists of genes repressed by EWSR1::FLI1, which are activated in EWSR1::FLI1-low states. Interestingly, EWSR1::FLI1-low cells have already been tied to more migratory and metastatic phenotypes, and the data here suggest these cells are more responsive to external signals from TGF-β, and this may be mediated through FOSL2-mediated gene regulation. While there are some minor additional validation studies that can be performed to strengthen a few individual analyses, this is a technically rigorous study, with a variety of different analytical techniques used to address similar questions, and this approach elevates confidence in the answers provided. This is further strengthened by the diverse set of model systems used, including patient-derived cell lines, cell line xenograft models, patient-derived xenografts, mining available single-cell data from patient samples, and validation of the gene modules identified in a larger set of patient microarray samples. In whole, this study provides a valuable resource for understanding heterogeneity, plasticity, and gene expression networks in Ewing sarcoma. This may be useful for future studies of metastatic disease and may also provide a framework for similar questions in other fusion-driven sarcomas.

      Strengths:

      There are a few core strengths in this study. First is the number and diversity of Ewing sarcoma models studied, spanning commonly used cell lines, patient-derived xenografts, and patient samples. The second is the large array of rigorous and orthogonal approaches used to uncover the identity and function of various gene modules. This includes an array of informatics techniques, as well as specific modulation of cell line models in culture. A third is confirmation that different gene expression programs are present in the same tumor using spatial transcriptomic analysis. Lastly, the authors have made all of their data and code accessible, enabling continued use of this dataset as a resource for others.

      Weaknesses:

      As highlighted by the authors, this study is somewhat limited by the small number of single-cell data from patient samples that are publicly available. Much of the analysis comes from cell lines. Additionally, they focus only on one type of signal that may modulate cell plasticity, and there are likely to be many others. Lastly, there are a few weak spots in the data. Some of this likely arises from the underlying complexity of the data, the generally sparse nature of scATAC data, and the biological heterogeneity present in the cell lines studied. The most pronounced weakness was in the analysis of transcription factors that dictate gene expression in the distinct modules, as well as the response to TGF-β. While some specific transcription factors showed module-specific expression consistent with the computational prediction in Figure 2, others did not likely due to additional factors not tested here. Likewise, the same transcription factors did not always show consistent enrichment in the gene modules that responded to TGF-β treatment when analyzed across cell lines. On the whole, these are relatively minor weaknesses and do not diminish the value of this study.

      We thank Reviewer 2 for their thorough and balanced assessment. We agree that the study's strengths lie in the breadth of model systems and orthogonal analytical approaches, and we appreciate the reviewer's acknowledgement that the identified weaknesses are relatively minor.

      In response to the reviewer's suggestions, we have made several substantive improvements. First, we have included a new Western blot panel in Figure 1 showing EWS::FLI1 protein levels across cell lines, which provides important context for interpreting the long-read fusion transcript detection data. Second, we have revised text throughout the Results section to improve precision — in particular, clarifying that our chromatin analyses assess accessibility at published EWS::FLI1 binding sites rather than binding per se, and ensuring that our stated hypotheses match the metrics presented in the corresponding figures. Third, we have improved figure color schemes and labeling to aid interpretation and corrected errors in panel labeling.

      Regarding the reviewer's observation about transcription factor enrichment patterns across cell lines (particularly RUNX3 in the TGF-β response analysis and FOSL2’s modest expression at the protein level in CHLA9), we acknowledge that not all TFs showed perfectly consistent module-specific behaviour across every cell line. As the reviewer notes, this likely reflects the underlying biological complexity and additional regulatory factors not tested here. We have made attempts to temper our language accordingly.

      We believe that the revised manuscript, with its additional experimental data, improved figures, and clarified text, addresses the concerns raised by both reviewers and strengthens the overall impact of our findings.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Specific comments:

      (1) Figure 1A: Suggest adding cell labels on the UMAP plot - difficult to tell with different colors - which cluster is which cell line.

      We have added labels to Figure 1A (UMAP plot) as suggested.

      (2) Figure 1B: Although this may have been slightly addressed later in the manuscript, since CHLA9 and 10 are from the same patient, are the authors surprised to see the differences in EWS::FLI1 motif accessibility, or are these findings further reinforcement of inherent heterogeneity? Should one expect the ChromVAR deviation z-score at least to overlap between CHLA9 and 10, since they're from the same patient, if not, can the authors explain why?

      We were indeed surprised to see differences in EWS::FLI motif accessibility inferred from our data particularly from the isogenic lines CHLA9 and 10. While it is impossible to know for sure, we suspect that some treatment related changes increased accessibility in CHLA10. However, when considering global accessibility signatures, a subset of cells from CHLA10 (C23) was most correlated with signatures from CHLA9. We have addressed this by stating “Unsurprisingly, of all cell lines, CHLA10 C23 cells exhibited the greatest correlation to profiles from CHLA9,” in the 11th paragraph of the results section.

      (3) Figure 1B: Can the authors comment on the differences in trends of fusion transcript to chromatin enrichment (e.g., A673 and TC71 inverse trends vs. other cell lines that are direct positive or negative trends).

      We are hesitant to draw too many conclusions from transcript counting vs chromatin enrichment given some outliers, but in general, we found that high-fusion transcribing cell lines A4573, SKNMC, RDES were associated with module 3, while the lower transcribing cells CHLA9/10, TC32 and PDX305 were associated with module 2. We did not choose to emphasize this correlation too strongly in the paper as there may be other factors such as additional mutations (such as BRAF<sup>V600E</sup> in A673) that may have either been present in the original tumour or after serial passaging that may play a role.

      (4) Figure S1D: Please make the pt. smaller in order to better visualize the differences and highlight the stark differences of the ChromVAR binding site score.

      We have adjusted the point size of the embeddings in Fig S1D to enhance visualization.

      (5) Figure S1D: Is it surprising to see a large proportion of PDX305 and minor proportions of CHLA10 and CHLA9 with low ChromVAR binding site score?

      Our analyses indeed show that PDX305, CHLA9, and CHLA10 utilise fusion-repressed gene programs, so lower enrichment of EWS: FLI1 microsatellites is consistent with our findings.

      (6) Figure 2A: Unclear how k = 3 was ultimately selected - was it purely a visualization of how well separated each of the cell lines is? Can the authors further clarify how k = 3 was determined to be the most optimal? Is there a UMAP that demonstrated a difference in clustering between groups 1, 2, and 3? Is there a threshold cut-off to assign groups into 1, 2, or 3?

      Thank you for pointing out this omission. In Fig 1 we show that unsupervised clustering of DA peaks across cell lines grouped EwS cell lines into 2 clusters. When using peak-2-gene linkages (Fig 2), we chose a k =3 to explore the potential genes/CREs explaining the grouping found in Fig 1 because A673 was a clear outlier. We have added the following sentence to the results section paragraph 6. “We selected k = 3 to extend the two-cluster structure observed in Fig. 1F–G, reasoning that A673 represented a clear outlier whose distinct regulatory program would be obscured at lower k.”

      (7) Figure 2G: Understanding the substantial heterogeneity of each cell and TF binding motifs - given that CHLA9 is considered group 2, is it unexpected that FOSL2 was not highly expressed compared to the other EwS cell lines within group 2: (TC32, CHAL10, PDX305).

      Thank you for astutely pointing out that CHLA9 did defy the trend for FOSL2 protein expression compared with other group 2 lines. We specifically did not comment on this finding in the manuscript as we could not account for its status as an outlier. We address this in the public response above.

      (8) Figure S3B: Did the authors perform a similar computational analysis (performed for Figure 4), looking specifically at CHLA9 (Clusters 24 and 25) to determine if there are any overlaps with CHLA10 cluster 23's pathway activity/MSigDB/GO: Biological process terms? If there are potential overlaps, can the authors potentially infer tumor clonal evolution from CHLA9 to CHLA10?

      While we didn’t do pathway analysis, Figure S3 shows strong pearson correlation of C23 from CHLA10 with both C24 and C25 from CHLA9. We have highlighted this with a red box in the figure to make this more obvious to the reader.

      (9) Figure 4: Given the heterogeneity of CHLA-10. Did the authors observe any differences in morphology within CHLA-10 between the predominant modules (modules 2 and 3), given such stark transcriptional heterogeneity?

      We did not observe any morphologic differences within CHLA10 but acknowledge this would be an interesting avenue of further investigation.

      (10) Figure 5B: Can the authors also plot out module 3 gene expression to see if the cluster enrichment is unique from module 2 gene expression with TGFB1 and vehicle?

      We thank the reviewer for this suggestion and have replaced Fig S4B with violin plots so the enrichments are clearer. Module 3 expression in cluster 3 cells from CHLA10 is among the lowest in the cell line.

      (11) Figure 5H: Can the authors provide a higher zoom/resolution of the H&E stain of the ROIs in order to see if there are indeed more stroma/fibrosis in ROI9 and ROI10, and if there are differences in tumor cell morphology within different ROIs that harbor different modules?

      We have now included higher resolution and magnification of IHC panels in Figure 5H. These images are included as Supplemental Fig 4C. It is notable that although no dramatic differences in tumour cell morphology are visualized, ROI9 and ROI10 comprise small islands of viable tumour surrounded by necrosis.

      (12) Figure 6: Within Volchenboum/Lawlor's dataset, of the 46 clinically annotated primary tumors, 10 of the COG samples contained substantial stromal elements, while all the tumors in the European cohort were >70% viable tumors. Can the authors separate out the stromal-rich (n = 10) samples and analyze the 36 tumor-enriched samples to see if the survival curve is the same as what is shown in Figure 6G/H and S4E?

      We thank the reviewer for this suggestion. We observed no differences when stratifying the patients as outlined. Indeed, in the original Volchenboum et al, manuscript it was demonstrated that no prognostic gene signature was identifiable when the stromal-rich tumours were removed from the cohort.

      (13) Figure 6: Additionally, can the authors run a similar computational analysis to determine the predominant modules within bulk sequencing of the Volchenboum/Lawlor dataset between each tumor?

      The computational approach deconvolution was developed for use with bulk RNA-seq and is based on count data. The linear mixture model at the heart of deconvolution methods requires that the measured signal is proportional to abundance across the full range, and Affymetrix microarray data violate that assumption in a gene-specific, nonlinear way that can't be fully corrected post hoc.

      (14) Page 15, line 2: Wrong GSE data cited - currently cited as GSE61357 - should be GSE63157 instead.

      This has been corrected, thank you.

      (15) Please add statistical comparisons for Figure 3E-H.

      We thank the reviewer for highlighting this omission. We used the software package ggpubr to perform Wilcoxon rank-sum tests comparing mean module scores between conditions. We have added statistical labels to the plots. All comparisons were statistically to the level indicated in the figure.

      (16) Page 11 Line 21: Figure S3 D-E is not about EMT, migration, and TGFB signaling - I believe the authors are referring to Figures 4D-E

      Thank you, we have corrected these errors

      Reviewer #2 (Recommendations for the authors):

      (1) Data

      (a) In Figure 1 and the associated text, there is an analysis of cells expressing EWSR1::FLI1 performed using a locus-specific amplification and long-range sequencing. On page 7, lines 1-5, there is some discussion about how some of these track with overall transcript levels, while others don't. Additionally, a very low fraction of cells is shown to be EWSR1::FLI1 positive. This analysis might also be strengthened by a Western blot to show protein levels and how they vary across the cell lines, which may help explain additional differences in the data. The Abcam antibody ab133485 works well for western blotting of EWSR1::FLI1. While this isn't additional single-cell data, the percentage of cells with transcript detected is not really equivalent to the total expression level. This seems particularly valuable to do, as prior publications (Pishas, et. al., Mol. Cancer. Ther., 2018) show that of the cell lines tested here, TC32 has relatively high EWSR1::FLI1 protein levels, while A673 has relatively low expression. This contrasts with the percent of positive cells here.

      We thank the reviewer for this suggestion and have included a new panel, Figure 1E containing the western results for cells harvested in log-phase growth.

      (b) Figures 5D-F are not obviously referenced in the section about Figure 5 (page 12 line 4 through pg. 13 line 20). One question to pose to the authors is how to interpret the data here for TC71 in light of the fact that they were relatively insensitive to TGF-β. This shows up obviously in the Western blot in 5G.

      Thank you, for drawing attention to this omission. We have corrected the figure reference to include Figure 5D-F. In regards to our interpretation of TC71’s lack of response to TGF-B, we direct the authors attention to our statement: “The muted response of the TC71 cell line (module-3 dominant) to the influence of TGF-b (Fig. 5A-C, Fig. S4B) suggests that pre-existing transcriptional states may condition the sensitivity to TGF-β signalling.”

      (c) For the discussion of Figure 5, the authors say on page 12, line 21, that "RUNX3 was enriched in non-responsive clusters." I'm not entirely convinced that the data support this statement as written. This is true in A673s, but there appears to be no difference between the maximally responsive and maximally non-responsive clusters in CHLA10. TC71 was not particularly responsive, but showed the opposite effect.

      We agree this was not clearly written. We have de-emphasized RUNX3 findings here. Our new conclusion in the results section paragraph 15 is: “Correlation analyses revealed a reciprocal enrichment pattern of many key TFs from Figure 2, where correlation of FOSL2 (but not RUNX3) gene expression and accessibility was generally highest in TGF-β responsive clusters.”

      Can data for A673 cells be included for Figure 5G? Like CHLA10, this cell line had the pattern of FOSL2 enrichment that is concordant with that described in the text.

      We thank the reviewer for this suggestion and have included A673 in the western blot.

      (2) Figures

      (a) In Figure 2A, it is very difficult to distinguish the colors for CHLA9 and TC32 in the left panel. Can the color scheme here be changed to make these easier to distinguish?

      We thank the reviewer for pointing this out and we have added labels to the UMAP. We hope this makes the visualization more obvious.

      (b) Similarly, the KLF4 and SP1 lines in Figure 2D are a little close and might benefit from having more distinct colors.

      We thank the reviewer for pointing this out. We have changed the KLF4 to a green hue.

      (c) The lower panel of Figure 5I has 2 samples labeled "11" and no sample labeled "12".

      Thank you for catching this error. It has been corrected.

      (3) Text

      (a) One section early in the response was a little bit confusing and could benefit from some revision to improve clarity. On page 6, lines 9-11, this reads a little bit like they looked at cell-line specific EWSR1::FLI1 binding, but that wasn't the assay that was performed. Perhaps there is a better way to describe this than simply "enrichment of EWS::FLI1 sites."

      We thank the reviewer for their efforts to improve clarity of our work.

      We changed the following sentence in the 2nd paragraph of the result section:

      "We visualized the enrichment of EWS::FLI1 sites across all cell lines and discovered distinct EwS cell-line specific usage (Fig. 1C & Fig. S1D)."

      And revised to:

      "We then assessed chromatin accessibility at these published EWS::FLI1 binding sites across all cell lines and discovered that accessibility at these loci varied in a cell-line-specific manner (Fig. 1C & Fig. S1D)."

      (b) Then at the start of the next paragraph (page 6 line 12), the authors talk about "differences in EWS::FLI1 motif and binding site enrichment" and my first thoughts were whether this was differences in which sites were bound or differences in the strength of enrichment. More precise language would be helpful.

      Here in the 3rd paragraph of the result section we changed:

      "Given the differences in EWS::FLI1 motif and binding site enrichment"

      To:

      "Given the heterogeneity in the magnitude of EWS::FLI1 motif enrichment"

      (c) Related to the comment about EWSR1::FLI1 positive cells vs. protein levels above, the hypothesis on page 6, line 13 says that there was a hypothesis that different Ewing sarcoma lines have different levels of EWSR1::FLI1 transcript. But the metric shown in Figure 2D is the percentage of cells with detectable transcript, not transcript levels. Be specific about what the data are showing here.

      We agree with the need to improve clarity. In the 3rd paragraph, we changed: "we hypothesized that EwS cell lines have different levels of EWS::FLI1 transcript."

      To:

      "We hypothesized that EwS cell lines differ in the proportion of cells expressing high levels of the EWS::FLI1 fusion transcript."

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript puts forward a statistical method to more accurately report the significance of correlations within data. The motivation for this study is two-fold. First, the publication of biological studies demands the report of p-values, and it is widely accepted that p-values below the arbitrary threshold of 0.05 give the authors of such studies justification to draw conclusions about their data. Second, many biological studies are limited by the number of replicate samples that are feasible, with replicates of less than 5 typical. The authors report a statistical tool that uses a permute-match approach to calculate p-values. Notably, the proposed method reduces p-values from around 0.2 to 0.04 as compared to a standard permutation test with a small sample size. The approach is clearly explained, including detailed mathematical explanations and derivations. The advantage of the approach is also demonstrated through analysis of computer-generated synthetic data with specified correlation and analysis of previously published data related to fish schooling. The authors make a clear case that this method is an improvement over the more standard approach currently used, and also demonstrate the impact of this methodology on the ability to obtain p-values that are the standard for biological research. Overall, this paper is very strong. While the subject matter seems somewhat specialized, I would make the case that this will be an important study that has broad general interest to readers. The findings are very general and applicable to many research contexts. Experimentalists also want to report accurate p-values in their work and better understand how these values are calculated. Although I believe the previous statement is true, I am not sure that many research groups doing biological work are reading specialized statistics journals regularly. Therefore a useful and broadly applicable statistical tool is well placed in this journal.

      Strengths:

      The proposed method is broadly applicable to many realistic datasets in many experimental contexts.

      The power of this method was demonstrated with both real experimental data and "synthetic" data. The advantages of the tool are clearly reported. The zebrafish data is a great example dataset.

      The method solves a real-life problem that is frequently encountered by many experimental groups in the biological sciences.

      The writing of the paper is surprisingly clear, given the technical nature of the subject matter. I would not at all consider myself a statistician or mathematician, but I found the text easy to follow. The authors did an impressive job guiding the reader through material that would often be difficult to grasp. The introduction was also well-written and clearly motivated the goals of the study.

      We appreciate the reviewer’s summary of our study and its strengths.

      Weaknesses:

      A few changes could be made if the manuscript is revised. I would consider all of these points minor, but the paper could be improved if these points were addressed.

      (1) The caption of Figure 2 doesn't seem to mention panel D. Figure A-2 also does not mention C in the caption.

      We apologize for this error, and thank you for catching it! The figure legends had missing or incorrect panel labels. This error has been corrected.

      (2) Figure 2D is a little hard to follow. First, the definition of "Power" is not clear, and I couldn't find the precise definition in the text. Second, the legend for the different lines in 2D is only given in Figure A-2. Perhaps a portion of the caption for Figure 2 is missing?

      We have added a definition of power in the main text:

      “Although the permutation test, simultaneous permute-match test, and sequential permute-match test are all valid, they vary in power – the probability of detecting true dependence.”

      We have clarified the use of “power” in legend of Fig 2 and clarified that the color key for Fig 2D is in Fig 2A. The relevant excerpt of the Fig 2 legend is copied here:

      “(D) Statistical power for the permutation test and various permute-match tests as a function of the replicate number n, significance level α, and strength of dependence r<sub>X, Y</sub>. Power was estimated as the proportion of simulations in which dependence was detected, calculated from 5000 simulations at each value of r<sub>X, Y</sub> between r<sub>X, Y</sub> = 0 and 0.54 in steps of size 0.01. At r<sub>X, Y</sub> = 0, there is no dependence, so the curve at that point indicates the false positive rate rather than power. We chose the Pearson correlation coefficient as our correlation function ρ. See (A) for the color legend.”

      We have also added dotted lines connecting the legend in panel A to the curves in panel D.

      (3) The concept of circular variance for the fish data was heard to understand/visualize. The equation on line 326 did not help much. If there is a very simple picture that could be added near line 326 that helps to explain Ct and theta, that could be a big help for some readers who do not work on related systems. The analysis performed is understandable, the reader just has to accept that circular variance captions the degree of alignment of the fish.

      We have replaced references to circular concentration with “mean resultant length”, which is the standard jargon for this term in circular statistics, and we have added an illustration.

      (4) For the data discussed in Figure 3, I wasn’t 100% sure how the time windows were selected. In the caption, it says “time series to different lengths starting from the first frame”. So the 20 s time window was from t=0 to t= 20 s. Would a different result be obtained if a different 20 s window was chosen (from t = 4 min to t = 4 min 20 s just to give a specific example). I suppose by chance one of the time windows would give a pvalue less than the target 0.05, that wouldn’t be surprising. Maybe a random time window should be selected (although I am not indicating what was reported was incorrect)? A little more discussion on this aspect of the study may be helpful.

      As suggested by the reviewer, we have redone the analysis of Figure 3D with random segments. This provides a more complete picture of how the chance of detecting a significant correlation varies with segment length. The main conclusion is unchanged: Perfect match tests reliably detect dependence across a wider range of segment lengths than the naive parametric alternative.

      The relevant panel and an excerpt from the legend text are copied below.

      “(D) Permute-match tests detected a significant correlation between speed and alignment more consistently than the parametric test. For a grid of lengths between 20 and 600 seconds we sampled 500 random segments of each length, each drawn from the first 600 seconds, and determined for each segment whether the parametric test and/or the two possible permute-match tests detected a significant (p ≤ 0.05) correlation. In the edge case of the maximum 600-second length, all 500 “random” segments were identical.”

      Reviewer #2 (Public review):

      Summary:

      This paper presented a hypothesis testing procedure for the independence of two timeseries that was potentially suitable for nonlinear dependence and for small-sample cases. This should bring potential benefits for biology data.

      Strengths:

      The test offers good flexibility for different kinds of dependence (through adjusting \rho), and seems to have good finite sample performance compared to the literature. The justification regarding the validity of the test procedure is clear.

      We appreciate the reviewer’s summary of key aspects of our manuscript.

      Weaknesses:

      (1) The size of the test is not guaranteed to (asymptotically) equal \alpha, which may damage the power.

      We thank the reviewer for raising the issue of test size and power. We agree that a conservative test (one whose size can fall below alpha) may sacrifice power.

      Our objective is distribution-free false-positive rate (FPR) control. That is, we wish to keep the FPR at or below alpha for every distribution of X and Y, because in our regime (nonstationary time series with few independent replicates) the scientist often cannot verify distributional assumptions. Inspired by the reviewer’s comment, we now show (new Proposition 14) that the perfect match probability can be made arbitrarily close to 1/n<sup>!</sup>. As a consequence, any reported perfect match p-value below 1/n<sup>!</sup> would break the distribution-free validity of the test.

      A test that exploits distributional structure could likely access lower p-values; we have now explored how the empirical FPR of the permute-match test varies with the data-generating process (see our response to reviewer 2's recommendation 1 below).

      (2) The computational time can be an issue for a moderately large sample size when calculating the X / Y-perfect match. It will be beneficial to include discussions on the implementations of the test.

      We agree this is an important consideration. We have added the following text to the Discussion:

      “The test appears computationally tractable for relevant sample sizes: Our implementation of the permute match procedure completed a single test of dependence in the setting of Fig 2 with an average runtime of 3 seconds when n = 10 on a 2023 14-inch MacBook Pro with an M2 Pro processor and 16 GB RAM (see Source data 1). For n > 10, a standard permutation test already can report a p-value below 3 × 10<sup>−8</sup> so the perfect match test is likely unnecessary for typical applications.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      A few more minor notes/comments:

      As a personal preference, I like it when figures printed in grayscale retain their meaning (when possible). Just FYI, Figure 3D in grayscale is uninterpretable. Not saying a change is needed, just pointing it out.

      We changed Fig 3D to address a comment above, and think the new version better distinguishes between the permute-match and permutation test results in greyscale.

      On line 304, is that a lower bound or an upper bound? Maybe the issue is the probability mentioned on line 305 is not clear.

      As this point is not the main focus of the investigation, we have rephrased it to make it less technical and eliminate the issue of which bound is in question.

      “It seems likely that the tests could be further modified to report an even lower p-value when an X- and Y -perfect match occur simultaneously, as in Fig 3B. However, we have not investigated further and this problem is left for future efforts.”

      I am not sure if this is a weakness, but the p-value changing depending on the choice of whether to apply the X-perfect match or Y-perfect match test first is fascinating. The authors did discuss this very issue at several points in the manuscript. It is slightly unsettling to me that there isn't an exact p-value for a given set of data. This one point gives me a new perspective on statistics.

      In full transparency, I don't believe I have to background to thoroughly review the appendix. I did read through it and did not notice any errors, but I couldn't confidently say there are not any small mathematical errors or any logical flaws in the proofs. Some sections were not easy to follow (my own shortcomings, the writing appeared sufficient for more of an expert to understand).

      We greatly appreciate the reviewer’s time and effort tackling an appendix outside their comfort zone.

      Reviewer #2 (Recommendations for the authors):

      (1) In the numerical experiment session, the authors should include the null situation, i.e., the performance of the test when X and Y are independent. This helps assess the size of the test.

      We have added a section on size to our results section, copied below:

      “The permute-match test’s false positive rate depends on the process tested. The permute-match test is conservative – meaning that its false positive rate can fall below the significance level – because both the permutation test and perfect match test are conservative. As discussed elsewhere [29], the permutation test is conservative when α is not one of its possible p-values and when ties may occur between the original correlation and shuffled correlations. Checking for a perfect match is similarly conservative. The actual probability of a false-alarm perfect match event can vary depending on the process being tested. To see this consider the permute-match test in the setting where n = 3, where α = 0.05, and where r<sub>X,Y</sub>= 0 (independent X and Y). Note that in this case, obtaining a Y -perfect match (and thus p = 1/n<sup>n</sup>) is necessary and sufficient to detect dependence since α is too low for detection by either the permutation test or the p = 2/n<sup>n</sup> leg of the permute-match test. In the linear system of Fig 2, we observed among 5000 simulations a detection rate of 0.0148, significantly below the upper bound of 1/3<sup>3</sup> (Figure 2 - Source data 1; one-tailed exact binomial test, p < 10<sup>−20</sup>). Conversely, in the nonlinear system of Fig S2, this same event (Y -perfect match under n = 3 and r<sub>X,Y</sub> = 0) occurs with a detection rate of 0.0328, not significantly 9 below the upper bound of 1/3<sup>3</sup> (Figure S2 - Source data 1; one-tailed exact binomial test, p = 0.059). Thus, depending on the underlying process studied, the actual chance of a perfect match happening under independence may be near or significantly below the theoretical upper bound.”

      (2) Some insights regarding the choice of rho should be provided. Especially, are there any examples that the classical test, such as the Pearson correlation or Granger causality test does not work?

      We have redone the example of Appendix 4 with Pearson correlation, showing that Pearson correlation has substantially lower power than cross-map skill in this case (compare figures S2 and S3).

      (3) Line 34 - 35, page 2: Correlations and causality should be separately considered. This sentence talks more about causality rather than correlation.

      We appreciate the reviewer’s perspective and agree that correlation and causality are distinct.

      We feel that pointing out the issue of spurious correlations is helpful to orient our readers, especially those from a broad scientific audience. In the text, we define “correlation” as a descriptive statistic (rather than normalized covariance), and later distinguish it from “dependence”, which has causal implications due to Reichenbach’s common cause principle. We believe this distinction provides a useful backdrop for practitioners who use statistical methods but are perhaps new to thinking deeply about statistical dependence.

      (4) Please add some discussions on the situation that X_i depends on Y_{i - j} for some j > 0, which is associated with the setting of Granger causality test.

      We have added the following to the discussion:

      “No distributional assumptions are required, and the correlation function ρ can be completely arbitrary. For instance, ρ could include a lag to detect delayed dependence, or even evaluate the correlation strength at several lags and report the strongest among them [39].”

    1. Author response:

      We are glad the reviewers found the tandem fluorescent timer approach valuable and the leading-edge Arm stabilization finding significant.

      We agree with the Assessment that our evidence for three specific claims Wingless-independence, JNK-mediated regulation of Arm stability, and a direct mechanical/force-transmission role for stabilized Arm is currently incomplete, and we will revise the text throughout to reflect this more precisely rather than overstating the current data. In addition, we commit to two new experiments, both using existing reagents and fly stocks, that speak directly to the two most experimentally tractable points raised by the reviewers:

      (1) Re-staining our existing JNK-RNAi and JNK-overexpression embryos for E-cadherin (reagent already validated in Figure 4), to test whether JNK acts directly on junctional architecture or only indirectly, via broader epithelial disruption.

      (2) Imaging ArmTimer in a wingless loss-of-function background, to directly test Wingless-dependence of leading-edge Arm stabilization as the reciprocal of our existing overexpression data.

      We address each public review point below and outline the accompanying text revisions.

      On Wingless independence (Reviewer #1; Reviewer #2, Weaknesses #2 and #5; Recommendation #1):

      We agree that our current evidence unquantified Wg overexpression restricted to the amnioserosa (C381-Gal4) and uniform overexpression, alongside the absence of detectable Wg-stripe-associated Arm-Timer signal supports a more limited conclusion than "Wingless-independent" as currently stated. We will revise our language throughout the Abstract, Results, and Discussion to state that canonical Wg overexpression does not detectably enhance leading-edge Arm stabilization or perturb dorsal closure under our conditions, rather than asserting pathway independence. As noted above, we commit to imaging ArmTimer in a wg mutant background to test this directly, complementing our overexpression data with the reciprocal loss-of-function manipulation.

      On the related point that the Timer's failure to detect a Wg-stripe-associated stabilization signal could reflect a sensitivity limitation rather than a true absence of stabilization (Reviewer #2, Weakness #5): we agree and will state this explicitly rather than treating absence of signal as evidence of absence. This does not undermine the positive leading-edge finding, which is not defined relative to the stripe comparison: all embryos, channels, and time points were imaged and rendered using identical laser power and brightness/sensitivity settings, and the leading-edge RFP signal clearly exceeds background under those same acquisition conditions. We also note that detection limits of this kind are a recognized challenge for endogenously tagged reporters of canonical Wnt/β-catenin signaling generally, including in mammalian systems, and cite two studies already in our bibliography that report the same class of limitation: de Man et al. (2021, eLife 10:e66440) and Ambrosi et al. (2022, eLife 11:e64498). We have added this clarification, with these citations, to the Results (paragraph describing Figure 3).

      On the mechanistic link between Arm stability and destruction-complex activity (Reviewer #1):

      We agree that because both ΔArm and Axin overexpression converge on the destruction complex, our data cannot yet fully separate "escape from degradation" from "impaired α-catenin/junctional coupling" as the operative mechanism. We will revise the Discussion to state this ambiguity explicitly and will treat the ArmTimer-AA result (partial α-catenin-binding disruption via phosphosite mutation, independent of destruction-complex regulation) as the strongest current evidence isolating the junctional-coupling mechanism. We also thank Reviewer 1 for pointing us to Pokutta, Choi, Ahlsen, Hansen & Weis (2014, J Biol Chem 289:13589-13601), which structurally and thermodynamically characterized the mammalian cadherin·β-catenin·α-catenin complex and showed that α-catenin binding to β-catenin is a distinct, allosterically regulated interface cadherin binding increases β-catenin's affinity for α-catenin roughly 10-fold, and α-catenin homodimerization independently competes with β-catenin binding. We have added this citation to the Discussion as structural support for treating cadherin engagement, α-catenin coupling, and destruction-complex regulation as mechanistically separable interfaces, and note that the crystallized β-catenin·α-catenin interface provides a structural template for future experiments for example, structure-guided point mutations at the homologous interface residues in Arm, or in vitro binding assays comparing wild-type and threonine-mutant (T111A/T121A) Arm affinity for α-catenin.

      We do not, however, believe a destruction-complex-independent stabilizing allele of Arm is a tractable experiment to close this gap directly: any allele that stabilizes Arm without engaging the destruction complex is, by definition, a Wnt pathway gain-of-function allele, since destruction-complex-mediated degradation is the very regulatory step that canonical Wnt signaling controls. Nor would restricting the allele to a transcriptionally inactive form of Arm cleanly resolve the confound: Wnt/TCF target loci include dedicated repressive TCF-binding sites (Blauwkamp, Chang & Cadigan, 2008, EMBO J 27:1436-1446), so a transcriptionally "dead" stabilized Arm could still alter transcription by disrupting TCF-mediated repression. We therefore treat this as a genuine, currently unresolvable confound of the overexpression approach, and rely instead on the CRY2 optogenetic and ArmTimer-AA results as the strongest available evidence isolating a junctional-coupling contribution. We have added this reasoning, with both citations, to the Discussion.

      On JNK acting on Arm directly vs. indirectly (Reviewer #1; Reviewer #2, Weakness #3 and Recommendation #2):

      This is the most actionable point raised by both reviewers. As noted above, we commit to re-imaging and re-staining our existing JNK-RNAi and JNK-overexpression embryos for E-cadherin to determine whether junctional/polarity architecture is broadly disrupted under these conditions (indirect mechanism) or whether E-cadherin localization is comparatively preserved while Arm stabilization is specifically altered (direct mechanism). In the meantime, we note that a direct mechanism is biochemically plausible: in mammalian cells, JNK phosphorylates β-catenin directly and regulates adherens junction integrity, and JNK activity separately controls the binding of α-catenin to the junctional complex (Lee, Koria, Qu & Andreadis, 2009, FASEB J 23:3874-3883; Lee, Padmashali, Koria & Andreadis, 2011, FASEB J 25:613-623). We cite these as precedent that a direct route from JNK to junctional β-catenin/α-catenin regulation exists in another system, while being explicit that this does not establish the same mechanism in Drosophila dorsal closure that will be tested directly by the E-cadherin re-staining experiment. We have added these citations and this caveat to the Discussion.

      On the Dsh-DEP-to-JNK mechanistic link (Reviewer #1, Public Review #2 and Recommendation #5):

      We agree that our data show the Dsh-DEP requirement and the JNK requirement for dorsal closure as parallel, independent findings rather than a demonstrated linear pathway in our system. To provide context for why we consider a DEP-to-JNK connection a reasonable working hypothesis, we searched the literature in both Drosophila and vertebrates and will cite six additional studies establishing this link: Axelrod et al. (1998) and Axelrod (2001), establishing that DEP-dependent membrane recruitment and unipolar localization of Dishevelled are specifically required for planar polarity signaling, distinct from Wingless signaling; Paricio et al. (1999) and Fanto et al. (2000), showing Dishevelled acts through Misshapen and Rac1/RhoA to the same JNK module used in dorsal closure; and Moriguchi et al. (1999) and Yamanaka et al. (2002), showing biochemically in vertebrates that the DEP domain of Dvl-1 selectively activates JNK independent of β-catenin/TCF-LEF activity, and that this JNK requirement is conserved in Xenopus convergent extension, the vertebrate process most functionally analogous to dorsal closure. We will state explicitly that this precedent, while now cross-species, comes from planar-cell-polarity and convergent-extension assays rather than dorsal closure itself, so it supports the plausibility of a Dsh/Dvl-DEP-to-JNK connection without establishing that the identical pathway operates in our system.

      On the Dsh DIX/DEP domain-separability argument (Reviewer #1, Recommendation #4):

      We thank the reviewer for pointing us to Gammons, Renko, Johnson, Rutherford & Bienz (2016, Mol Cell 64:92-104), which showed that the Wnt signalosome itself is assembled by head-to-tail DEP domain swapping between Dishevelled molecules, and that this DEP-dependent oligomerization is directly required for canonical Wnt pathway activity not restricted to the non-canonical/planar-polarity branch as we had implied. We agree this evidence undercuts our previous interpretation of the DshΔDEP dorsal closure phenotype as evidence for a strong non-canonical/polarity-specific role for the DEP-dependent branch of Dsh. We have revised the Discussion accordingly: we now state that the DEP domain is required for the morphogenetic program culminating in dorsal closure, cite Gammons et al. directly, and note that this requirement does not by itself establish a non-canonical/polarity-specific role, since we cannot rule out a contribution from DEP-dependent canonical Wnt signalosome assembly.

      On force transmission (Reviewer #2, Weakness #1):

      We agree that we have not directly measured force or tension at the leading edge, and that our current data (colocalization with actin/E-cadherin, and functional requirement shown via CRY2 optogenetics and mutant analysis) are consistent with, but do not directly demonstrate, a role in force transmission. We do not have the in-house expertise to perform direct force/tension measurements (e.g., laser ablation, junctional tension assays), so we will not be adding such an experiment in this revision. Instead, we have revised the language throughout the manuscript including two Discussion section headings that previously stated a mechanical role for stabilized Arm as established fact to consistently present the mechanical/force-transmission role as a hypothesis raised by our data, not a demonstrated conclusion, and we retain a clear statement that direct force measurement (ideally in collaboration with groups with the relevant biophysical expertise) is future work rather than a claim we are making in this manuscript.

      On the phosphomimetic threonine mutant (Reviewer #1, Recommendation #2):

      ArmTimer-AA (T111A, T121A) was generated with the expectation that the tyrosine phosphosite mutants (ArmTimer-EE, ArmTimer-FF) would be the primary drivers of any dorsal closure phenotype, given their proposed role in E-cadherin binding; the pronounced zippering defect we observed in ArmTimer-AA was therefore an unanticipated finding rather than a predicted result. We have not generated the reciprocal phosphomimetic ArmTimer-EE(Thr) (T111E, T121E) allele. Generating and characterizing this allele is a substantial undertaking we estimate over a year including allele generation, validation, and phenotypic characterization and we will state this explicitly in the Discussion as planned future work rather than part of the current revision.

      On confirmation of myristoylated-Dsh membrane targeting (Reviewer #1, Recommendation #3):

      We cannot confirm that myristoylation localizes all Dsh protein to the membrane. However, this strategy has extensive prior genetic validation using the identical Src-derived myristoylation sequence: it was originally used to tether Armadillo and shown sufficient for constitutive Wnt pathway activation (Zecca, Basler & Struhl, 1996; Tolwinski & Wieschaus, 2001, 2004), and the same approach was subsequently applied to GSK3 and Dishevelled, in each case producing the expected pathway-activation phenotypes (Mannava & Tolwinski, 2015; Kaur et al., 2017). We have added these citations to the Results where the Myr-Dsh constructs are introduced.

      On overexpression-based perturbations versus endogenous regulation (Reviewer #2, Weakness #4):

      We would like to clarify that most of the Arm alleles used in this study including all of the point-mutant Timer alleles (ArmF1a, ArmTimer-FF, ArmTimer-EE, ArmTimer-AA) central to our mechanistic conclusions were generated as knock-ins at the endogenous ‘arm’ locus via MiMIC/RMCE, not overexpressed. The two exceptions are ΔArm and ArmS56A, expressed from UAS constructs because both are gain-of-function alleles anticipated to be lethal if expressed from the endogenous locus, based on prior experience with similarly stabilizing mutations. Axin overexpression was used because no Axin mutant or knock-in allele was generated for this study; we agree an endogenous Axin allele would be the ideal complement and will state this explicitly as a limitation, while noting that our CRY2 optogenetic perturbations of Arm and α-catenin which act acutely on the endogenous proteins provide an orthogonal line of evidence supporting the same conclusions. We have added this clarification to the Discussion.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the paper, the authors propose a new RNA velocity method, TSvelo, which predicts the transcription rate linearly based on the expression of RNA levels of transcription factors. This framework is an extension of its recent work TFvelo by including unspliced reads and designing a coherent neuralODE framework. Improved performance was demonstrated in six diverse datasets.

      Strengths:

      Overall, this method introduces innovative solutions to link cell differentiation and gene regulation, with a balance between model complexity (neuralODE) and interpretability (raw gene space).

      Comments on revised version:

      The authors have added comprehensive analyses in this revision, and all of my concerns have been very well addressed. Here, I just want to re-emphasize the original points 1 and 3.

      (1) The analysis and clarification are very helpful - thanks! I found that Fig. R1 and R2 are very insightful, as DoRothEA-only returns much worse performance. Please consider adding these two figures to the supp figure and possibly highlighting your setting for edge pruning (down-weights); therefore, the model is more likely to be affected by false negatives than false positives in the TF-target prior.

      We thank the reviewer for the positive feedback and for recognizing the value of the additional analyses. We have added the previous Fig. R1 and Fig. R2 to the Supplementary Information as Fig. S13 and Fig. S14, respectively, and have referred to them in the revised manuscript.

      We have also expanded the description of the TF–target prior used in TSvelo in the “Acquiring Prior Knowledge of Gene Regulatory Relations” subsection of the Methods. As noted by the reviewer, TSvelo is expected to be less sensitive to false-positive TF–target interactions because unsupported edges can be down-weighted during training. In contrast, missing true regulatory interactions are not represented in the prior network and therefore cannot contribute to the learned regulatory dynamics, making the model potentially more sensitive to false negatives.

      (3) Please consider adding some discussion on the challenges in capturing cell cycle transitions.

      We thank the reviewer for this suggestion. We have added a brief discussion in the Discussion section on the challenges of modeling cell-cycle transitions. In particular, cell-cycle progression is often characterized by cyclic dynamics and overlapping transcriptional programs, which can complicate the inference of directional state transitions and regulatory relationships.

      Reviewer #3 (Public review):

      Despite the abundance of RNA velocity tools, there are still major limitations, and there is strong skepticism about the results these methods lead to. In this paper, the authors try to address some limitations of current RNA velocity approaches by proposing a unified framework to jointly infer transcriptional and splicing dynamics. The method is then benchmarked on 6 real datasets against the most popular RNA velocity tools.

      Comments on revised version.

      The Authors addressed all my comments suitably. I'd like to thank them for the time they spent addressing them: the revised paper is much more convincing.

      I have 2 very minor follow-up concerns:

      (1) I appreciated the simulation study, however, no null simulation is present.

      We know RNA velocity tools are inclined to provide false positives: trajectories even when the data doesn't have any.

      I'd be helpful to add null simulations where the data has no trajectories and see if methods erroneously identify any.

      We thank the reviewer for this helpful suggestion. We have added null simulations to evaluate TSvelo and baseline approaches on data without underlying dynamic structure. Specifically, we generated a null dataset including 200 genes and 600 cells by independently sampling spliced (S) and unspliced (U) counts, thereby removing any coherent transcriptional relationship between them.

      When applying scVelo and UniTVelo to this data, no genes passed the velocity gene selection step under the default likelihood-based filtering, and no velocity field could be obtained. We further tested TSvelo, Dynamo, and cellDancer on the same null data and observed that all three methods still produce trajectory-like patterns despite the absence of true dynamics (See Supplementary Information as Fig. S17).

      Including TSvelo, many RNA velocity and trajectory inference approaches assume that they are applied to datasets reflecting underlying dynamic biological processes. We agree that incorporating additional checks during preprocessing could help prevent applying velocity analysis to non-dynamic datasets. We have added this discussion to the revised manuscript.

      (2) Several of the novel analyses are only reported in the Supplementary material and only references in the main text (e.g., "A validation of TSvelo on simulated data is provided in Fig. S1 and Fig. S2 in the Supplementary Information."). This is pity!

      If allowed, I'd add some comments about the new analyses (simulations, computational benchmarks, etc...) also in the main text.

      We thank the reviewer for this suggestion. We agree that several analyses presented in the Supplementary Information provide important support for our conclusions. To improve their visibility, we have expanded the corresponding descriptions in the main text and briefly summarized the key findings of the relevant Supplementary Figures instead of only citing them. These revisions have been made for Fig. S1, Fig. S2, Fig. S10, Fig. S12, Fig. S13, Fig. S14 and Fig. S17. In particular, we have incorporated a summary of the simulation results at the end of the subsection “Estimate RNA Velocity with TSvelo” in the Results section, and added a discussion of the computational benchmarking analyses in the Discussion section. We hope these changes improve the accessibility of these results while maintaining a concise presentation of the main findings.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      I suggest the paper to undergo (very) minor revisions as detailed in the Public Review.

      Simone Tiberi, The University of Bologna

      We sincerely thank all reviewers for their thoughtful suggestions, which have helped improve the clarity and overall presentation of the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Seegren and colleagues demonstrate that in a mouse model of neonatal E. coli meningitis, loss of endothelial toll-like receptor 4 (TLR4) leads to a marked decrease in transcriptional dysregulation across multiple leptomeningeal cell types, a decrease in vascular permeability, and a decrease in macrophage abundance. In contrast, loss of macrophage TLR4 had less pronounced effects. Using cultured wild-type and TLR4knockout endothelial cells, the authors further demonstrate that TLR4-NF-κB signaling leads to reversible internalization of the tight junction protein claudin-5, establishing a potential mechanism of increased vascular permeability. Finally, the authors use RNA sequencing of wild-type and TLR4-knockout endothelial cells to define the TLR4dependent cell-autonomous transcriptional response to E. coli.

      Strengths:

      (1) The authors address an important, well-motivated hypothesis related to the cellular and molecular mechanisms of leptomeningeal inflammation.

      (2) The authors use model systems (mouse conditional knockouts and cultured endothelial cells) that are appropriate to address their hypotheses. The data are of high quality.

      Weaknesses:

      (1) The authors perform single-nucleus RNA-seq on dissected leptomeninges from control and E. coli-infected mice across three genotypes (WT, Tlr4MKO, and Tlr4ECKO). A major discovery from this experiment, as summarized by the authors, is: "Tlr4ECKO mice exhibited a global attenuation of infection-induced transcriptional responses across all major leptomeningeal cell types, as judged by the positions of cell clusters in the UMAP." This conclusion could be considerably strengthened by improving the qualitative and quantitative analysis.

      Thank you for this comment. We agree that the UMAP-based interpretation would benefit from additional qualitative and quantitative support. We have expanded the snRNA-seq analysis with additional images and supplemental figures (Figure 1 – figure supplement 3, Figure 1 – figure supplement 5, and Figure 1 – figure supplement 6). The first and third of these new supplemental figures show dot plots for each major leptomeningeal cell type, for each genotype, for the two experimental conditions (infected vs. uninfected), and for individual genes in three immune-related gene sets (NF-kB and TNF-α, JAK-STAT, and IFN-ɣ), providing a more explicit comparison of infection-induced transcriptional responses across genotypes. The second of these new supplemental figure shows principal component analysis (PCA) of the individual snRNA-seq datasets (one mouse per dataset) for each genotype and experimental condition, demonstrating that the observed transcriptional shifts are consistent across biological replicates. Finally, Figure 1 – figure supplement 7, which was included in the original submission, shows changes in the most up- and down-regulated genes (based on adjusted p-value or fold change) in endothelial and myeloid cells across individual mice and genotypes/conditions, further supporting the genotype-dependent effects at the level of individual animals.

      (2) The authors interpret E. coli infection-induced increases in leptomeningeal sulfo-NHSbiotin as evidence of compromised BBB integrity (i.e., extravasation from the vasculature) (Results, page 7), but another possible route in this context is sulfo-NHS-biotin entry from the dura across a compromised arachnoid barrier. The complete rescue in Tlr4ECKOs is strongly suggestive that the vascular route dominates, but it would strengthen the work if the authors could assess arachnoid barrier fidelity (e.g. via immunohistochemistry). At a minimum, authors should mention that the sulfo-NHS-biotin signal in this context may represent both vascular and arachnoid barrier extravasation.

      Thank you for this comment. We agree that our data cannot rule out leakage across the arachnoid barrier during infection. While the rescue observed in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice strongly supports a dominant vascular contribution, we acknowledge that the sulfo-NHS-biotin signal may reflect permeability at both the vascular and arachnoid barriers. We do not think there is a clear way to directly test this possibility functionally, since the arachnoid barrier appears intact by confocal microscopy. Subtle differences in barrier cell morphology might be detectable by electron microscopy, but this would not definitively address whether infection permits molecular passage across the arachnoid barrier. We have followed the reviewer’s suggestion and revised the Results section to reflect this interpretation. Specifically, we added the following: “The simplest interpretation of these data is that the site of sulfo-NHS biotin leakage is primarily vascular. However, we cannot exclude some contribution from increased arachnoid barrier permeability.”

      (3) The authors state that "deletion of TLR4 prevented both NF-κB nuclear translocation and Cldn5 internalization in response to E. coli (Figure 4A-D)" (Results, page 9). In Figures 4C and D, however, there is no indicator of a statistical test directly comparing the two genotypes. A comparison of within-genotype P-values should not be used to support a genotype difference (PMID: 34726155).

      Thank you for pointing out this omission. We have updated the figures so that the between-genotype p-values are shown for those panels (including this panel) that had not previously shown them.

      (4) In the first paragraph of the Results, the authors summarize the meningeal layers as (1) pia, (2) subarachnoid space, (3) arachnoid, and (4) dura, and then state "The second and third layers constitute the leptomeninges." This definition of leptomeninges seems to omit the pia, which is widely considered part of the leptomeninges (PMID: 37776854).

      Thank you for pointing out this error, which has now been corrected.

      (5) The Cdh5-CreER/+;Tlr4 fl/- mouse lacks TLR4 in all endothelial cells (i.e., in peripheral organs as well as CNS/leptomeninges), and, as the authors note, the periphery is exposed to E. coli. It would be helpful if the authors could comment in the Discussion on the possibility that peripheral effects (e.g., peripheral endothelial cytokine production, changes to blood composition as a result of changes to peripheral endothelial permeability) may contribute to the observed leptomeningeal phenotypes.

      Thank you for raising this point. We agree that peripheral responses could contribute to the observed leptomeningeal phenotypes in this model. We have added two sentences to the second paragraph of the Discussion to address this: “We note that these experiments do not distinguish between local vs. distal anatomic sources of LPS or downstream effector molecules, such as cytokines, that activate the leptomeningeal inflammatory response (Huang et al., 2021). Histologic observations of RFP-expressing E. coli in the brain, liver, and lungs, together with positive blood cultures, indicate substantial systemic dissemination in this model. Thus, the inflammatory responses of leptomeningeal cells likely reflect exposure to bacterial products and inflammatory mediators derived from both local meningeal and peripheral sources.”

      Reviewer #2 (Public review):

      Summary:

      The authors use a postnatal mouse model of E. coli bacterial meningitis and a mouse brain endothelioma cell line combined with cell-type-specific gene deletion to study the function of endothelial TLR4, a cell surface receptor that recognizes gram positive bacterial wall components, in the local leptomeningeal (LPM) response with a focus on endothelial barrier breakdown mediated by TLR4. Single-cell transcriptional profiling and imaging studies using whole-mount preps of the LPM support that LPM endothelial, CD206+ local macrophage and LPM fibroblast and arachnoid barrier cell inflammatory response and is abrogated in endothelial-specific KO of TLR4, pointing to a role for endothelial TLR4 in local LPM response. Culture studies using Bend3.1 cells (a mouse brain endothelioma cell line) support a direct role for TLR4 in the bacteria-mediated inflammatory response and in internalization of Cldn5 via the endosomal-lysosomal pathway, resulting in loss of barrier integrity

      Strengths:

      The local LPM cell response in meningitis and the role of specific LPM cells in inflammation and CNS barrier breakdown have not been extensively studied, despite ample evidence for primary immune response in the meninges in human patients and in animal models. The authors employ a robust, multi-model approach using both in vivo and in vitro models with cell-type-specific knockout to study the function of TLR4 in brain endothelial cell response. The authors nicely combine functional barrier assays with IF for junctional localization in their experimental design, and they delve into potential mechanisms of Cldn5 internalization using markers of endosomal-lysosomal pathway localization. The authors also describe a new type of barrier assay using a streptavidin-coated plate upon which barrier-forming cell cultures can be placted, this could be a very useful alternative or complement to other size-selective barrier assays and presumably could work for other barrier forming cells types, likely epithelial cells.

      Weaknesses:

      (1) There are no measures of bacterial burden in peripheral organs, blood, in the LPM or brain in the TLR4 endothelial cKO mice. Lack of TLR4 in endothelial cells could prevent bacterial 'access' into the LPM and brain, essentially preventing meningitis and leading to a lack of inflammatory responses in the LPM-located cells simply because there is no bacteria present. Bacteremia may also be reduced, as might inflammatory responses in peripheral organs with TLR4-deficient peripheral endothelium. Bacterial counts and inflammatory measures in peripheral organs and blood are important to better understand the mechanism(s) underlying the reduced inflammatory profile in LPM cells and no LPM endothelial breakdown in the Tlr4 endothelial cKO mice. In other words, does deleting TLR4 in EC protect against the development of meningitis by somehow blocking bacteria access to the LPM (this would be supported by low or no CFU counts in infected Tlr4 endothelial cKO) or is it what the authors appear to propose in Figure 1J that TLF4 in EC is the only cell responding to the bacteria to trigger the immune cascade in the LPM? More data is needed to resolve this, as this is a major claim of the paper.

      Thank you for this comment. We agree that it is important to distinguish whether the reduced inflammatory response in Cdh5-CreER; Tlr4CKO (Tlr4<sup>VEKO</sup>) mice reflects altered bacterial burden versus altered host sensing. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4CKO mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs of infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice in E. coli burden; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and the time of sacrifice 24 hours later (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid leptomeningeal cells does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (2) The authors look at the underlying cortical response (cerebral vasculature for ICAM and immune cells) but do not use markers that could identify microglia (Iba1), the primary resident immune cell (CD206 is not useful, at this stage, in perivascular macrophages that are extremely sparse in the postnatal brain). This would be important to better study the impact on CNS resident immune cell morphological activation.

      Thank you for this comment. In response, we have analyzed Iba1 staining in the cortex in infected vs. uninfected mice. This is shown in Figure 2 – figure supplement 3. These data demonstrate a several-fold increase in Iba1 immunostaining in infected compared to uninfected cortex, consistent with increased microglial activation in response to infection. There is no statistically significant difference between infected WT and infected Cdh5-CreER; Tlr4CKO mice in Iba1 staining in cortex.

      (3) The authors suggest that Cldn5 junctional localization is selectively disrupted upon bacterial exposure, mediated by TLR4 - they suggest this based on studying PECAM, GLUT1, ZO-1 and B-catenin (all normally junction or cell surface located in cultured Bend3.1) in relationship to Cldn5 localization (normally high) - it is possibly these are also impact by bacteria exposure (maybe through different mechanisms?) - a better measure would be to use the similar cyto/PM measure they do for Cldn5 in Fig. 4D and to evaluate this or to use intensity measurements.

      Thank you for this comment. As the reviewer noted, the analysis of Cldn5 localization with vs. without E. coli exposure and in WT vs. Tlr4KO bEnd.3 cells (shown in Figure 4B and D) – uses Cell Trace to partition the image into cytoplasmic vs. plasma membrane territories. For the analyses in Figure 5, we wanted to compare the localization (and potentially re-localization) behaviors of a variety of subcellular markers with the localization and re-localization of Cldn5 following E. coli exposure. By directly measuring the % overlap of the two immunostains, we get that data. We note that the goal of this analysis is to assess relative co-localization with Cldn5 rather than absolute subcellular partitioning of each marker. While this analysis could have been extended to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories, it is clear by visual inspection of Figure 5A-C that beta-catenin, ZO-1, and PECAM1 remain plasma membrane-associated with E. coli exposure, and GLUT1 goes from the part of the plasma membrane not involved in cell-cell contact without E coli exposure to cytoplasmic with E. coli exposure (as judged by the appearance of a nuclear “shadow” after E. coli exposure). Thus, we do not believe that additional cytoplasmic vs. plasma membrane quantification for these markers would alter the interpretation. The main reason that we did not extend this analysis to include independent quantification of the subcellular localization of each of those other markers with respect to cytoplasmic vs. plasma membrane territories is because that would introduce the Cell Trace localization as an additional variable.

      (4) The discussion could benefit from delving more into the prior literature on E coli mediated breakdown of junctions in cultured human microvascular brain endothelial cell model and critical host-pathogen interactions of the bacteria with ECs (PMID: 14593586), and how this might involve TLR4.

      Thank you for this comment. Two paragraphs addressing the prior literature have now been added to the discussion.

      (5) It would be important to discuss how their results relate to earlier studies on TLR4-/- and TLR2-/- global knockout mice and protection vs vulnerability to development of meningitis (see PMCID: PMC3524395) - this paper showed that TLR4 global KO mice have increased susceptibility to die from meningitis and have much higher CFU counts in the CNS. In this manuscript and their prior work (Wang et al., 2023), this group shown that both global TLR4-/- mutants and their EC-specific KO have reduced barrier permeability, but we don't have any information about CFU or susceptibility to death from meningitis in their models.

      Thank you for these comments. The model we use – subcutaneous injection of E. coli (a clinical isolate from an infant with meningitis) at postnatal day (P)5 – results in the death of the infected mouse within 2 days (shown in Figure 1 – figure supplement 3 in Wang et al. 2023). Our analyses of infected mice were conducted 24 hours after infection. As noted in the reply to comment #1, in the revised manuscript we present a clinical assessment of WT vs. Cdh5-CreER; Tlr4CKO mice 24 hours after infection based on (1) a quantitative microscopic analysis of E. coli burden in the brain (visualized based on RFP fluorescence in the E. coli used here), (2) quantifying CFUs in blood and (3) mouse weights at P5 and P6, a sensitive indicator of overall health since this is a time when mice are normally gaining weight rapidly (~25% weight gain per day). These data (shown in Figure 2 figure supplement 4) indicate that bacterial burden and disease severity are similar between genotypes in our model. In Wang et al., 2023, we did not conduct a quantitative clinical assessment of WT vs. Tlr4-/- mice following infection, but by visual inspection, infected WT and Tlr4-/- mice appeared to have similar downhill clinical trajectories. We have expanded the Discussion to relate these findings to prior studies of global TLR4 and TLR2 knockout mice, noting that differences in experimental models and the distinction between global versus VECadCreER-specific deletion may account for the differing outcomes reported.

      Comment on the paper listed by the reviewer (PMCID: PMC3524395).

      The cited study demonstrates that global TLR4 deficiency leads to increased bacterial burden and mortality, indicating an essential role for TLR4 in host defense and bacterial clearance. In our study of Cdh5-CreER; Tlr4CKO mice, bacterial burden and disease severity at 24 hours post-infection are similar between WT and Cdh5-CreER; Tlr4CKO mice, indicating that Cdh5-CreER; Tlr4CKO does not alter the clinical course at this time point. This difference is noted in the Discussion section.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the molecular underpinnings of immune responses in the leptomeninges in neonatal bacterial meningitis. Bacterial meningitis is a major disease burden, particularly for neonates, and it has previously been noted that the meningeal immune environment in infants is permissive to opportunistic infection (Kim et al., Sci Immunol, 2023). There is less known about the contribution of the stromal compartment to meningeal immune responses. Seegren et al. interrogate the role of leptomeningeal endothelium in host defence in E. coli infected neonatal mice using mouse genetic tools to delete the LPS receptor Tlr4 from either endothelial cells (using Cdh5-CreER) or macrophages (using LysM-Cre). The authors use snRNAseq, cleared cortical mounts, and in vitro work to define the impact of E. coli infection on leptomeningeal endothelial cells. This study uses a range of innovative techniques to probe the role of the stromal compartment in meningitis.

      Strengths:

      This study makes excellent use of cleared cortical mounts to examine the biology of the leptomeninges, in particular, changes to the endothelium, with unprecedented detail. In combination with high-quality sequencing data provide new insights into the impact of meningitis on the leptomeninges. The data presented by the authors is of very high quality.

      Weaknesses:

      The weaknesses of the study were in terms of interpretation and perhaps study design.

      (1) Most importantly, the authors need to provide additional validation of their conditional knockout models. The authors need to confirm that the Cdh5-CreER does not impact leptomeningeal fibroblasts and to confirm gene deletion in macrophages.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclearlocalized GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite correct. The Results section of the revised manuscript includes an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      (2) The authors could also strengthen the paper by providing data on the impact of these conditional knockout models on the course of meningitis and bacterial burden.

      Thank you for this comment. We agree that these additional analyses strengthen the manuscript. We have fleshed out these issues by conducting the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences in bacterial burden between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) Finally, it is perhaps not surprising that Tlr4 is required for meningitis responses with E. coli. However, it is unclear if these findings can be generalised to other, more common, meningitis infections (streptococcal/pneumococcal).

      At present, it is an open question whether TLR4 plays as a large a role in meningitis caused by other gram-negative bacteria and whether TLR2 plays a similarly large role in meningitis caused by gram-positive bacteria. In the Discussion, the last two sentences under “Limitations of the study” summarize this point: “Finally, the present study focused on E. coli K1, the dominant Gram-negative neonatal pathogen. Future work could assess TLR signaling in response to other bacterial pathogens, such as Group B Streptococcus.”

      (4) There are additional minor issues; for instance, the arachnoid fibroblast 2 population appears to closely resemble dural border cells.

      Thank you for this comment. That is correct, and we have changed the nomenclature to “dural border cells”.

      (5) The cell line model (bEnd.3) is a relatively low-fidelity model of BBB endothelial cells, and this should be acknowledged.

      Thank you for this comment. That is correct. Despite being brain-derived, bEnd.3 cells have lost many BBB-specific attributes. Their responses might best be considered as generic endothelial responses rather than brain-specific endothelial responses. This is now stated in the Results section: “Although they are brain-derived, bEnd.3 cells lack many BBB-specific attributes and, therefore, they likely exhibit generalized endothelial responses rather than brain-specific responses to bacterial exposure.”

      With these caveats, it is difficult to be certain that the endothelium alone is the driver of meningeal immune responses in meningitis, and what the impact of these is.

      We agree with this critique. As noted above, the expression of Cdh5-CreER in essentially all endothelial cells and in a subset of dural border cells and leptomeningeal fibroblasts means that the comparison of TLR4 CKO with Cdh5-CreER vs. Lyz2-Cre is assessing phenotypes driven by TLR4 signaling in endothelial plus a subset of other non-myeloid cells vs. TLR4 signaling in myeloid cells. We have revised the text to reflect this more precise understanding of Cdh5-CreER specificity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Transcriptomic analysis: The analysis and display of the single-nucleus RNA-seq data should be improved. The authors could perform a more granular, unbiased clustering of each cell class in the combined dataset and then compare the proportion of each experimental group (genotype x control/infected) in each cluster. At present, it appears the differentially-expressed genes (DEGs) shown in Figure 1 were identified using the Seurat FindMarkers function with default parameters (Methods). This considers each cell as an independent experimental unit and is therefore not appropriate for a comparison of control versus infected groups (see e.g., PMID 34584091, 35880426. The authors should implement a statistical analysis strategy that considers true biological replicates (mice, as shown in Supplementary File 1).

      We do not fully agree with this critique. We agree that biological replication at the level of individual mice is important for interpreting these data, but within each mouse, the characteristics of individual cells is also of interest, including the degree of heterogeneity, the sample size for a given cell cluster, and the statistical significance of any observed changes in transcript abundance. As requested, we have prepared a new supplemental figure (Figure 1 – figure supplement 5) showing a principal component analysis of the scRNA-seq data for each mouse (one mouse was used for each snRNA-seq dataset) and for each of the principal leptomeningeal cell types. This analysis shows, for example, that the three infected Cdh5-Cre; Tlr4flox/- mice have transcriptomes for each of the six cell clusters that are very similar to the transcriptomes of the two uninfected WT and the two uninfected Cdh5-CreER; Tlr4flox/- mice. Thus, the genotype- and condition-dependent effects are consistent across biological replicates. At the most granular level, Figure 1 – figure supplement 7, which was part of the original submission, shows for the most up- and down-regulated genes (based on adjusted p-value or based on fold-change) in endothelial cells and in myeloid cells how individual transcript abundances change for each mouse and for each genotype/condition.

      (2) The authors use immunohistochemistry to assess claudin-5 "disorganization and redistribution" (Results, pages 7-8 and Figures 3A-B). They state that "Tlr4ECKO mice showed minimal changes in the distribution of Cldn5, implying that cell autonomous endothelial TLR4 signaling regulates tight-junction organization." It is not clear, however, that the quantified parameter (Cldn5+ area relative to total area) would be an accurate readout of claudin-5 organization/distribution (i.e., subcellular localization) as it would also be sensitive to claudin-5 expression, vascular density, and vessel diameter. The authors use a similar assessment of ZO-1 to suggest that changes to claudin-5 are not due to a "generalized disassembly of TJs" and could also use this to argue that the above potential confounds (vascular density, vessel diameter) do not change, but the data in Figure 3 - Figure Supplement 1B, lower panel, show that infection does cause an increase in ZO-1 area relative to total area (P = 0.0004). Thus, the statement in the results "Zonula Occludens-1 (ZO-1) [...] remained unchanged during infection (Figure 3 - figure supplement 1)" is not accurate. The authors should revise this section to ensure their conclusions are aligned with the presented data.

      Thank you for this comment. The reviewer is correct that the Cldn5 area measurement is unable to deconvolve the various factors that might contribute to it (vessel density and diameter, and Cldn5 distribution). This part has been rewritten. “Consistent with prior findings (Wang et al., 2023), both WT and Tlr4<sup>MKO</sup> mice showed an increase in the area occupied by Cldn5 in the leptomeninges following infection, likely referable to both increased vessel diameter and a redistribution of Cldn5 within ECs (Figure 3A-B; Figure 3 – figure supplement 1C).”

      The reviewer is also correct about our initial description of the ZO-1 data. What we meant to write and what the revised manuscript now shows is: “The area occupied by Zonula Occludens-1 (ZO-1), a tight junction scaffold protein, showed a modest but statistically significant increase in WT leptomeningeal vessels but no significant change in Tlr4<sup>VEKO</sup> leptomeningeal vessels during infection (Figure 3 – figure supplement 1A and B).”

      (3) In the Methods, under Mouse Models and E. coli Infection, the authors state, "The Cdh5-CreER line (Monvoisin et al., 2023) was the same line used in Wang et al. (2023)." However, there is no mention of Cdh5-CreER in Wang et al. (2023). Could authors please clarify? Also, because this appears to be an inducible Cre, the authors must include details on the dose and timing of tamoxifen or 4-OHT used in this study.

      Thank you for catching that error. We meant to reference Wang et al (2025), not Wang et al (2023). [Wang et al (2025) is: Wang Y, Rattner A, Li Z, Smallwood PM, Nathans J. (2025) Vascular endothelial-specific loss of TGF-beta signaling as a model for choroidal neovascularization and central nervous system vascular inflammation. Elife 14:RP107018.] This has now been corrected.

      We have now included the details related to 4HT injection in the Methods section “Mouse Models and E. coli Infection”. These are intraperitoneal injection at P2 with 40- 50 µL of 2 mg/ml 4HT.

      (4) The legend for Figure 1A is "Schematic of the leptomeninges", but the figure shows the entire brain-skull interface, including underlying cortex, leptomeninges, dura, and skull.

      Thank you. Corrected.

      (5) Page 7, typo: "In the brain, CD206+ cell were too sparse ..." Should be "cells".

      Thank you. Corrected.

      (6) Page 14, typo: "... could represents a double-edged ..." Should be "represent".

      Thank you. Corrected.

      Reviewer #2 (Recommendations for the authors):

      (1) Perform CFU counts from LPM, dura, brain, peripheral organs (liver) in infected v mock mice from control v TLR4 EC-cKO.

      Thank you for this comment, with which we agree. We have addressed this by quantifying bacterial burden and assessing disease severity in WT and Cdh5CreER; Tlr4floxed mice. Specifically, we performed CFU measurements in blood, monitored mouse weights at P5 and P6, and histologically surveyed the E. coli-RFP signal (i.e., E. coli burden) in brain, liver, and lung. These analyses show that bacterial burden and disease progression are comparable between WT and Cdh5-CreER; Tlr4floxed mice at 24 hours post-infection. These data are presented in Figure 2 – figure supplement 4 and described in the Results.

      (2) Lyz2Cre/+ is used to delete TLR4 from macrophages, but recombination efficiency (in LPM BAMs) is described as only partial, suggesting that TLR4-response in LPM BAMs (and potentially macrophages in the dura) is at least partially intact. It undercuts conclusions that can be made using this line.

      Thank you for this comment. We have conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The Results section text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Also, as noted in the reply to comment 4 below, a direct analysis of Tlr4 recombination efficiency is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (3) Inflammatory responses [qPCR] from peripheral organs and also physiological measures in the pups [weight post-infection, time to moribund or death curves] in control v TLR4 EC-cKO and TLR4 mac-cKO.

      Thank you for this comment. We have not conducted a qPCR analysis of inflammatory gene expression in peripheral organs because (1) the dramatic upregulation of these transcripts in the leptomeninges, (2) the presence of E. coli in blood and peripheral organs, and (3) the clinical assessment (cessation of weight gain) all predict that such an analysis would reveal a large up-regulation of inflammatory gene expression throughout the body. More specifically, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (4) The conditional macrophage line is problematic due to the partial recombination. I question the utility of including this unless they can come up with a way resolve the response of recombined TLR4 macrophages vs ones that are not (could they use the single cell data to pick this a part? Are TLR4-null cells and TLR4 'wt' cells transcriptionally similar in the infected condition, suggesting TLR4 is not doing much in the macs, potentially due to alternate TLRs?). There are good BAM Cre lines that have been described [Lyve1-cre would be good for LPM BAMS, the other is Pf4-cre, see https://pmc.ncbi.nlm.nih.gov/articles/PMC7375817/ - just as an FYI for the future].

      Thank you for this comment. As noted in the reply to point 2 (above), we have conducted a more in-depth analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup> directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      We agree that, based on Figure 6 in the cited paper [McKinsey et al (2020) A new genetic strategy for targeting microglia in development and disease eLife 9:e54590], the Pf4-Cre line may be superior to the Lyz2<sup>Cre</sup> line that we used for recombination in leptomeningeal myeloid cells. Unfortunately, we missed this paper in our literature searches, probably because it focuses on a microglial CreER line, P2ry12-CreER, and the Pf4-Cre line is not mentioned in the title or abstract. Our decision to use the Lyz2<sup>Cre</sup> line was based on an extensive comparison among myeloid Cre lines showing that Lyz2<sup>Cre</sup> was the most efficient [Abram CL, Roberge GL, Hu Y, Lowell CA. 2014. Comparative analysis of the efficiency and specificity of myeloid-Cre deleting strains using ROSA-EYFP reporter mice. J Immunol Methods 408:89-100.] However, the Abram et al study did not look at the leptomeninges. Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4<sup>VEKO</sup> and Tlr4<sup>MKO</sup> mice.”

      (5) Figure 1 - Figure Supplement 2 - the authors nicely break down the pathway response [NFKB and TNF] in EC and macs, it would be great to have similar information for the fibroblasts (in the main figure or the supplement). Does their inflammatory response show a similar pattern?

      Thank you for this suggestion. We have now done that analysis and present it in Figure 1 – figure supplement 3. For completeness, we also performed the same type of analyses for JAK-STAT signaling and IFN-gamma response and these are shown in Figure 1 – figure supplement 6. The principal conclusion is that across all major leptomeningeal cell types, the Cdh5-CreER; Tlr4floxed samples (i.e., Tlr4 KO’d in non-myeloid cells) show much reduced transcriptome changes with infection.

      (6) What is ASC and Cd206 quantification measuring, and how does this relate to 'activation' - is this the intensity of signal or a morphological change? What is the precedence for using ASC (citations)? In their prior work, they showed no change in CD206 number, so a significant increase upon infection here, it's confusing exactly what is being studied. Also, loss of Lyve1 is a well-accepted measure of activation that they have previously used, adding that it could be helpful. This is not a major issue since they have robust data that the macrophages are not transcriptionally activated. Clarification of what exactly is being measured would be sufficient (in the text).

      CD206 immunostaining, which reveals myeloid cell morphology, shows that, with E. coli infection, myeloid cells convert from a more compact morphology to a more expanded morphology. This is now explained more fully in the Results section.

      Regarding ASC, changes in the state of ASC aggregation and ASC subcellular localization have been used by others to monitor immune cell responses to inflammatory signals (Sester et al., 2016; Franklin et al., 2018). While this change in subcellular localization may explain part of the increase in immunostained area in myeloid cells in the infected mice (Figure 2D), the increase in the area of ASC immunostaining largely reflects a shift of myeloid cells from a compact to a more extended morphology. This is now explained more fully in the Results section. We have also added two references (Sester et al., 2016; Franklin et al., 2018) that described how ASC distribution changes with inflammation.

      Regarding LYVE1, we observe a decrease in LYVE1 transcript abundance in myeloid cells with infection, as predicted. Given the large amount of other data that document myeloid activation with infection, we have elected not to include this.

      (7) The authors suggest the internalization of Cldn5 is not due to NFKB downstream signaling that includes transcriptional mechanisms because it happens as early as 1 hour, prior to NFKB localization to the nucleus. However, a lot of their experiments, including on endosomal-lysosomal protein co-localization are done at 4 hours, when their RNAseq data show robust NFKB-mediated gene upregulation and (though not tested) potentially protein production of factors that can act back on the cells, including to impact endo-lysosomal processing. Without studies at earlier timepoints post-bacteria exposure, separating these two mechanisms is difficult.

      Thank you for this comment. We have explored this question by looking at Cldn5 internalization in bEnd.3 cells at 1 hour after E. coli exposure, and the data clearly show that internalization occurs within 1 hour. Additionally, we have conducted this experiment in the presence of 1 uM ACHP, an IKK inhibitor that blocks NF-кB migration to the nucleus. ACHP treatment shows no effect on the rapid internalization of Cldn5, implying a mechanism independent of NF-кB control of gene expression. These data are shown in a new figure (Figure 6) in the revised manuscript.

      (8) Figure 2 - CD206 are quite sparse however, Iba1 would work well to look at microglial activation.

      Thank you for this suggestion, which we have followed. To assess microglial activation, we have immunostained for Iba1 and quantified the data. These are now included in Figure 2 – figure supplement 3. The data show that there is an increase in Iba1 immunostaining following E. coli infection in both WT and Cdh5-CreER; Tlr4floxed mice, with more in the former than the latter, but the difference is not statistically significant.

      (9) Suggest performing the LAMP+ co-localization experiment at <1hr, prior to NFKB nuclear localization and transcriptional changes. This would better support it, this is (or is not) independent of the NFKB. Could also test this with an NFKB inhibitor, do they still see the CLDN5 internalization when NFKB is blocked?

      Thank you for these suggestions. We have done both of these analyses, and the results are presented in Figure 6. The results show that (1) Cldn5 is internalized within 1 hour and (2) its internalization is independent of NF-кB signaling inhibition by 1 uM ACHP. Since ACHP treatment shows no effect on the rapid internalization of Cldn5, that implies a mechanism independent of NF-кB control for gene expression.

      Reviewer #3 (Recommendations for the authors):

      Major points

      (1) The most important caveat is that the Cdh5-CreER model is known to recombine in leptomeningeal fibroblasts (10.1038/s41586-023-06993-7, 10.1101/2025.05.13.653681), and Cdh5 expression in these populations is now well described (10.1038/s41467-02341580-4, 10.1016/j.neuron.2023.09.002). Although the authors did not observe recombination in their reporter (details of the tamoxifen injection protocol should be provided), it is imperative to validate the specificity of their model to Tlr4 in endothelial cells, leveraging their sequencing data and providing additional IHC or ISH to confirm this. Alternatively, Tlr4 could be deleted in a more specific model, e.g., the Pdgfb-iCreERT2 or Slco1c1-CreERT2. It is also important to do the same with the LysM model, to confirm that the lack of impact of macrophage Tlr4 is not due to failure to delete the gene. This is again important to the interpretation of the study, since the authors propose that the endothelium, specifically, is the driver of the meningitis response.

      We are very grateful for this critique. After several years of using the Cdh5-CreER line in other parts of the CNS, where its expression is endothelial-specific, we applied it to the meninges without realizing that its specificity is broader in that tissue. Our initial analysis with a Cre reporter line that uses a membrane tdTomato appeared to confirm endothelial-specific recombination in the meninges. Following receipt of the reviews of this manuscript, we repeated this analysis with two Cre reporter lines that use a nuclear-localised GFP, and we immunostained for each of several transcription factors to assess various meningeal cell types and quantified GFP co-localization (Figure 1 – figure supplements 1 and 2). This quantitative Cre reporter analysis shows CreER expression from the Cdh5-CreER transgene in all or nearly all endothelial cells and in a subset (~20%) of dural border cells and/or leptomeningeal fibroblasts, but not in myeloid cells. Additionally, our snRNA-seq analysis of Cdh5 transcripts shows expression in endothelial cells, dural border cells, and leptomeningeal fibroblasts, but not in myeloid cells (Figure 1– figure supplement 4), which agrees with several recent publications (Mapunda et al., 2023; Pietilä et al., 2023; Smyth et al., 2024). Thus, our initial interpretation that the phenotypes in the Cdh5-CreER; Tlr4floxed mouse were a consequence of recombination exclusively in endothelial cells was not quite right. The Results section of the revised manuscript has an expanded description of Cre and CreER expression specificity analysis, with supporting data in Figure 1 – figure supplements 1 and 2. Throughout the text of the revised manuscript, we are careful to note that the Cdh5-CreER; Tlr4floxed mouse has Tlr4 deletion in a subset of dural border cells and leptomeningeal fibroblasts. To reflect this fuller understanding of the specificity of Cdh5-CreER, we have changed the name of the Cdh5-CreER; Tlr4floxed mice in the text and figures from TLR4ECKO (“endothelial cell KO”) to TLR4VEKO (“VE-cadherin CreER KO”).

      We have also conducted a more detailed analysis of Lyz2<sup>Cre</sup> specificity by immunostaining for multiple markers and quantifying the results (Figure 1 – figure supplement 2). We now think that the more cursory analysis in the original submission was inaccurate. The more in-depth analysis shows that Lyz2<sup>Cre</sup>-directed Cre-recombination with 90-100% efficiency in CD206+ cells and with 50-70% efficiency in ASC+ and PU.1+ cells, the range depending on whether tdTomato or GFP colocalization was being scored (Figure 1 – figure supplement 2). The text in the Results section now states: “In the text that follows, we will refer to Lyz2<sup>Cre</sup>-recombined cells simply as “myeloid cells”, although they should be understood as CD206+ myeloid cells.”

      Regarding the efficiency of recombination of the floxed Tlr4 target, a direct analysis is technically challenging due to the low abundance of TLR4 and the failure, in our hands, of commercial anti-TLR4 antibodies to produce clear immunostaining. The phenotype of Cdh5-CreER; Tlr4floxed mice – a dramatically reduced infection-associated transcriptional response – argues that the floxed Tlr4 target was recombined at appreciable efficiency in those mice (Figure 1D and 1E). For Lyz2<sup>Cre</sup>; Tlr4floxed mice the principal phenotype is an up-regulation of infection-associated transcripts in a subset of dural border cells in the absence of infection; the transcriptional response to infection was largely unaffected in all leptomeningeal cell types (Figure 1D and 1E). We have added a comment in the results section noting that we do not have a measure of the efficiency of recombination of the floxed Tlr4 target in vivo: “The low abundance of TLR4 and the limitations of commercial anti-TLR4 antibodies precluded a direct immunohistochemical assessment of TLR4 loss in Tlr4VEKO and Tlr4MKO mice.”

      (2) The authors did not examine the consequences of Tlr4 cKO on the course of meningitis or bacterial burden. Knowing the impact of this would strengthen the paper and allow us to determine if the endothelial responses are helpful or harmful in meningitis progression.

      For the revised manuscript, we have conducted the following comparisons between infected and uninfected WT and infected and uninfected Cdh5-CreER; Tlr4floxed mice: (1) quantifying E. coli in the blood of infected mice by counting colonies on agar plates; (2) quantifying E. coli in the brain by measuring the red fluorescent protein (RFP) signal (the infecting E. coli carry an RFP-expression plasmid); (3) histologically surveying liver and lung for RFP+ E. coli; (4) monitoring the weights of infected and uninfected mice. These data are presented in Figure 2 – figure supplement 4 and in the Results section, and they can be summarized as follows. (1) E. coli is consistently detectable in the blood, brain, and peripheral organs in infected mice and is not detectable in control mice; (2) there are no statistically significant differences between infected WT and infected Cdh5-CreER; Tlr4floxed mice; (3) infected mice of both genotypes stop gaining weight between the time of infection (P5) and 24 hours later at the time of sacrifice (P6). Our conclusion is that loss of TLR4 in endothelial cells and in a subset of other non-myeloid cells in the leptomeninges does not alter the overall clinical course of the infection despite changes in leptomeningeal gene expression and vascular permeability.

      (3) TLR4 is a known receptor for LPS. It is unsurprising (especially in the in vitro experiments) that Tlr4 knockout reduces NF-kB signalling and other downstream changes to endothelial cells. Furthermore, it is uncertain if the infection was left to continue, similar changes to the endothelium would nonetheless occur through other mediators such as IL1B and TNFa.

      We agree that it makes logical sense that Tlr4 KO decreases NF-кB signaling. The interesting next question is: what are the mechanistic underpinnings of the responses that are downstream of TLR4 and NF-кB? The cell culture experiments with WT vs. Tlr4KO bEnd.3 cells identify one set of cell biological responses related to Cldn5 and junctional integrity, and the NF-кB inhibition experiment (Figure 6) implies that rapid internalization of Cldn5 occurs in the absence of NF-кB mediated transcriptional changes. Regarding the possibility that other mediators such as IL1B or TNFα might, at least partially, make up for the lack of TLR4 signaling later in the infection, that is an open question at present.

      (3) The arachnoid fibroblast 2 cluster should be renamed to dural border cells based on their high expression of Slc4a10, Adamtsl3, Tmeff2, etc which are all highly enriched in dural border cells. I suspect this cluster is also highly enriched for Slc47a1, probably the most specific marker for these cells (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002).

      Thank you for this comment. The reviewer is correct. These are dural border cells and they express Slc47a1, as seen in a new supplemental Figure 1 – figure supplement 4, which shows UMAP plots for many leptomeningeal cell type-specific genes. We have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (4) It would be helpful to provide higher resolution images of Cldn5 in the leptomeningeal mounts. At the current resolution, it is difficult to tell if there is a similar internalisation/disruption phenotype to what is observed in vitro. Notably, this finding is similar to another recent publication on Cldn5 recycling (in the context of stroke) (10.1186/s40478-025-02125-6).

      Higher resolution images of Cldn5 in leptomeningeal vessels without or with E. coli infection are now shown in Figure 3 - figure supplement 1C. There is a visual impression of greater area occupied by Cldn5, which is confirmed by quantification (Figure 3A and B). This effect appears to be due to both an average increase in vessel diameter and a redistribution of some of the Cldn5 away from plasma membrane junctions. Thank you for pointing out the interesting and relevant Cottarelli et al (2025) paper, which we had not read. This is now referenced.

      Minor points

      (1) Typo: prominant should be spelled prominent.

      Thank you for catching that one. It is now corrected.

      (2) Strictly speaking, the arachnoid layer is not epithelial (despite Cdh1 expression). They are fibroblasts that acquire barrier-forming properties.

      Thank you for that comment. That appears to be the consensus view, and we will go along with it.

      (3) Notably, LyzM Cre will also recombine in other myeloid populations, so I wouldn't describe it as a macrophage.

      Thank you for this comment. We agree, and we have therefore changed the text and figure labels from “macrophage” to “myeloid”.

      (4) It is interesting and notable that ICAM1 expression is observed in nonendothelial populations, in the IHC, too, perhaps.

      We agree. ICAM1 may be a broader marker/mediator of inflammation than is generally recognized.

      (5) In F1B, your labels on the right image to the arachnoid barrier and pial surface are presumably meant to refer to the image on the left with DPP4 and laminin labelling? The subarachnoid should be between the laminin and DPP4 layers (although it will be collapsed in your preparations).

      Thank you for catching this error. The vertical bars were sized erroneously, and the labels were also placed erroneously. These have now been corrected.

      (6) I would reference the papers that defined leptomeningeal cell type markers (10.1038/s41586-023-06993-7, 10.1016/j.neuron.2023.09.002) when you define your cell types.

      Thank you. We have done that, and we have updated our cell cluster assignment to align with the assignments in Pietilä et al (2023).

      (7) I would change references to the subarachnoid space in your figures to the leptomeninges (which include the SAS, but extend either side of it).

      Thank you. The labels have been changed to “leptomeninges”.

      (8) In Figure 2 - Supplement 1A, it looks like the populations are mislabelled.

      Thank you. This has been corrected to be consistent with the assignments in Figure 1B

    1. Author response:

      We are pleased that the reviewers found the repeated single-neuron sequencing and the finding of less than 1% transcriptomic variability to be original and striking, valued the single-neuron Dscam isoform repertoires and the scale of the functional screen, and judged the evidence solid to compelling. We provide below our provisional response and an outline of the revisions we plan.

      Overall plan: We intend to submit a revised version that addresses the public reviews and the recommendations to the authors. Because our conclusions rest on data already in the manuscript, the revisions are clarifications, added analysis of existing data, tempered language, and improved figures, rather than new experiments. Given the focused nature of these revisions, we would be happy for the editors to assess the revised version without re-involving the reviewers.

      One factual note for the Assessment and public reviews: The morphological RNAi screen comprised 213 cell-surface receptor genes; the figure of “140 genes” in one public review is the number that produced strong-to-severe phenotypes (Grade 3–5 at >40% penetrance), not the number screened. We will make this unambiguous in the revised text.

      Main changes in the revision:

      (1) We will explain that the RNAi screen was performed blind and independent of the RNA sequencing experiments. This was intentional, so that functional perturbation and transcriptomic identity would serve as independent lines of evidence, but could be compared with each other.

      (2) We will revise the Methods and Results to clarify how the morphological RNAi screen and behavioral subset should be interpreted, including conservative treatment of the negative-control distribution and mild-to-moderate phenotypes.

      (3) We will soften language that overstated certainty. Differentially expressed molecules are now described as prioritized candidates and convergent evidence, not definitive determinants.

      (4) We will reframe Gr59d/PXGS experiments as morphological rewiring and ectopic branching, and no longer imply a pSc-like conversion or functional rewiring.

      (5) We will add a limitations paragraph addressing RNAi off-target/background concerns, the absence of direct aPa functional testing, and the need for future mechanistic validation.

      (6) We will disclose or remove any figure panels that overlap with the companion PXGS manuscript and revise legends/labels to make it more clear.

      We hope these revisions make the logic of the study clearer and align the strength of the claims with the evidence.

      Below is our more detailed (provisional) response (not sure if this is required at this stage):

      Response to the eLife Assessment:

      Clearer articulation of the experimental logic. The Assessment’s central request (Reviewer 2) concerns the relationship between the transcriptomic experiments and the functional screen. The two were performed independently on purpose; the RNAi screen was assembled from a comprehensive survey of the literature rather than from the results of our differentially expressed genes from single-cell RNA sequencing. Thus, the RNAi screen was performed and graded blind in parallel with the single cell sequencing, with the gene identities unmasked only after both were complete. This was intentional, so that the sequencing (i.e., which molecules differ between neurons) and the screen (which molecules are functionally required) would provide mutually unbiased corroborating evidence rather than self-referential support/circular reasoning. We will state this more explicitly in the Introduction, in the Results where the screen is introduced, and in the Methods.

      Additional controls in the RNAi screen. We will treat the 39 genes that were identified in our single-cell RNA sequencing to be not expressed in the pSc neuron as a randomized negative control set in our RNAi experiments. We will state more explicitly the empirical RNAi false positive rate for a miswiring phenotype is 6/39 = 15%, likely due to RNAi off-target effects. We will also more clearly state that our claims about cell surface receptor functions are restricted to strong-to-severe phenotypes at high penetrance reproduced by at least two independent RNAi lines and corroborated independently (differential expression and/or single neuron qPCR).

      A more complete characterization of the re-wiring. We will state more clearly that mis-expressing the pSc-enriched cell surface receptors within Gr59d neurons partially shifts the arbour toward a pSc-like pattern (e.g., increased ectopic branching), and does not reproduce the full anatomical wiring, and that functional/behavioral re-wiring was not tested.

      Response to Reviewer 1:

      Reviewer 1 found the work valuable and data-rich, and the Dscam expression bias interesting. They noted over-confident language and asked how rigorously the differentially expressed genes were identified.

      Over-confident language. We will rewrite the two flagged sentences. The claim that the ~10 differentially expressed molecules are “likely the most important” will become a correlational statement, while also noting the lack of an aPa-specific Gal4 driver for direct testing. Our sentence that, “Our RNA sequencing data is biologically inadequate without a functional characterization of each molecule within the specific neuron” will be replaced with a clearer statement that gene expression data can nominate candidates, and functional perturbation of each gene is required to demonstrate necessity and sufficiency (i.e., biological function); which is exactly why we paired the RNA sequencing with an independent RNAi screen.

      Rigor of the differential-expression calls. We will more clearly state the statistical criteria in the Results (absolute log2 fold change ≥ 2 and Benjamini–Hochberg-adjusted p < 0.05). We will also note the small replicate numbers for the pooled pSc versus aPa comparisons, and emphasize that the central gene calls are independently supported by the blind RNAi screen and, for five genes, by single neuron qPCR. The full statistical workflow is in the Methods.

      Response to Reviewer 2:

      Reviewer 2 considered the findings potentially important but raised concerns about the experimental logic, the rigour of the screen, the completeness of the re-wiring, figure quality, and overlap with our PXGS companion paper. We will address each.

      Experimental logic. Beyond the design of our independent, blinded RNAi screen described above, we will add to the Introduction the rationale for sequencing single identified neurons (rather than subclasses) along with the two-pronged strategy, and add a summary paragraph at the start of the Discussion.

      Developmental stage choices. We will clarify our justification for the P14 pupal stage (the period when the mechanosensory neuron is actively elaborating its arbour while also enabling dissection). We will also clarify the rationale and caveats for comparing the pupal pSc neuron with the adult Gr59d neuron (i.e., the wiring occurs at the pupal stage, but the pupal Gr59d neurons could not be isolated at sufficient quality; the pSc pupal samples are less age-synchronized, so we simply used the comparison to identify the genes shared with the adult comparison).

      Transcriptome precision controls. We will state that the ten single pSc neurons passed the same quality controls for neuronal markers (elav, nSyb) and glial markers (Repo, moody < 20 reads) as all single-neuron libraries, which argues against any contamination by the attendant glial cell, and the aPa transcriptome is used as a different identity comparison.

      Off-target rate. As stated above, we will add the false positive rate for RNAi and restrict our confidence claims to those genes/cell surface receptors with multiple lines of evidence (e.g., strong phenotype, multiple RNAi lines, RNA sequencing, etc).

      Rewiring completeness and the behavioral prediction. As stated above, we will clarify that true re-wiring of the Gr59d neuron requires a future experiment, where a bitter tastant stimulus would elicit a grooming response.

      Response to Reviewer 3:

      We thank Reviewer 3 for judging our work to be fundamental in significance and the evidence compelling, with no major criticisms. Our clarifications above will further reinforce our hypothesis that the differential expression of specific cell surface receptors “do, in fact, control synaptic patterns,” which the reviewer highlighted.

      We are grateful for the reviewers’ time and for eLife’s model. We believe the planned revisions substantially clarify the experimental logic and tighten the claims, and we look forward to submitting the revised version.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      (1) Zonation definition under injury has been shown to be sustained broadly, but is not sufficiently validated and quantified, especially considering the resolution of the 10x Visium system and the potential variation of outcomes based on how to define zones.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g.,Alb, Mup20, Cyp2f2,Pck1,Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2,Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g.,Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) The model is built entirely in APAP injury, which specifically targets pericentral hepatocytes. It remains unclear whether the proposed mechanism applies to other liver injuries (e.g., partial hepatectomy, CCl4).

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Baseline proliferation appears higher than expected in homeostasis (Figure 1B), and fold change analysis (not absolute counts) may be needed to assess zonal proliferation suppression (Figure 1D).

      We thank the reviewer for this insightful comment. The baseline proliferation rates observed in our study are consistent with previously reported zonal distributions (PMID: 33632817; PMID: 33632818), with approximately 70% of proliferating hepatocytes located in zone 2, 20% in zone 3, and 10% in zone 1 under homeostatic conditions. To further address the reviewer’s concern, we performed a fold-change analysis of Ki-67<sup>+</sup>hepatocytes across different zones. This analysis revealed that only the mid (zone 2) and pericentral regions exhibited significant changes, whereas no statistically significant differences were observed in the other zones (as shown in Author response image 1). Importantly, when considered together with the absolute cell counts, these results indicate that the apparent suppression of proliferation is most pronounced in the mid zone, likely due to its relatively higher baseline proliferation under homeostatic conditions. In contrast, this effect is less evident in the fold-change analysis, as zones with low baseline proliferation show limited dynamic range for detecting relative changes.

      Author response image 1.

      Fold changes of Ki67-positive cells across liver zones (PC, Mid, PP) at 0, 3, 6, 12 and 24 h post-APAP. Fold change was the number of Ki67-positive cells in the three regions at each time point after APAP treatment divided by the number of positive cells in each region at 0 hour post-APAP. (a) denotes significance between PC and Mid regions, (b) denotes significance between PC and PP regions, and (c) denotes significance between Mid and PP regions.

      (4) AAV-based overexpression raises potential confounds (altered CYP activity before injury) and shows incomplete penetrance that is not quantified (Figure 5 - Figure 6).

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (5) The functional link between proliferation suppression and improved survival is inferred, but direct survival /injury readouts are limited.

      We thank the reviewer for this insightful comment. To more directly evaluate the functional link between proliferation control and liver injury, we manipulated Btg2, a downstream effector of the Atf4–Chop axis and a known inhibitor of cell proliferation. Knockdown of Btg2 using AAV8–CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial knockdown efficiency, we observed a clear exacerbation of liver injury, as evidenced by an approximately 2-fold increase in serum ALT levels and a ~1.5-fold expansion of necrotic areas. In parallel, hepatocyte proliferation was significantly increased (~1.8-fold increase in Ki67⁺ hepatocytes) compared to control mice (Revised Figure 6K–N). Conversely, Btg2 overexpression produced the opposite phenotype, markedly attenuating liver injury while suppressing hepatocyte proliferation (Revised Figure 6G–J). Together, these gain- and loss-of-function data provide direct evidence linking proliferation control to injury severity, thereby supporting a causal relationship between suppressed proliferation and improved liver outcomes (Revised manuscript, page 14, lines 399–406).

      Reviewer #2 (Public Review):

      (1) Starting with the basics, one wonders why midlobular hepatocytes manage to mount a defensive response to APAP but pericentral hepatocytes don't. Is this because midlobular hepatocytes express the relevant Cyps (2e1, but also 1a2 and 3a11) at lower levels, which mitigates toxicity and buys them time? This would be supported by F2A but not by F3B, at least not for the most important Cyp2e1. A moderate difference is shown for Cyp1a2 expression in F3D, but is that enough to explain the different fates? Or are additional post-transcriptional effects on these Cyps at work?

      We thank the reviewer for this important question. We fully agree that the differential susceptibility between mid‑zone and pericentral (PC) hepatocytes is likely rooted in the zonal gradient of cytochrome P450 expression. Our spatial transcriptomics data (Revised Figure 2A) show that mid‑zone hepatocytes express Cyp2e1, Cyp1a2, and Cyp3a11 at levels intermediate between PC and periportal (PP) zones. This intermediate expression may generate sufficient NAPQI to activate stress signaling but not so much as to cause immediate mitochondrial collapse, thus “buying time” for adaptive responses. We also appreciate the reviewer’s observation that Cyp2e1 mRNA levels remain highest in the PC zone even after APAP (Revised Figure 3B). However, mRNA abundance does not necessarily reflect functional protein level. In the PC zone, massive necrosis rapidly compromises cellular integrity; as shown in Revised Figure 3D, Cyp1a2 protein declines sharply around the central vein, and we observed similar degradation for Cyp2e1 (data not shown). Consequently, despite sustained Cyp2e1 transcripts, PC hepatocytes are unable to mount an effective stress response because they are already undergoing cell death. By contrast, mid‑zone hepatocytes retain sufficient metabolic capacity to activate the Atf4‑Chop axis while preserving cellular function.

      (2) The evidence presented in support of cell cycle arrest of midlobular hepatocytes is not fully convincing: there is no overt difference in S and G2/M gene scores in F2F; the marker genes used for S phase and G1 to S progression in F2G are unusual. Along these lines, one wonders if spatial transcriptomics confirmed the Ki67 immunostaining results in F1 also for specific zones, not only overall, as shown in F2E?

      We thank the reviewer for these important observations. We agree that the current spatial transcriptomics (ST) data alone do not provide sufficiently strong support for this conclusion. The limited sensitivity of ST for detecting rare proliferative events further constrains its utility in this context. At baseline, only ~1% of ST spots are Ki67-positive (Revised Figure S1I), and this fraction becomes even lower during the early phase following APAP injury. As a result, there are insufficient Ki67+ spots to robustly assess zonal distribution using ST, which precludes a reliable spatial validation of proliferation patterns at this resolution. For this reason, our primary evidence for zonal proliferation dynamics relies on Ki67 immunohistochemistry (Revised Figure 1), which provides single-cell resolution and higher sensitivity. These data show a marked reduction in Ki67+ hepatocytes specifically in the midlobular zone at 3-6 hours post-APAP, supporting a transient suppression of proliferation in this region. In addition, we agree that the transcriptional evidence for cell cycle arrest was not strong the S and G2/M scores showed no overt difference, and the gene sets used were suboptimal. We have therefore moved these analyses to the supplement and toned down the claims. We have also clarified this limitation in the manuscript (Revised manuscript, page 18, line 518-524)

      (3) The authors conclude in line 364 that halting of proliferation by Btg2 favors survival, which raises the question of whether Btg2 knockout causes death in midlobular hepatocytes in F6K. Data addressing this question, that is, the localization and extent of tissue necrosis and ALT levels after APAP, are missing. The efficiency of the knockout of Btg2 is also not given.

      We thank the reviewer for this insightful comment. We have included the missing data. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup>hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N, Revised manuscript, page 14, line 399-406).

      (4) Related to the previous question, the BTG2 immunostaining in F6F is not convincing when compared to F6D. One also wonders if it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3?

      We thank the reviewer for this insightful comment. We have included an inset of the original image to better show BTG2 staining in revised Figure 6F. During our study, we tested BTG2 expression in mice transduced with AAV‑TBG‑EGFP or AAV‑TBG‑BTG2 for three weeks without APAP challenge. We observed that BTG2 in these non‑injured livers was predominantly cytoplasmic (Author response image 2), contrasting with the nuclear localization seen after APAP treatment (Figure 6F). Regarding whether it is necessary to apply APAP to find induction of BTG2 by AAV-Ddit3, we think Ddit3 promotes BTG2 expression (as shown in revised Figure F6F), but APAP is necessary for its nuclear translocation.

      Author response image 2.

      Immunohistochemical detection of Btg2 in liver tissue from mice transduced with AAV-TBG-EGFP or AAV-TBG-Btg2 for 3 weeks without APAP treatment.

      (5) Related to the previous question, the proposed Atf4-Ddit3 axis is challenged by the lack of midlobular induction of Atf4 in the APAP scRNA-seq data published by another group, presented in S4F and G. Further analysis of AAV-Atf4 samples generated for F5 could address whether it is really Atf4 that acts on Ddit3 in APAP toxicity.

      We thank the reviewer for this insightful comment. We agree that Atf4 was not among the top 30 active transcription factors in our initial analysis; however, when we extended the list to the top 50, Atf4 was included. We have therefore updated Revised Figures S4F and G to show the top 50 transcription factors. We also appreciate the reviewer’s suggestion to further investigate whether Atf4 directly acts on Ddit3 in the context of APAP toxicity. While this still shows a less pronounced midlobular enrichment for Atf4 compared with Ddit3, we sought additional evidence for a functional Atf4-Ddit3 link. In primary hepatocytes treated with APAP, we observed nuclear co‑localization of Atf4 and Ddit3 (Author response image 3A) and increased nuclear protein levels of both factors (Author response image 3B), supporting their potential cooperative role. We agree that direct analysis of AAV‑Atf4 samples generated for Figure 5 would provide more definitive evidence; unfortunately, co‑staining for Atf4 and Ddit3 on those tissue sections didn’t work well.

      Author response image 3.

      Subcellular localization of Atf4 and Chop in primary hepatocytes following APAP treatment. (A) Immunofluorescence staining of Atf4 and Chop in primary hepatocytes treated with 10 mM APAP for 6 hours or left untreated (UT). Nuclei were counterstained with DAPI. Scale bar as indicated. (B) Primary hepatocytes were treated with 0, 5, or 10 mM APAP for 6 hours. Cytoplasmic and nuclear fractions were isolated and analyzed by western blot. Lamin B1 and α-Tubulin were used as markers for the nucleus and cytoplasm, respectively

      (6) Related to the previous question, the ATF4 immunostaining in F5A doesn't look convincing, with many brown pigments appearing to be outside of the nucleus.

      We thank the reviewer for this helpful comment. To better demonstrate ATF4 nuclear localization, we have added enlarged insets of the original representative images in revised Figure 5A. These magnified views more clearly show nuclear ATF4 staining after APAP treatment, addressing the concern about extranuclear signal.

      (7) It is not ruled out that AAV expression of Atf4 or Btg2 reduces hepatocyte sensitivity to APAP by affecting the expression of the Cyps needed for activation. In other words, does AAV-Atf4 or AAV-Btg2 change the expression of any of the Cyps relevant to APAP in the 3 weeks before APAP application (F5B)?

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      (8) It is laudable that the authors tried to extend their findings to humans by using snRNA-seq data from a published study (line 391), but it is unclear why they didn't analyze all 10 patients in that study but instead focused on 2 and stated that this small sample number prevented drawing definitive conclusions and could therefore only be mentioned in the discussion.

      We thank the reviewer for this clarification. The analysis mentioned in line 391 originally referred to spatial transcriptomics (ST) data from two ALF patients, not snRNA-seq. For the snRNA-seq dataset, we analyzed all 10 patients, but snRNA-seq lacks spatial resolution and cannot reliably assign zonal identity. We stipulate that snRNA-seq requires viable cells and thus likely excludes necrotic/peri-necrotic areas. Therefore, direct zonal comparison with our ST data was not possible. We have now clarified this in the revised manuscript (Revised manuscript, page 18, line 510-519).

      Reviewer #3 (Public Review):

      The main concern is that the overexpression of ATF4 and DDIT3 is causing reduced cell death and damage by APAP. This makes it harder to understand if these genes are truly increasing survival or if they are just reducing the injury caused by APAP. It may be better to perform overexpression immediately after, or at the same time as APAP delivery. Alternatively, loss-of-function experiments using AAV-shRNAs against these targets could be useful.

      We thank the reviewer for raising this important point. We agree that overexpression prior to APAP administration leaves open the question of whether the observed protection reflects true cytoprotection or simply reduced initiation of injury. To address this, we pursued loss‑of‑function approaches. Due to their very low basal expression, AAV‑shRNA‑mediated knockdown of endogenous Atf4 and Ddit3 proved inefficient. We therefore targeted Btg2, a downstream mediator of Ddit3 that inhibits proliferation. Knockdown of Btg2 resulted in a significant increase in APAP‑induced liver injury, as evidenced by elevated ALT levels and expanded necrotic areas (Revised Figure 6K-N). These results indicate that the ATF4‑DDIT3‑BTG2 axis limits hepatocellular damage, consistent with a protective role. We have clarified this point in the revised manuscript (page 15, line 407-414)

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Clarify how zones were defined when necrosis disrupted pericentral areas. Provide marker validation across time and whether necrotic spots are excluded or not from zonal analysis.

      We thank the reviewer for these insightful suggestions. In this study, under normal conditions (APAP 0h), each liver lobule was divided into three zones based on unbiased gene expression profiles. The PP zone was defined by enrichment of PP signature genes (e.g., Alb, Mup20, Cyp2f2, Pck1, Apoa4). The PC zone was defined by high expression of PC markers (e.g., Gs, Cyp2e1, Oat, Cyp1a2, Apoe). The Mid zone comprised regions with intermediate expression of PC and PP markers and elevated levels of Igfbp2 and Hamp (Revised Figure 2A and S1C). Following APAP-induced injury (3 h, 6 h), the PC zone remained identifiable based on residual enrichment of PC signature genes (e.g., Cyp1a2, Glul) despite necrosis and reduced overall transcription, while PP gene expression remained largely unchanged. The Mid zone was defined as the transcriptional cluster between PC and PP regions exhibiting marked reprogramming, (e.g., Sqstm1, Igfbp1) (Revised Figure 2A and S1C). To validate and quantify our zonation approach, we compared it with classical nine even layers from central vein (CV) to portal vein (PV). Immunostaining and quantification for Cyp2f2 (a PP marker), p62 (the protein product of Sqstm1, a Mid marker during early liver injury), Glutamine Synthetase (GS, the protein product of Glul, a PC marker) further corroborated zone definitions at each time point, showing correspondence of our PC (layers 1–2), Mid (layers 3–6), and PP (layers 7–9) (Revised Figure 2 B-D) (Revised manuscript, page 5, lines 119–131, page 6-7, line 174-182).

      (2) Test whether the ISR-Btg2 program applies in other models; even targeted validation via qPCR and IF would be valuable.

      We thank the reviewer for this insightful comment. To test whether the proposed mechanism applies to other liver injuries, we employed mouse models of partial hepatectomy (PHx) and carbon tetrachloride (CCl4)-induced acute liver injury. In our CCl4 model (administered intraperitoneally in corn oil, with samples collected 18 h post‑injection), the ISR was activated around injury sites, accompanied by decreased proliferation, as evidenced by increased expression of p‑eIF2α, Atf4, Chop, and Btg2, along with reduced Ki67 expression (Revised Figure S6A–G). In PHx model (examined 24 h after surgery), ISR activation was similarly observed around ischemic injury sites, with increased p‑eIF2α, Atf4, Chop, and Btg2 expression and undetectable Ki67 expression (Revised Figure S7A–G). Together, these additional models suggest that the proposed mechanism may be applicable to other types of liver injury (Revised manuscript, page 15, Line 410-426).

      (3) Proliferation quantification in liver sections in Figure 1: how to define the zones and why, at the basal level, there is a high proliferation rate in the mid zone? From Figure 1B-C, all three zones showed decreased hepatocyte proliferation, although the mid zone had a higher baseline. Will the mid-zone stand out by converting to the fold change of Ki-67+ hepatocytes decrease?

      We thank the reviewer for these insightful comments. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x<sup>2</sup> + z<sup>2</sup> – y<sup>2</sup>) / (2z<sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885). Consistent with previous reports (PMID: 33632817; PMID: 33632818), we observed a higher baseline proliferation rate in the mid zone, where approximately 70% of proliferating hepatocytes reside under basal conditions, compared to 10% in zone 1 and 20% in zone 3. However, when analyzing the fold change in Ki-67+ hepatocytes, only Mid and PC region showed significant difference in Ki-67+ hepatocytes, other zones showed no significant differences (as shown in the fold-change results in Author response image 1), indicating that the mid zone does not stand out in the fold change analysis. See Author response image 1.

      (4) The authors need to strengthen the causal chain with rescue experiments, e.g., Atf4/Chop overexpression and Btg2 knockdown. Link proliferation suppression to survival/ALT directly.

      We thank the reviewer for these constructive comments. Besides existing data from Figure 5 (Atf4 overexpression), we included Btg2 knockdown data in the revised Figure. Knockdown of Btg2 using AAV8‑CasRx achieved a moderate (~30%) reduction in Btg2 expression (Revised Figure S5D). Despite this partial efficiency, we observed a significant increase in serum ALT levels (~2‑fold), expansion of necrotic areas (~1.5‑fold), and a marked increase in Ki67<sup>+</sup> hepatocytes (~1.8‑fold) compared to control mice (Revised Figure 6K–N) (Revised manuscript, page 14, lines 399–406).

      (5) Transduction efficiency, distribution, and expression levels via the AAV overexpression need to be quantified. Key CYP genes in the APAP metabolic pathway need to be assessed to exclude confounds.

      We thank the reviewer for raising these important points. We measured basal Cyp2e1 protein levels by western blot in AAV‑EGFP, AAV‑Atf4, and AAV‑Btg2 mice without APAP treatment. Compared to AAV‑EGFP controls, Cyp2e1 expression was modestly reduced in the Atf4 and Btg2 groups, respectively (Revised Figure S5A). Although we assessed protein abundance rather than enzymatic activity directly, Cyp2e1 protein levels under basal conditions generally correlate well with activity. Published studies demonstrate that robust protection against APAP hepatotoxicity typically requires >50% suppression of CYP2E1 activity (PMID: 35145060; PMID: 30151903). The minor reductions we observed are therefore far below the threshold needed to explain the 70–90% decreases in serum ALT conferred by Atf4 or Btg2 overexpression (Revised Figures 5D and 6I). Accordingly, altered CYP2E1 activity is unlikely to represent a significant confound in our model.

      We quantified transduction efficiency by immunohistochemical detection of the respective transgene proteins and determined the percentage of positive hepatocytes. At a dose of 1.2 × 10<sup>11</sup> viral genomes per animal, average transduction rates were 32% (EGFP), 18% (Atf4), and 23% (Btg2) (Revised Figure S5B). Individual animal transduction efficiency showed a negative correlation with serum ALT levels (e.g. Pearson r = –0.7681, p = 0.0260 for Atf4; Revised Figure S5C), demonstrating that greater transgene expression associates with stronger protection. Although these average transduction rates appear modest relative to the 70–90% reduction in ALT, this apparent disproportion is consistent with the known tendency of AAV‑TBG vectors to transduce hepatocytes preferentially in the pericentral region—the same zone where APAP‑induced necrosis initiates. Pericentral enrichment of transgene expression could thus provide disproportionate protection by targeting the most vulnerable cells. These data are now included in Revised Figure S5A–C and detailed in the Results (page 14, lines 383–399).

      (6) The authors claim that the requirement of the Atf4/Chop at the early stage of APAP injury protects hepatocytes from proliferation for survival. What is the consequence if we remove the protective mechanism?

      We thank the reviewer for this insightful question. In our model, early induction of Atf4 and Chop functions as a cell survival checkpoint. Removal of this protective mechanism is predicted to result in two deleterious outcomes: (1) Acute exacerbation of necrosis due to the inability of hepatocytes to manage stress-induced bioenergetic demands, and (2) Impaired long-term regeneration due to depletion of the surviving cell pool. We directly tested the acute prediction (< 24 h) in Author response image 4. We deleted Ddit3 specifically in hepatocytes. Initial attempts using AAV-CasRx failed due to negligible baseline Atf4/Chop expression in healthy liver, preventing effective knockdown. We therefore generated hepatocyte-specific Ddit3 knockout mice (Alb<sup>∆Ddit3</sup>; Author response image 4B). Immunohistochemistry confirmed APAP-induced Chop induction occurs primarily in the centrilobular zone by 6 h (Author response image 4A). Following a two-dose APAP regimen (Author response image 4C), Alb<sup>∆Ddit3</sup> mice displayed significantly larger areas of centrilobular necrosis compared to Ddit3<sup>fl/fl</sup> controls (Author response image 4D; **p < 0.01). Thus, hepatocyte-intrinsic Chop limits acute APAP injury, consistent with its proposed early protective role.

      Author response image 4.

      Hepatocyte-specific deletion of Ddit3 exacerbates APAP-induced liver injury. (A) Immunohistochemical staining of Chop in liver sections at 0,3 and 6 h post-APAP. Red arrows indicate Chop-positive hepatocytes. Scale bar = 50μm. Quantification of zonal distribution of Chop-positive cells in liver sections at 6 h post-APAP is conducted . The statistic is the percentage of Chop-positive hepatocytes in each layer over the total number of Chop-positive hepatocytes. n=3 mice. (B)The construction, genotyping strategy and genotyping results of Alb<sup>∆Ddit3</sup> mice. P: positive control; WT: Wild-type; Neg: Blank control(ddH<sub>2</sub>O). (C) Schematic figure illustrating the experimental strategy for the administration of two doses of APAP to Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice. (D) H&E staining showing liver morphology from Ddit3<sup>fl/fl</sup> and Alb<sup>∆Ddit3</sup> mice at 6 h post-second dose of APAP. Injured area is outlined by black dashed lines. Scale bars = 200 μm. The percentage of injury area is quantified. n = 3- 4 mice/group. Data are represented as means ± SD; *p < 0.05; **p < 0.01; ***p < 0.001; ****p < 0.0001; ns, not significant.

      (7) Is there any human relevance to the sensitivity of APAP injury regarding the Atf4/Chop axis?

      We thank the reviewer for this insightful comment. During our study, we analyzed a spatial transcriptomics dataset from APAP patients. In one of two analyzed patients, mid-zone hepatocytes exhibited transcriptional signatures remarkably consistent with our murine findings, including: (1) upregulation of Atf4-Chop pathways, and (2) downregulation of cell proliferation genes (Author response image 5). This suggests that this axis may also be involved in the response to APAP injury in humans. However, given the limited sample size, definitive conclusions cannot be drawn at this stage. We have now included this point in the Discussion section (Revised manuscript, page 18, line 510-519).

      Author response image 5.

      Spatial transcriptomics (GSE223561) reveals zonal gene expression changes in APAP patients. Heatmap of ISR, cell death, and cell cycle gene expression across zonal regions in healthy versus APAP‑treated human livers. 

      (8) Several IHC stainings have a weak signal and need inserts to zoom in for a clear view of the positive signals. Figure 5A, E, G, and Figure 6D, F.

      We thank the reviewer for this observation. We agree that the immunostaining signals for several target genes are relatively weak, which reflects their low endogenous expression levels. To address this, we have included higher-magnification insets in the indicated panels (Revised Figure 5A, E, G and Figure 6D, F) to show the positive signals.

      Reviewer #2 (Recommendations for the authors):

      (1) What is the functional classification of DEG in F2A based on? GO terms?

      We thank the reviewer for this constructive question. The functional classification of differentially expressed genes (DEGs) in F2A is based on Gene Ontology (GO) terms. For each DEG, we retrieved its associated GO annotations across the three main categories (biological process, cellular component, molecular function). In cases where a gene was assigned multiple GO terms, we prioritized the most representative or significantly enriched term for functional interpretation. This clarification has been incorporated into the revised figure legend and the according GO number has been included in the figure.

      (3) The rationale for focusing on CHOP is not clear because Ddit3 is not shown in the spatial transcriptomics in F2A and is not significant in F2B, contradicting what is stated in line 206.

      We thank the reviewer for raising this important point. We apologize that Ddit3 was missing from the original figure. In the revised manuscript, we have included an updated version of Figure 2A, which now shows that Ddit3 is indeed one of the differentially expressed genes (DEGs) in the Mid zone at both 3 and 6 hours post-APAP. We agree with the reviewer that, as shown in Figure S1G (previous Figure 2B), Ddit3 did not reach statistical significance, due to its relatively low expression level in that analysis. Nevertheless, when we examined transcription factor (TF) activity in the Mid zone during early AILI, Ddit3 and Atf3 ranked as the top two most highly expressed TFs among the top ten with the highest activity, whereas Atf4 ranked seventh (Revised Figure 4B and Figure S3B). Given that Ddit3 frequently co-worked with Atf4 and that the Atf4–Ddit3 axis plays a well-established role in cellular stress adaptation, we considered this pathway to be biologically relevant and worthy of further investigation.

      (3) The term "redistribution" used in line 197 to describe the expression of Cyp2e1 and other Cyps in the midlobular zone seems inappropriate, considering that they just continue to be expressed there, whereas pericentral hepatocytes are dying in F3B; the same applies to "Gene Expression Shift" in F3H.

      We thank the reviewer for this important clarification. We have revised the text (Revised manuscript, page 9, line 234-236) to state that selective loss of Cyp‑expressing pericentral hepatocytes leads to the mid‑zone becoming the primary site of residual Cyp activity. The figure label has been changed from “Gene Expression Shift” to “Peri‑necrotic Cyp retention” and the legend now explicitly notes that this is an apparent zonal shift due to necrosis, not active redistribution.

      Reviewer #3 (Recommendations for the authors):

      (1) Please do not use abbreviations like AILI. This makes the paper more difficult to read.

      We thank the reviewer for pointing this out. We have replaced AILI with the full term “APAP-induced liver injury” to ensure easiness for readers.

      (2) It will be important to clarify how pericentral, mid, and periportal were defined. In Figure 1, it appears that some of the pericentral hepatocytes that are Ki67 positive are quite mid-zonal. It would be important to have rigorous definitions for the location determination.

      We thank the reviewer for this constructive comment. To define the pericentral (PC), mid, and periportal (PP) zones, we adopted the classical nine‑layer model of the hepatic lobule described by Lin et al. (PMID: 29618815). Layers 1–2 were designated as the PC zone, layers 3–6 as the mid zone, and layers 7–9 as the PP zone. For quantitative zonal distribution of protein‑positive nuclei (e.g., Ki67, CHOP, ATF4), we calculated a position index (P.I.) based on distances to the nearest central vein (CV) and portal vein (PV), using the law of cosines: P.I. = (x <sup>2</sup> + z <sup>2</sup> – y <sup>2</sup>) / (2z <sup>2</sup>), where x = distance to CV, y = distance to PV, and z = distance between CV and PV. This quantification method has now been included in the Methods section (Revised manuscript, page 33, line 880-885).

      We thank the reviewers for their rigorous critique again. We thank eLife for fostering an environment of fairness and transparency that enables authors to communicate openly and present their data honestly.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1:

      I recommend that the title be changed to not focus on sex differences to avoid misunderstanding.

      We thank Reviewer #1 for this suggestion and agree that the original title could create a misleading impression. We have updated the title from "Sex-specific behavioral and thalamo-accumbal circuit adaptations after oxycodone abstinence" to "Thalamo-accumbal circuit adaptations following extended oxycodone abstinence" to more accurately reflect the scope of the findings.

      The authors should also address the lack of difference physiologically compared to the behavior as a caveat more clearly in the discussion.

      We thank the reviewer for this important suggestion. We have revised the Discussion to explicitly address this dissociation. Specifically, we added the following to the PVT-NAcSh synaptic strength section: " The absence of sex differences in PVT-NAcSh synaptic measures suggests that this pathway, as characterized here, represents a shared neuro-adaptation to prolonged oxycodone abstinence rather than a substrate for the heightened relapse vulnerability observed in females. The mechanisms driving sex-specific relapse likely involve additional circuit elements, such as sex hormone-dependent modulation, upstream inputs, or cell-type specific plasticity." This point is also summarized in the abstract.

      Reviewer #2:

      A major weakness of this study is the disconnect between the behavioral and neurophysiological data reported. While a striking sex difference in relapse-like behavior is observed, there are no statistically significant sex differences in any of the neurophysiological data reported. Moreover, without an experiment to functionally test the role of the PVT-NAc projection in relapse-like behavior following prolonged oxycodone, these two arms of the study seem divorced.

      We respectfully disagree with the characterization that the behavioral and neurophysiological data are "divorced." The two arms of the study converge on a consistent and meaningful finding: PVT-NAcSh synaptic strength increases specifically after prolonged abstinence, this is the same time point at which enhanced cue-induced relapse is observed in both sexes. The absence of sex differences in synaptic measures does not weaken this convergence; it refines it by suggesting that circuit-level potentiation is a shared neuro-adaptation, while the sex-specific behavioral phenotype likely reflects additional modulatory mechanisms acting on this shared substrate. We have revised the Discussion to explicitly address this dissociation, as noted in our response to Reviewer #1 above. We acknowledge that functional manipulation of the PVT-NAcSh circuit would be required to establish causality, and we state this clearly in the manuscript.

      In the introduction the authors state they aim to test the hypothesis that increased synaptic strength in PVTNAcSh projections are necessary for drug-seeking. This study does not include the required experiments to test this hypothesis.

      We have revised the relevant section in the Introduction to accurately reflect the scope of our study: " We aimed to determine whether synaptic strength in PVT-NAcSh projections is affected following oxycodone abstinence and whether such changes are associated with cue-induced relapse and drug-seeking. Additionally, we examined whether there are sex-specific differences in either cue-induced relapse or PVT-NAcSh synaptic transmission after either 1 (acute) or 14 (prolonged) days of forced abstinence. Our results demonstrate that sex-specific enhancement in cue-induced relapse emerges after prolonged abstinence but not during acute abstinence from oxycodone self-administration. Although both males and females show increased cue-induced relapse after prolonged abstinence, females exhibited a greater relapse rate compared to males. Both sexes showed similar increases in PVT-NAcSh synaptic strength after prolonged abstinence, while synaptic strength was not altered after acute abstinence compared to saline controls. Together, these findings reveal a time-dependent increase in PVT-NAcSh synaptic strength and a sex-specific effect of prolonged abstinence on cue-induced relapse, while synaptic enhancements after prolonged abstinence were not sex-specific." This revision avoids implying a necessary or causal role for the circuit, which we did not test.

      Reviewer #3:

      The PVT-NAcSh synaptic strengthening after prolonged abstinence is statistically indistinguishable between sexes, while females but not males show a time-dependent escalation in oxycodone seeking from 1 to 14 days of abstinence. The Discussion proposes hormonal modulation or differences in upstream inputs as possible explanations, but none of these are tested and the gap is left unresolved.

      We agree that the mechanistic basis of the behavioral sex difference remains an open question that the current study does not resolve. As noted in our response to Reviewer #1, we have revised the Discussion to explicitly acknowledge this dissociation and to clarify that PVT-NAcSh synaptic strengthening represents a shared neuro-adaptation rather than a mechanism specific to the female behavioral phenotype. We maintain that identifying this dissociation is itself a scientifically meaningful finding.

      The intrinsic excitability recordings come from NAcSh MSNs with no confirmation that those neurons receive direct PVT input, which was raised in the original review, acknowledged in the revision, and not experimentally addressed.

      We have added the following clarification to the excitability section of the Discussion: “It should be noted that the intrinsic excitability recordings were designed to characterize general properties of NAcSh MSNs following oxycodone abstinence, independent of their synaptic inputs. As such, the excitability data should be interpreted as reflecting changes in the NAcSh MSN population broadly rather than in PVT-connected neurons specifically. The standing theory suggests that MSN excitability decreases as a homeostatic response to increased glutamatergic input [23,43,58]. Our data do not support a compensatory decrease in excitability in either sex at either abstinence time point”. These recordings were never intended to be circuit-specific; the experiment was designed to characterize NAcSh MSN excitability at the population level, which is a valid and informative question in its own right.

      The male prolonged-abstinence excitability trend has approximately 20% statistical power and is non-significant, yet the Discussion interprets it as a potential neuro-adaptation that could facilitate signal flow through the PVT-NAcSh circuit and contribute to relapse, which goes well beyond what the data support.

      We have revised the relevant Discussion text to ensure the male excitability trend is interpreted appropriately. The revised text now reads: "In males, a non-significant trend toward increased excitability was observed after prolonged abstinence, with a large effect size (Cohen's d = 1.18); however, given that this group was substantially underpowered (approximately 20% power), this finding should be interpreted with caution and cannot be taken as evidence of a neuro-adaptation. Whether this trend, if confirmed in future studies with larger cohorts, reflects a broader MSN population response or is specific to PVT-connected neurons remains an open and interesting question." The speculative mechanistic interpretation previously present in this section has been removed.

      The failure to distinguish between D1 and D2 MSNs remains a significant limitation given that cell-type specific plasticity at PVT-NAc synapses has been shown to be directly relevant to opioid seeking in prior work.

      We agree that distinguishing between D1 and D2 MSNs would provide important mechanistic insight, and we acknowledge this explicitly as a limitation and a future direction in the Discussion. The use of transgenic Cre rat lines for cell-type-specific recording in a self-administration model requires significant additional infrastructure and was beyond the scope of the present study. This is precisely the direction our laboratory is currently pursuing, and the present findings provide empirical motivation for those experiments.

      The Conclusion builds a mechanistic framework around D2 MSNs, PV interneurons, and D1 MSNs that is drawn from studies using different drugs or experimental designs, and none of these cell-type-specific mechanisms are tested in the present experiments.

      We thank the reviewer for this important critique. We have revised the opening of the Conclusion to clarify that the cell-type-specific framework is grounded in prior literature and represents a hypothesis for future investigation rather than a conclusion drawn from the present data. The revised text now reads: " When considered alongside prior work, our findings highlight the need to examine the anatomical and cell-type specific organization of PVT inputs to the NAcSh in the context of opioid relapse. Based on existing literature, PVT projections onto D2 MSNs and PV interneurons may contribute to relapse vulnerability, while adaptations involving D1 MSNs may underlie incubation of craving, though these mechanisms remain to be directly tested in the oxycodone self-administration model used here". We believe this framing accurately represents the relationship between our findings and the broader literature without overstating what the present data demonstrates.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Tittelmeier et al. explored the role of sphingolipid metabolism in maintaining endolysosomal membrane integrity and its downstream effects on tau aggregation and toxicity, using both worms and human cell models. The authors showed that knockdown of sphingolipid metabolism genes reduced endolysosomal membrane fluidity, as revealed by FRAP and C-Laurdan imaging, leading to increased vesicle rupture. Furthermore, tau aggregates accumulated in endolysosomes and exacerbated membrane rigidity and damage, promoting seeded tau aggregation, likely by enabling tau seed escape into the cytosol. Importantly, unsaturated fatty acid supplementation restored membrane fluidity, suppressed tau propagation, and alleviated neurotoxicity in C. elegans. These findings provide insight into how lipid dysregulation contributes to tau pathology and highlight membrane fluidity restoration as a potential therapeutic avenue for Alzheimer's disease.

      Strengths:

      The study addresses the connection between sphingolipid metabolism, endolysosomal membrane integrity, and tau pathology, which is a relevant topic in the context of Alzheimer's disease and related tauopathies.

      The use of both C. elegans and human cell models provides cross-species perspectives that help frame the findings in a broader biological context.

      The combination of FRAP and C-Laurdan dye imaging offers a biophysical approach to investigate changes in membrane properties, which is a technically interesting aspect of the study.

      The observation that unsaturated fatty acid supplementation can modulate membrane fluidity and influence tau-related phenotypes adds an element of potential therapeutic interest.

      The study presents multiple experimental approaches to address the proposed mechanism, and efforts were made to examine both membrane behavior and tau aggregation dynamics.

      We thank the reviewer for this positive assessment of the study.

      Weaknesses:

      In Figure 3, the authors used C-Laurdan imaging to assess membrane fluidity and showed that knockdown of SPHK2, the human ortholog of sphk-1, led to increased membrane rigidity. However, the authors did not co-stain with a lysosomal marker, making it unclear whether the observed effect is specific to lysosomal membranes or reflects general membrane changes. Co-staining with LysoTracker or applying segmentation masks to isolate lysosomal signals would significantly improve interpretation.

      We agree with the reviewer that it is important to isolate lysosomal signals for interpreting the C-Laurdan data. We therefore repeated and extended the C-Laurdan experiments in combination with LysoTracker staining and selectively analyzed LysoTracker-positive regions. These analyses showed pronounced increases in GP values in LysoTracker-positive vesicles after SPHK2 knockdown, supporting the conclusion that SPHK2 depletion increases endolysosomal membrane rigidity. We also performed LysoTracker-based analysis in the tau-fibril and fatty-acid experiments to better assess lysosome-associated membrane properties (see new and updated Figures 3B-E, Figures 5A and B, Figures S3A, B, F, G, Figures S4A-F, and Figures S5A-D for details). The respective Results sections have been revised accordingly.

      Line 173 states that Lipofectamine 2000 increases membrane fluidity based on GP index changes, but this is incorrect. A higher GP index indicates increased membrane order (i.e., reduced fluidity), so the statement should be revised. Additionally, Lipofectamine 2000 can itself alter membrane rigidity, posing a risk of false-positive interpretations. To confirm the role of SPHK2 in this phenotype, the authors should use a CRISPR/Cas9 knockout model instead of relying solely on siRNA transfection, which may be confounded by the delivery reagent. Without lysosomal co-staining and SPHK2 KO validation, the authors cannot conclusively claim that SPHK2 loss affects endolysosomal membrane integrity.

      We thank the reviewer for pointing out the incorrect wording regarding Lipofectamine. A higher GP index indicates increased membrane order/rigidity, not increased fluidity. Since the revised main figure now includes the SH-SY5Y data (Figure 3A-D), in which Lipofectamine alone did not significantly alter GP values (see Figure S3A, B), we removed the misleading statement from the Results.

      We also agree that Lipofectamine can affect membrane properties and therefore needs to be carefully controlled. In all siRNA-mediated experiments, SPHK2 siRNA was compared to a matched control siRNA condition exposed to the same transfection reagent. We state in the Methods that cells were transfected with either SPHK2 or scrambled control siRNA and that the medium was exchanged after 6 h “to minimize lipofectamine impact on the endolysosomal system”. Thus, the effect attributed to SPHK2 KD is assessed relative to the appropriate Lipofectamine-containing control condition.

      Importantly, we have now added lysosomal co-staining to address the reviewer’s concern about compartment specificity. This analysis showed that SPHK2 KD resulted in a pronounced increase in GP values within LysoTracker-positive compartments, demonstrating increased membrane rigidity at lysosomes. Thus, the revised data support the conclusion that SPHK2 KD increases lysosome-associated membrane rigidity, rather than only causing nonspecific effects on other cellular membranes.

      We also clarified the relationship between the current siRNA-based assay and our previous CRISPR inhibition-based analysis. In the revised Results, we now write: “While SPHK2 KD alone significantly increased galectin puncta above the matched control, its effect was more modest than in our previous CRISPR inhibition-based analysis [20]. This difference likely stems from the earlier readout required for the combined siRNA/tau fibril assay, when transient Lipofectamine-associated effects still increased the control background.”

      This addresses why the SPHK2 KD effect appears smaller in the current siRNA/tau-fibril assay than in our previous CRISPR inhibition-based analysis. The previous study, which is now peer-reviewed and published in the journal Autophagy, used a CRISPR inhibition-based strategy to reduce SPHK2 levels, which resulted in a highly significant increase in sfGFP-LGALS3 foci formation compared to the control [1]. Thus, the SPHK2 phenotype is not supported solely by the current siRNA experiment.

      In addition, we sought to genetically validate the RNAi phenotypes using mutant strains. However, mutant strains were not available for all sphingolipid metabolism hits analyzed in this study. We therefore used the sphk-1 mutant strain available at CGC (CZ24969; sphk-1(ju831)) to validate one of the key SL metabolism hits independently of RNAi. The revised manuscript states: “As genetic validation independent of RNAi, we tested an available sphk-1 mutant strain, which also showed a robust increase in hypodermal sfGFP::LGALS3 foci (Figure S1A).” This result supports the conclusion that genetic perturbation of sphingosine kinase activity compromises endolysosomal integrity in vivo.

      Together, the revised manuscript addresses the reviewer’s concerns by correcting the GP interpretation, controlling the siRNA experiments against matched Lipofectamine-treated controls, adding LysoTracker-based lysosome-associated C-Laurdan analysis, relating the current siRNA data to our previous CRISPR inhibition-based analysis, and providing genetic validation for the available sphk-1 mutant.

      The section titled "Fibrillar tau increases membrane rigidity and exacerbates endolysosomal damage" (lines 177-215) requires substantial revision. The narrative jumps abruptly between worms and cell models, making it hard to follow the logic. The use of the F3ΔK281::mCherry strain is introduced without explanation or context. It is unclear whether this strain is relevant to lysosomal membrane rupture, as no reference or justification is provided. The authors should clarify whether this reporter is intended to detect lysosomal membrane permeabilization (LMP). If so, it would be more appropriate to use established LMP reporters, such as lysosome-targeted fluorescent sensors, galectin-based reporters, or dextran leakage assays. Based on the current data in Figure 3G, it is difficult to draw firm conclusions regarding membrane rupture levels.

      We agree that this section required clarification, and we have substantially revised the Results to improve the logic and separation between model systems.

      First, we now introduce the C. elegans reporter strain earlier in the manuscript, in the first Results section. In the revised text, we explain both the tau construct and the actual lysosomal damage reporter: “In this strain, endolysosomal membrane damage is monitored in the hypodermis by expression of human galectin-3 fused to superfolder-GFP (sfGFP::LGALS3). The animals also express an aggregation-prone tau fragment fused to mCherry (F3ΔK281::mCherry) in touch receptor neurons, which is transmitted to the hypodermis, as described previously [20].” We also clarify the principle of the Galectin reporter: “Under steady-state conditions, sfGFP::LGALS3 remains diffusely distributed throughout the cytosol. Upon endolysosomal damage, luminal β-galactosides become exposed and recruit sfGFP::LGALS3 into visible puncta, providing a sensitive readout of vesicle rupture.” Thus, F3ΔK281::mCherry is not the reporter for lysosomal membrane permeabilization; the membrane-damage readout is sfGFP::LGALS3 puncta formation.

      Second, we reorganized the manuscript to separate the human cell experiments from the C. elegans experiments more clearly. The revised section “Fibrillar tau and SPHK2 KD act in concert to exacerbate endolysosomal damage and seeded tau aggregation” now focuses on human cell data. The C. elegans experiments are now presented in a separate section, “Tau transmission sensitizes endolysosomal membranes to sphingolipid perturbations in vivo.” We believe that this revised structure now clearly distinguishes the role of the Galectin reporter from the tau transmission model, separates the human cell and C. elegans data, and avoids the abrupt transitions between model systems noted by the reviewer.

      To support the conclusion that sphingolipid metabolism gene knockdown alters membrane properties, the study would benefit from direct lipidomic analysis. Measuring changes in sphingolipid profiles in both C. elegans and cell models would provide biochemical evidence for the proposed disruption of lipid homeostasis. Given the availability of lipidomics platforms, this type of analysis should be feasible in both worms and human cells and would significantly strengthen the mechanistic claims regarding membrane fluidity and integrity.

      Because we did not perform lipidomics in the present study, we have revised the wording throughout the manuscript to avoid implying that we directly measured lipid composition. Instead, we now refer to “genetic perturbation/disruption of sphingolipid metabolism” or “knockdown of enzymes involved in sphingolipid metabolism” when describing our experimental interventions.

      We agree that lipidomic analyses will be important in future studies to define how perturbation of sphingolipid metabolism changes lipid composition in C. elegans and human cells. However, lipidomics itself would not directly establish which lipid changes causally drive the membrane rigidification observed in our study. Membrane fluidity is a biophysical property determined by the combined composition of the membrane, including lipid abundance, saturation, acyl-chain length, head groups, sterol content, and membrane-associated proteins. Thus, even if lipidomics identified changes in sphingolipid profiles, these changes could not be directly translated into a predictable effect on membrane fluidity without additional biophysical validation, using Laurdan dye imaging or FRAP. Moreover, whole-cell or whole-animal lipidomics would not resolve whether the relevant lipid changes occur specifically at endolysosomal membranes, which are the focus of our study.

      We have now clarified this point in the Discussion. Specifically, we state that “even detailed lipidomics would not by itself identify which lipid changes are responsible for the observed membrane rigidification” and that future lysosome-enriched or organelle-specific lipidomic approaches should be combined with direct manipulation of candidate lipid species, followed by measurements of membrane fluidity and rupture, to determine which lipid changes causally contribute to endolysosomal membrane rigidification. In the present study, we therefore focused on direct quantitative biophysical readouts of membrane properties in C. elegans. We used FRAP of the lysosomal membrane protein LAAT-1::mCherry to assess lateral mobility within lysosomal membranes and showed that knockdown of sphingolipid-metabolism genes increased the time to half-maximal recovery, indicating reduced lysosomal membrane fluidity. Notably, knockdown of genes involved in both sphingolipid biosynthesis and sphingolipid degradation increased membrane rigidity. This makes it unlikely that the observed rigidification is caused by accumulation or depletion of a single shared lipid species. Rather, perturbations at different steps of sphingolipid metabolism may lead to distinct lipidomic changes that nevertheless converge on a common biophysical outcome: reduced endolysosomal membrane fluidity. In parallel, we used C-Laurdan imaging to quantify membrane order and found that SPHK2 knockdown in SH-SY5Y human neuroblastoma cells increased GP values, consistent with increased membrane rigidity. Two-channel thresholding of LysoTracker-positive compartments further showed that SPHK2 knockdown increased GP values in lysosome-associated regions.

      Thus, although lipidomics will be valuable to define the underlying lipid changes in future work, the current data already provide convergent quantitative evidence from independent membrane-fluidity readouts across C. elegans and human cell models. This cross-model consistency strengthens the robustness and reproducibility of the central conclusion that perturbation of sphingolipid metabolism alters endolysosomal membrane properties and promotes membrane rupture.

      The conclusions of the study rely heavily on imaging-based assays, including FRAP, C-Laurdan, and fluorescence microscopy. While these approaches provide valuable spatial and qualitative insights, they are inherently indirect and subject to interpretive limitations. To strengthen the mechanistic claims, the authors should incorporate additional biochemical or quantitative approaches. For example, lipidomics would allow direct measurement of membrane lipid composition changes, and western blotting or quantitative proteomics could assess levels of membrane-associated proteins involved in endolysosomal function or stress responses. Including such data would significantly improve the robustness and reproducibility of the study's conclusions.

      We agree that lipidomic and proteomic analyses will be important in future studies to define which sphingolipid species and/or membrane-associated proteins contribute to the observed rigidification of endolysosomal membranes. In response to this point, we have revised the wording throughout the manuscript to more precisely distinguish our experimental interventions from inferred changes in lipid composition. Because we did not directly measure lipid composition in the present study, we now refer more specifically to “genetic perturbation/disruption of sphingolipid metabolism” or “knockdown of enzymes involved in sphingolipid metabolism” when describing our data, rather than implying that global sphingolipid homeostasis was directly quantified. We retain “sphingolipid imbalance” only in interpretive or model-based statements where appropriate.

      However, we respectfully disagree that the current data are only qualitative. FRAP and C-Laurdan GP imaging are established quantitative biophysical approaches: FRAP provides quantitative parameters such as the time to half-maximal recovery and the mobile fraction, whereas C-Laurdan GP provides a ratiometric measurement of membrane lipid order and packing. Similarly, the Galectin puncta assay is an established quantitative readout of lysosomal membrane permeabilization. Thus, while these approaches are imaging-based, they provide quantitative readouts of membrane mobility, membrane order, and membrane rupture, respectively.

      We also note that lipidomic and proteomic profiling, although valuable, would not by itself establish which lipid or protein changes causally drive the membrane rigidification observed in our study. Membrane fluidity is an emergent biophysical property determined by the combined composition of the membrane, including lipid abundance, saturation, acyl-chain length, head groups, sterol content, and membrane-associated proteins. Therefore, an increase or decrease in a given lipid or protein species cannot be directly translated into a predictable change in membrane fluidity without additional biophysical validation. This point is further supported by our observation that knockdown of genes involved in both sphingolipid biosynthesis and sphingolipid degradation increased endolysosomal membrane rigidity. These perturbations would be expected to affect lipid composition in different, possibly even opposing, ways, making it unlikely that the shared rigidification phenotype is caused by accumulation or depletion of one single lipid species. Rather, distinct lipidomic changes may converge on a common biophysical outcome: reduced endolysosomal membrane fluidity.

      We have clarified this point in the Discussion and now state that future lysosome-enriched or organelle-specific lipidomic/proteomic approaches should be combined with direct manipulation of candidate lipid or protein species, followed by measurements of membrane fluidity and rupture, to determine which changes causally contribute to endolysosomal membrane rigidification. Such experiments would address the distinct question of which molecular components mediate the effect. By contrast, the central aim of the present study was to test whether genetic perturbation of enzymes involved in sphingolipid metabolism alters membrane fluidity and thereby promotes endolysosomal rupture and tau seeding.

      For this question, direct biophysical measurements of membrane fluidity and quantitative readouts of membrane rupture are the most relevant assays. We therefore used complementary quantitative approaches in two distinct model systems: FRAP of the lysosomal membrane protein LAAT-1::mCherry in C. elegans and C-Laurdan GP imaging in human cells. The fact that perturbing sphingolipid metabolism reduced endolysosomal membrane fluidity in C. elegans and increased lysosome-associated membrane rigidity in human cells supports the robustness and reproducibility of the central conclusion across independent model systems. In the revised manuscript, we further strengthened the human-cell data by adding SH-SY5Y neuroblastoma cells as a neuronal-like model and by combining C-Laurdan imaging with LysoTracker-based analysis to assess lysosome-associated membrane properties.

      To further address causality, we manipulated membrane fluidity independently of sphingolipid metabolism enzymes using fatty acid supplementation. Increasing membrane rigidity with PA exacerbated tau-induced endolysosomal rupture and seeded aggregation, whereas increasing membrane fluidity with ALA reduced tau-induced membrane rigidification, endolysosomal rupture, and seeded aggregation. Thus, the revised manuscript combines genetic perturbation of sphingolipid metabolism, quantitative membrane-fluidity measurements, whole-cell and lysosome-associated C-Laurdan analysis, and Galectin-based rupture assays across complementary models.

      Regarding lysosomal function, we agree that functional readouts are informative, but lysosomal membrane rupture and global lysosomal degradative capacity are related but not identical readouts. This distinction is supported by Yong et al., who reported that lipid dysregulation can induce lysosomal membrane permeabilization and lysosomal accumulation of endogenous protein aggregates without broadly impairing core lysosomal or proteasomal functions [2]. Accordingly, the absence of overt defects in general lysosomal activity would not necessarily exclude membrane damage.

      The human cell experiments were performed exclusively in HEK293T cells, which are not physiologically relevant for modeling Alzheimer's disease or lysosomal function in neurons. Given that the study aims to draw conclusions related to tau aggregation and lysosomal membrane integrity, the use of a more disease relevant cellular model is essential. There are several established AD-relevant cell models, including iPSCderived neurons, neuroblastoma lines expressing tau, or microglial models, which would better reflect the cellular context of tauopathies. Validation of key findings in at least one of these systems would substantially enhance the biological relevance and translational impact of the study.

      We have expanded and clarified the human cell data in the revised manuscript. Specifically, we now include SH-SY5Y human neuroblastoma cells for key C-Laurdan experiments assessing membrane rigidity after SPHK2 knockdown. We also show that recombinant tau fibrils increased membrane rigidity in SHSY5Y and HEK293T cells, including in LysoTracker-positive compartments.

      Importantly, the HEK293T cells are used for specific, established quantitative assays rather than as a model of neuronal toxicity. In particular, HEK293T sfGFP-LGALS3 cells are used to quantify galectin puncta formation as a readout of endolysosomal rupture, and HEK tau-Venus biosensor cells are used to quantify seeded tau aggregation. Thus, SH-SY5Y cells and HEK293T cells are used for complementary purposes: SHSY5Y cells provide a more neuronal-like human cell context for membrane-rigidity measurements, whereas HEK293T reporter/biosensor cells provide robust quantitative assays for galectin puncta formation and tau seeding.

      In addition, tau-associated neuronal dysfunction and toxicity were assessed in vivo, in functional C. elegans touch receptor neurons. In the revised manuscript, we show that ALA supplementation mitigated the age-dependent touch-response deficit and reduced neurotoxicity in animals expressing F3ΔK281::mCherry in touch receptor neurons. We have also revised the wording throughout the manuscript to avoid implying that HEK293T cells are used to model neuronal toxicity.

      Finally, the relevance of these hits to human neuronal tau seeding is also supported by our previous study, in which conserved hits from the C. elegans screen, including sphingosine kinase perturbation, were validated in human iPSC-derived neurons for their effect on seeded tau aggregation [1]. We now cite this published study where appropriate. Together, the revised manuscript combines neuronal-like human SH-SY5Y cells, established HEK293T tau-seeding and galectin reporter assays, in vivo neuronal readouts in C. elegans, and prior validation in human iPSC-derived neurons, thereby strengthening the biological relevance of the conclusions while using each model for the assay in which it is most informative.

      The authors reported that PUFA supplementation rescues neurotoxic phenotypes by increasing membrane fluidity. However, the data supporting this claim rely entirely on confocal imaging, shown in both the main and supplemental figures. To substantiate the mechanistic link between PUFA treatment and improved lysosomal membrane properties, the authors should include functional assays demonstrating that PUFAs are indeed incorporated into lysosomal membranes. Additionally, lipidomics analysis would be valuable to identify which lipid species are altered upon supplementation and correlate these changes with the observed phenotypic rescue. Furthermore, the conclusion that PUFAs rescue "neurotoxic phenotypes" is not appropriate based on data derived solely from HEK293T cells, which are not neuronal. To make claims about tau-related neurotoxicity, the authors should validate their findings in a more relevant neuronal model, such as SH-SY5Y neuroblastoma cells expressing tau or iPSC-derived neurons. This would better reflect the cellular environment of Alzheimer's disease and provide stronger support for the proposed therapeutic potential of PUFA supplementation.

      We agree that PUFA supplementation can have effects beyond membrane fluidity and that our data do not directly demonstrate incorporation of ALA into lysosomal membranes. We have therefore revised the Discussion to explicitly acknowledge this limitation. In the revised text, we state that “PUFAs can also influence lipid signaling, oxidative stress responses, and broader membrane remodeling” and that we “cannot exclude additional direct or indirect effects of ALA.” At the same time, we note that the opposing effects of PA and ALA, together with the sphingolipid-metabolism knockdown data, support membrane fluidity as a major determinant of endolysosomal membrane integrity and rupture in our models. To strengthen the link between ALA and lysosome-associated membrane properties, we combined CLaurdan imaging with LysoTracker-based analysis. In the revised Results, we show that ALA prevented tau-induced membrane rigidification and that LysoTracker-based analysis indicated effects on lysosome-associated membrane properties. ALA also reduced tau-induced endolysosomal rupture and seeded aggregation in human cell models.

      Regarding lipidomics, we refer to our response above and to the revised Discussion. We agree that lipidomics would be valuable to identify ALA-induced lipid changes, but such data would not by itself establish how these changes affect membrane fluidity without additional biophysical validation.

      Finally, we clarify that our conclusion regarding tau-associated neuronal dysfunction and toxicity is not based on HEK293T cells. HEK293T cells were used for established quantitative assays of Galectin puncta formation and seeded tau aggregation. The neurotoxicity experiments were performed in vivo in C. elegans touch receptor neurons, where ALA supplementation reduced galectin foci formation, mitigated age-dependent touch-response deficit and reduced neuronal toxicity.

      While the authors demonstrate that ALA supplementation mitigates neurotoxicity in C. elegans expressing aggregated tau (F3ΔK281::mCherry), the current data are not sufficient to conclude that ALA directly rescues tau aggregation toxicity via a lysosome-specific mechanism. It remains unclear how lipid composition is altered upon ALA treatment and whether these changes correlate with functional improvement of lysosomal pathways. The manuscript does not provide mechanistic insight into how ALA enhances lysosomal health or attenuates endolysosomal damage. Moreover, supplementation with PUFAs like ALA can activate a wide range of cellular processes beyond lysosomal function, including alterations in membrane fluidity, signaling cascades, and oxidative stress responses. The authors should clarify how they distinguish the lysosome-related effects from these alternative pathways. For example, did they observe specific lysosomal markers or structural improvements in lysosomes upon ALA treatment?

      Additional data or controls would be necessary to support a lysosome-specific protective mechanism and to exclude the involvement of other PUFA-responsive pathways in the observed phenotypes.

      We agree that our data do not prove that ALA acts exclusively through a lysosome-specific mechanism or that ALA is directly incorporated into lysosomal membranes. We have therefore revised the manuscript to avoid this interpretation and explicitly acknowledge alternative PUFA-responsive pathways. In the revised Discussion, we state that “PUFAs can also influence lipid signaling, oxidative stress responses, and broader membrane remodeling” and that we “cannot exclude additional direct or indirect effects of ALA.” We further conclude more cautiously that the opposing effects of PA and ALA, together with the sphingolipid metabolism perturbation data, support membrane fluidity as a major determinant of endolysosomal membrane integrity and rupture in our models.

      To strengthen the lysosome-related aspect of the mechanism, we added LysoTracker-based analysis to the C-Laurdan experiments. In the revised Results, ALA prevented tau-induced membrane rigidification, and LysoTracker-based analysis indicated that ALA also affected lysosome-associated membrane properties. ALA further reduced tau-induced Galectin puncta formation and seeded tau aggregation in human cell models. These data support an effect of ALA on lysosome-associated membrane order and rupture, while not excluding additional effects through lipid signaling, oxidative stress responses, or other PUFA-responsive pathways.

      Regarding lipid composition, we refer to the revised Discussion and our response above. We agree that lipidomics would be valuable to identify ALA-induced lipid changes, but such analyses would need to be organelle-specific and combined with biophysical validation to determine how candidate lipid changes affect membrane fluidity and rupture.

      Finally, we clarify that our conclusion regarding tau-associated neuronal dysfunction and toxicity is based on the C. elegans experiments, not on HEK293T cells. HEK293T cells were used for quantitative Galectin puncta and tau-seeding assays, whereas neuronal dysfunction and toxicity were assessed in vivo in touch receptor neurons in C. elegans. In the revised Results, we state that ALA supplementation mitigated galectin foci formation, age-dependent touch-response deficit and reduced neuronal toxicity in animals expressing F3ΔK281::mCherry in these neurons.

      Reviewer #2 (Public review):

      Tittelmeier et al. investigated the role of sphingolipid (SL) metabolism in the maintenance of endolysosomal vesicle integrity. They find that both impaired SL biosynthesis and degradation in C. elegans, decrease the fluidity of endolysosomal membranes and promote their rupture, while it has little effect on plasma membrane fluidity. Endolysosomal membrane fluidity is also negatively affected in human cells upon knockdown (KD) of a gene (SPHK2) involved in the SL degradation pathway. Aggregated forms of tau in both models (C. elegans and human cells) can also cause rigidification of the endolysosomal membrane, with SL homeostasis disruption having an additive effect, exacerbating endolysosomal rupture. Notably, KD of SPHK2 also increased the formation of tau foci, suggesting that compromised endolysosomal integrity may promote tau aggregation. These data provide a clearer understanding of how genetic manipulation of SL metabolism affects endolysosomal membranes and their rigidification in the context of tau aggregation. Supplementation of polyunsaturated fatty acids (PUFAs), which has a beneficial effect on Alzheimer's patients, improved membrane fluidity and reduced tau propagation in human cells and tau-associated neurotoxicity in C. elegans, suggesting a possible mechanism of action.

      Overall, the conclusions of this paper are supported by the data, with a few aspects requiring further clarification and elaboration.

      (1) A reference to Figure S2E-G, which shows that KD of SL biosynthesis genes do not affect the plasma membrane, is missing from the main text.

      We thank the reviewer for pointing this out. We have added the reference to the respective figures in the main text when discussing the plasma membrane FRAP experiments.

      (2) In Figure 3C, lipofectamine alone shows that it increases membrane rigidity (increased GP values), not membrane fluidity.

      We thank the reviewer for pointing out this incorrect wording. A higher GP index indicates increased membrane order/rigidity, not increased membrane fluidity. Since the revised main figure now includes the SH-SY5Y data, in which Lipofectamine alone did not significantly alter GP values, we removed the misleading statement from the Results. Importantly, all siRNA-mediated knockdown experiments were compared to matched control siRNA conditions exposed to the same transfection reagent. Thus, the effect attributed to SPHK2 KD is assessed relative to the appropriate Lipofectamine-containing control condition.

      (3) In Figure 3F, the EV cntl condition expressing F3:mCh tau should have increased LGALS3 foci compared to the mCh EV cntl according to Ref (20) and its Figure 2G (at least for Day 5 animals), which would be indicative of the tau spreading in hypodermal tissue. What C. elegans age was examined in Figure 3F? Can the authors provide evidence of the transmission of the F3:mCh tau from the touch receptor neurons to the hypodermis in the EV [similar to Figure 2C & D from Ref (20)] and compare it to the KDs? Otherwise, it seems that KD of SL genes impacts not only endolysosomal rupture but significantly affects tau accumulation/spreading as well (e.g., shown later in HEK cells, where SPHK2 KD increases the formation of tau-Venus foci).

      We thank the reviewer for raising this important point. The analysis referred to by the reviewer has now been moved to the revised C. elegans section and is presented as Figure 4A and B. The experiments were performed in the reporter strain used in our genome-wide screen In Ref (20), now published in Autophagy [1]. We clarified the purpose of the reporter strain and the relationship between tau transmission and the galectin puncta readout. In the revised manuscript, we now state: “In this strain, endolysosomal membrane damage is monitored in the hypodermis by expression of human galectin-3 fused to superfolder-GFP (sfGFP::LGALS3). The animals also express an aggregation-prone tau fragment fused to mCherry (F3ΔK281::mCherry) in touch receptor neurons, which is transmitted to the hypodermis, as described previously [20].” We further clarify that sfGFP::LGALS3 puncta formation, not F3ΔK281::mCherry, is the readout of endolysosomal rupture: “Under steady-state conditions, sfGFP::LGALS3 remains diffusely distributed throughout the cytosol. Upon endolysosomal damage, luminal β-galactosides become exposed and recruit sfGFP::LGALS3 into visible puncta, providing a sensitive readout of vesicle rupture”.

      Furthermore, we now better explain that transmitted tau sensitizes endolysosomal membranes to additional perturbations rather than necessarily inducing a strong lysosomal rupture phenotype on its own. In the experiments shown in Figure 4A and B, we compare F3ΔK281::mCherry animals with matched mCherry-only control animals that also express sfGFP::LGALS3 in the hypodermis. We now state: “We compared animals expressing F3ΔK281::mCherry in touch receptor neurons, from where it is transmitted to the hypodermis, with matched controls expressing mCherry alone in the same neurons. In both strains, sfGFP::LGALS3 is expressed in the hypodermis to monitor endolysosomal membrane damage.” We have also clarified the age of the animals in the revised figure legends.

      To experimentally address whether the enhanced rupture phenotype could be explained by altered tau transmission, we quantified hypodermal F3ΔK281::mCherry levels after sphk-1 RNAi (new Figure 4C, D). Importantly, sphk-1 RNAi did not increase hypodermal F3ΔK281::mCherry levels, arguing that the enhanced rupture phenotype is not due to increased tau transmission. Moreover, C. elegans neurons are largely refractory to systemic RNAi under the conditions used here [3]. We have added this important information to the Discussion. Specifically, the revised manuscript states that “the enhanced rupture phenotype is unlikely to result from a direct effect of RNAi on neuronal F3ΔK281::mCherry expression, as C. elegans neurons are largely refractory to systemic RNAi under the conditions used here,” supporting the interpretation that the RNAi treatments primarily affect endolysosomal integrity in the recipient tissue rather than neuronal tau expression itself.

      Finally, we would like to clarify that the increased tau-Venus foci in HEK cells should not be interpreted as a direct induction of tau aggregation by SPHK2 KD. Only upon addition of recombinant tau fibrils did SPHK2 KD significantly increase tau-Venus foci formation (Figure 3 H, I). This is consistent with the control experiments performed in human iPSCs in our previous study and supports our interpretation that perturbation of sphingolipid metabolism increases susceptibility to seeded tau aggregation by promoting endolysosomal rupture and tau seed escape, rather than by directly increasing tau aggregation or tau transmission.

      (4) Sphingolipids are essential membrane components and signaling molecules. Does KD of SL genes in C. elegans and the subsequent endolysosomal rupture cause any major, intermediate, or minor defects/phenotypes (in non-aggregation prone models, w/t.)?

      We agree that sphingolipids are essential membrane components and signaling molecules and that perturbing sphingolipid metabolism can have broader physiological consequences. In the revised manuscript, we address this point in two ways.

      First, we directly tested whether SL gene knockdown can induce endolysosomal rupture independently of aggregation-prone tau by using matched control animals expressing mCherry alone in touch receptor neurons while also expressing sfGFP::LGALS3 in the hypodermis. In these animals, knockdown of most SL-related hits resulted in nearly all animals displaying hypodermal sfGFP::LGALS3 foci, indicating that perturbation of SL metabolism can compromise endolysosomal integrity in the absence of transmitted F3ΔK281::mCherry (Figure 4A, B).

      Second, we have added a Discussion paragraph to place these findings into a broader physiological context. We now clarify that endolysosomal membrane rupture and global lysosomal function are related but not identical readouts. In support of this distinction, we discuss work showing that lipid dysregulation can induce lysosomal membrane permeabilization and lysosomal accumulation of endogenous protein aggregates without broadly impairing core lysosomal or proteasomal function [2]. Thus, membrane damage can occur even when general lysosomal activity is not overtly disrupted.

      We also discuss a recent study published during the revision of this manuscript that independently identified SPHK-1 as an important regulator of lysosomal integrity in C. elegans, showing that strong sphk1 loss-of-function causes lysosomal sphingosine accumulation, membrane rupture, impaired degradative function, cargo accumulation, developmental defects, and reduced lifespan [4].

      Importantly, while that study focused on a strong loss-of-function mutation in a single SL-metabolism gene, our data show that knockdown of multiple SL-metabolism genes, including genes involved in both SL biosynthesis and degradation, converges on reduced endolysosomal membrane fluidity and increased rupture. This suggests that the observed membrane rigidification and rupture are not specific to one mutant background but represent a broader consequence of perturbing SL metabolism at multiple points. A systematic characterization of all organismal phenotypes caused by each SL gene knockdown was beyond the scope of the present study. Therefore, the revised manuscript now makes clear that the study focuses on endolysosomal membrane fluidity and rupture because these membrane-level changes are directly linked to tau seed escape and seeded tau aggregation, while broader physiological consequences may vary depending on the strength and context of the perturbation.

      Reviewer #3 (Public review):

      Summary:

      The authors set off with an analysis of the lysosomal integrity upon knockdown of genes of the sphingolipid metabolic pathway that they identified in a previous (yet unpublished) work of an RNA screen using a new C. elegans Tau model. They then used cell culture and C. elegans experiments to study the link between lysosomal rupture and Tau propagation.

      Strengths:

      The authors use two complementary model systems and use probes to assess membrane rigidity that allow a quick assessment of the membrane dynamics and offer the opportunity to treat the cells with lipids, RNAi. Tau seeds, etc.

      Weaknesses:

      The main weakness is that this work builds on not-yet-peer-reviewed manuscript that established a new C. elegans Tau model and RNAi screen that aimed to identify genes involved in the propagation of Tau.

      This reviewer misses essential information of the C. elegans Tau strain (not included in the method section): e.g., promoter used for the expression, information on the used Tau variant, expression pattern, and aggregation, etc.

      We thank the reviewer for raising this point. The related study establishing the C. elegans tau transmission model and RNAi screen has now been peer-reviewed and published in Autophagy [1]. We now cite the published article throughout the revised manuscript instead of the previous preprint.

      We also agree that the current manuscript should be understandable without requiring the reader to consult the previous paper for the basic logic of the model. We therefore added a clearer introduction of the reporter strain in the Results. Specifically, we now explain that the strain expresses the aggregation-prone tau fragment F3ΔK281::mCherry in touch receptor neurons, that this tau fragment is transmitted to the hypodermis, and that endolysosomal membrane damage is monitored in the hypodermis using sfGFP::LGALS3. We further clarify that sfGFP::LGALS3 remains diffuse under steady-state conditions and forms puncta upon endolysosomal membrane damage, when luminal β-galactosides become exposed. Thus, the revised manuscript now provides the key information needed to understand the experimental system used here, while the published Autophagy paper is cited for the full characterization of the tau transmission model, expression pattern, aggregation properties, and original genome-wide RNAi screen.

      Throughout the study, I missed data on:

      (1) Effect of the knockdown on Tau expression, localisation (with lysosomal membrane?), aggregation, and proteotoxicity. The effect of the RNAi-mediated knockdown could also simply lead to a reduced expression of Tau that, in turn, leads to suppressed propagation.

      We agree that it is important to distinguish effects on tau expression/transmission from effects on endolysosomal membrane integrity. In the C. elegans experiments, F3ΔK281::mCherry is expressed in touch receptor neurons and transmitted to the hypodermis, where sfGFP::LGALS3 reports endolysosomal membrane damage. We now describe this more clearly in the revised Results.

      The RNAi treatments target genes involved in sphingolipid metabolism under systemic RNAi conditions. Because C. elegans neurons are largely refractory to systemic RNAi in the absence of sensitizing backgrounds [3], which we did not use, a direct RNAi-mediated reduction of neuronal F3ΔK281::mCherry expression is very unlikely. We have added this point to the Discussion, stating that “the enhanced rupture phenotype is unlikely to result from a direct effect of RNAi on neuronal F3ΔK281::mCherry expression, as C. elegans neurons are largely refractory to systemic RNAi under the conditions used here.”

      Experimentally, we also tested whether sphingolipid perturbation alters transmitted tau levels (new Figure 4C, D). Specifically, we quantified hypodermal F3ΔK281::mCherry after sphk-1 RNAi and found no increase, arguing that the enhanced rupture phenotype is not due to increased tau transmission.

      Moreover, in the cell-based tau-Venus assay, SPHK2 knockdown alone did not induce detectable tau aggregation in the absence of exogenously added tau fibrils (Figure 3H, I). Only upon addition of recombinant tau fibrils did SPHK2 knockdown significantly increase tau-Venus foci formation. We now state this explicitly in the revised Results and conclude that disruption of sphingolipid metabolism is not sufficient on its own to initiate detectable tau aggregation under the conditions tested here but rather increases cellular susceptibility to seeded tau aggregation when tau fibrils are present.

      Together, these data argue against a direct effect of sphingolipid gene knockdown on tau expression or spontaneous tau aggregation. Instead, they support our interpretation that perturbation of sphingolipid metabolism compromises endolysosomal membrane integrity, thereby facilitating tau seed escape and seeded aggregation when tau seeds are present.

      (2) A quantification of RNAi knockdown is needed to judge the efficiency of the RNAi, in particular for the combinatorial RNAi experiments involving 2 and even 4 genes in parallel. Ideally, these analyses should be validated with mutants for these genes.

      We agree that RNAi efficiency can vary between clones and that this is particularly relevant for combinatorial RNAi experiments targeting two or more potentially redundant genes. We have added this limitation to the Results section and now state: “Because KD efficiency was not assessed for the individual RNAi clones or co-RNAi combinations, these experiments do not allow comparison of relative RNAi strength or inference of the relative importance of individual genes. Thus, the conclusions drawn from these RNAi experiments are qualitative: specific single or combined KDs can promote endolysosomal rupture, whereas the absence of a detectable phenotype after RNAi cannot exclude gene involvement, as KD may have been insufficient.”

      Where mutant strains were available, we performed genetic validation. Specifically, a sphk-1 mutant available at CGC (CZ24969; sphk-1(ju831)) also showed increased hypodermal sfGFP::LGALS3 puncta (new Figure S1A), supporting the RNAi-based conclusion that genetic perturbation of sphingolipid metabolism compromises endolysosomal integrity. Corresponding mutant strains were not available for the other selected hits. Importantly, most hits also induced sfGFP::LGALS3 foci in human HEK293T cells as assessed in our previous study [1], providing additional support that the observed effects are not random RNAi artifacts.

      Further:

      (3) Figure 4 H, I: Would Tau also aggregate in the absence of externally added Tau?

      No. In the tau-Venus biosensor cell line, SPHK2 knockdown alone did not increase tau-Venus foci formation (now Figure 3H, I). Tau-Venus foci increased only after addition of recombinant tau fibrils and were further enhanced by SPHK2 knockdown. We now state this explicitly in the Results.

      (4) How specific is the effect for Tau? It would help if the authors could assess other amyloid proteins.

      We agree that similar membrane-level mechanisms may apply to other amyloid assemblies. We have therefore added recent literature to the Discussion supporting the broader concept that intralysosomal amyloid assemblies can physically deform and rupture lysosomal membranes. The revised manuscript states: “This interpretation is consistent with recent ultrastructural studies showing that intralysosomal amyloid assemblies can physically deform and rupture lysosomal membranes.” We further clarify that “whether this mechanism is specific to tau or also applies to other amyloid assemblies remains to be determined.”

      Whether perturbation of SL metabolism similarly affects endolysosomal escape and seeded aggregation of other disease-associated amyloid proteins is an important question that we plan to address in future work. However, these experiments require additional disease-specific models, aggregation assays, and validation, and are therefore beyond the scope of the present revision.

      (5) The connection between sphingolipids and AD is not new. See He et al, 2010, Neurobiol. Aging + numerous publications and also not between Tau seeding and lysosomal rupture: Rose et al., PNAS 2024 (that has been cited by the authors).

      We agree and our manuscript does not aim to establish these associations as new. We state explicitly that alterations in sphingolipid metabolism have been reported in aging and AD, and that endolysosomal rupture is increasingly recognized as a critical step in tau seed escape and propagation.

      The novelty of our study lies in mechanistically connecting these two previously established areas. Specifically, we show that genetic perturbation of enzymes involved in sphingolipid metabolism reduces endolysosomal membrane fluidity, promotes membrane rupture, and thereby increases susceptibility to tau seed escape and seeded aggregation. We have revised the Introduction and Discussion to better emphasize this mechanistic contribution.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Figure formatting and annotation need improvement. Panel letters throughout the figures should be in uppercase, and gene names in pathway diagrams should be italicized for consistency. Several scale bars are missing, including in Figures 1C, 2A, and 2H, and should be clearly indicated in the figures and legends. In Figure 1C, the age of the worms used in the assay is not specified. While the Methods section mentions "age-synchronized animals," the precise age at the time of imaging or experimentation is not stated. It would strengthen the study to explore whether membrane integrity phenotypes vary between young adults (day 1) and older adults (day 7 or 10) across the different conditions. Figure 1B lacks sufficient detail describing the galectin puncta assay used. A brief explanation of the assay rationale and readout would help contextualize the findings.

      We thank the reviewer for pointing this out. We have revised the figures and figure legends accordingly by standardizing panel labels, adding scale bars where missing, and providing the age of animals used in the assays. We also expanded the description of the galectin puncta assay in the Results to explain the rationale and readout of sfGFP::LGALS3 puncta formation.

      Regarding the reviewer’s suggestion to compare young and aged animals, we agree that age-dependent changes in endolysosomal membrane integrity are an interesting question. However, the purpose of the present study was to investigate how perturbation of sphingolipid metabolism affects endolysosomal membrane fluidity and rupture under the assay conditions used in our original screen. A systematic comparison across aging is beyond the scope of the current revision. We have therefore clarified the animal ages used in the relevant figure legends and Methods.

      In Figure S1A, the authors show co-knockdown of multiple genes, including one condition with simultaneous RNAi against four targets. Because different RNAi clones can vary in knockdown efficiency, it is important to provide validation of gene knockdown levels (e.g., by qRT-PCR) shown in both panels a and b.

      We agree that RNAi efficiency can vary between clones and that this is particularly relevant for combinatorial RNAi experiments targeting two or more potentially redundant genes. We have added this limitation to the Results section and now state: “Because KD efficiency was not assessed for the individual RNAi clones or co-RNAi combinations, these experiments do not allow comparison of relative RNAi strength or inference of the relative importance of individual genes. Thus, the conclusions drawn from these RNAi experiments are qualitative: specific single or combined KDs can promote endolysosomal rupture, whereas the absence of a detectable phenotype after RNAi cannot exclude gene involvement, as KD may have been insufficient.”

      Where mutant strains were available, we performed genetic validation. Specifically, a sphk-1 mutant available at CGC (CZ24969; sphk-1(ju831)) also showed increased hypodermal sfGFP::LGALS3 puncta (new Figure S1A), supporting the RNAi-based conclusion that genetic perturbation of sphingolipid metabolism compromises endolysosomal integrity. Corresponding mutant strains were not available for the other selected hits. Importantly, most hits also induced sfGFP::LGALS3 foci in human HEK293T cells as assessed in our previous study [1], providing additional support that the observed effects are not random RNAi artifacts.

      In Figure 2E, the FRAP recovery curves show only ~60% recovery in controls after 25 seconds, and an even lower recovery (~40%) in hpo-8 and spp-10 RNAi conditions. The authors should discuss why the recovery is incomplete and what it implies about the mobile fraction of the protein or membrane components in these conditions.

      We agree that incomplete FRAP recovery is informative. For this reason, we report both the time to half-maximal recovery (thalf) and the maximal recoverable fluorescence signal. Increased thalf indicates reduced lateral mobility of LAAT-1::mCherry within the lysosomal membrane, consistent with reduced membrane fluidity. In addition, a reduced maximal recovery suggests that a larger fraction of the reporter is immobile or only slowly mobile during the time window analyzed. This may reflect stronger confinement of LAAT-1::mCherry within even more rigid membrane domains. However, because RNAi efficiency may differ between clones and we have not assessed their individual KD efficiency, we avoid overinterpreting differences in the absolute strength of recovery defects between individual KDs. Instead, we conclude that KD of sphingolipid-metabolism genes identified in our screen consistently reduces lysosomal membrane fluidity, as reflected by increased thalf and, in some cases, reduced maximal recovery.

      In Figure S3A, the Western blot for SPHK2 shows unequal loading between the control and siSPHK2 lanes. The blot should be normalized to a loading control and quantified to demonstrate knockdown efficiency.

      We have quantified SPHK2 levels relative to GAPDH across independent experiments and present the normalized quantification (Figure S3C-E).

      Key experimental details are missing from the manuscript. The strains of C. elegans and RNAi bacteria used were not described, and there is no information on biological replicates. The authors should clarify how many times each experiment was performed and provide more transparency on experimental reproducibility.

      We thank the reviewer for pointing this out. The C. elegans strains and RNAi bacterial clones used in this study were established and fully described in our previous study, which has now been published in Autophagy [1]. We now cite the published article throughout the revised manuscript and have added additional information in the Results section to explain the key features of the strains used here.

      We have also revised the Methods and figure legends to improve transparency regarding experimental details. The figure legends include the number of biological replicates, the number of animals or cells analyzed, and the statistical tests used for each experiment. In addition, the Statistical Analysis section in the Methods now summarizes how replicate numbers and sample sizes are reported across the study. Finally, the source details for the strains and RNAi clones used in this study are now provided in Tables S1 and S2, respectively. These revisions should improve the experimental clarity and reproducibility of the data shown.

      References:

      (1) Sandhof CA, Martin N, Tittelmeier J, Schlueter A, Pezzali M, Schoendorf DC, et al. A novel C. elegans model for MAPT/Tau spreading reveals genes critical for endolysosomal integrity and seeded MAPT/Tau aggregation. Autophagy. 2025;21(12):2963-81. Epub 20250904. doi: 10.1080/15548627.2025.2551676. PubMed PMID: 40851193; PubMed Central PMCID: PMCPMC12758218.

      (2) Yong J, Villalta JE, Vu N, Kukurugya MA, Olsson N, Lopez MP, et al. Impairment of lipid homeostasis causes lysosomal accumulation of endogenous protein aggregates through ESCRT disruption. eLife. 2024;12. Epub 20241223. doi: 10.7554/eLife.86194. PubMed PMID: 39713930; PubMed Central PMCID: PMCPMC11666243.

      (3) Calixto A, Chelur D, Topalidou I, Chen X, Chalfie M. Enhanced neuronal RNAi in C. elegans using SID-1. Nat Methods. 2010;7(7):554-9. doi: 10.1038/nmeth.1463. PubMed PMID: 20512143; PubMed Central PMCID: PMC2894993.

      (4) Li Y, Zhang J, Li M, Yang L, Wang X. Sphingosine kinase SPHK-1 maintains sphingolipid metabolism to protect lysosome membrane integrity in C. elegans. Mol Biol Cell. 2026;37(1):ar1. Epub 20251105. doi: 10.1091/mbc.E25-04-0182. PubMed PMID: 41191545; PubMed Central PMCID: PMCPMC12696880.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In their manuscript, Zhou and colleagues present a detailed look at how the JSP functions differently in the various cells of a breast tumor. The authors have effectively shown that the JSP acts as a double-edged sword, as it helps T cells fight cancer but also allows tumor cells to grow and avoid ferroptosis. These findings are important because they identify a useful biomarker to predict how TNBC patients might respond to PD-1 inhibitors.

      Strengths:

      This work is important because it provides a clear explanation for the conflicting roles of the JSP in the tumor environment. The evidence is solid, as it combines data from thousands of patients with single-cell analysis and lab experiments to confirm the role of STAT4 in cancer progression and immunity.

      Comments on revised version:

      The authors made a significant effort to improve the manuscript. My comments were sufficiently addressed.

      We sincerely appreciate your careful review and positive feedback. We are glad to hear that you are satisfied with the revised manuscript and acknowledge the scientific value and solid evidence of our work. Thank you again for all your efforts and valuable suggestions.

      Reviewer #2 (Public review):

      Summary:

      The JAK-STAT pathway (JSP) exhibits cell-type-specific functional heterogeneity in breast cancer. This study investigates the JSP in breast cancer and its response to anti-PD‑1 immunotherapy. JSP displays distinct cell‑type heterogeneity: it promotes malignant phenotypes and immunosuppression in tumor cells, while enhancing cytotoxicity and reducing exhaustion in T cells. Elevated JSP expression correlates with improved immunotherapy responses, especially in triple‑negative breast cancer. These findings highlight the paradoxical roles of JSP, indicating that broad inhibition may compromise anti‑tumor immunity.

      Strengths:

      The major strengths of this study include the comprehensive characterization JSP heterogeneity across epithelial, tumor, and T cells in breast cancer. The identification of JSP and STAT4 as predictive biomarkers for immunotherapy response, particularly in triple‑negative breast cancer, provides clinically relevant insights for patient stratification.

      Weaknesses:

      The corresponding content has been revised.

      We sincerely thank you for your detailed review and valuable comments. We greatly appreciate your recognition of the cell-type-specific heterogeneity of the JAK-STAT pathway and the clinical value of JSP and STAT4 as predictive biomarkers for immunotherapy in triple-negative breast cancer. We have thoroughly revised the manuscript according to your previous suggestions, and all raised concerns have been fully addressed.

      Reviewer #3 (Public review):

      Summary:

      This multi-omics study by Zhou et al elucidates the context-dependent roles of the Janus kinase-signal transducer and activator of transcription (JAK-STAT) pathway (JSP) across different cellular compartments in the breast cancer tumor microenvironment. While bulk JSP activity is associated with a favorable prognosis, single-cell analysis reveals a paradoxical landscape: high JSP in T cells drives anti-tumor cytotoxicity and reduces exhaustion, whereas high activity in tumor epithelial cells promotes malignancy and immunosuppression via the MIF-CD74 signaling axis. The JSP score (immune-related) serves as a robust predictive biomarker for response to anti-PD-1 immunotherapy, particularly in triple-negative breast cancer (TNBC). Furthermore, the study identifies the STAT4/SLC47A1 axis as a critical mechanism through which tumor cells resist ferroptosis, facilitating disease progression. These findings suggest that broad JAK-STAT inhibition may be counterproductive in cancer therapeutics; instead, therapeutic success depends on precise modulation and carefully timed interventions to preserve its T-cell-associated functions. This study may inspire future studies to explore specific factors that selectively modulate JAK-STAT activity in immune cells to achieve favorable therapeutic outcomes.

      Strengths:

      Significant therapeutics implications

      Weaknesses:

      Limited molecular mechanisms

      Comments on revised version:

      The authors have addressed my comments

      Many thanks for your careful evaluation and valuable suggestions. We highly appreciate your affirmation of the therapeutic significance of this study. We have fully revised the manuscript to enrich the molecular mechanisms, and all your comments have been properly resolved.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The most content has been revised.

      Minor corrections:

      (1) The icon about "prognosis" in graphic abstract is overly childish.

      The graphical abstract has been redrawn. The inappropriate prognosis icon is deleted accordingly.

      (2) Please double check the whole content to avoid typos. For instance, "2.2" and "2.3" have been repeated twice.

      We appreciate your reminder. The duplicate numbering of 2.2 and 2.3 resulted from Word’s automatic heading feature. We have disabled this function and fixed all repeated section numbers. In addition, we have carefully checked the full text and corrected all typos.

      (3) It will be more interesting if the oncogenic role of STAT4 could be verified via cell cloning assay.

      We appreciate your thoughtful comment. Considering the limited revision time, we cannot add the cell cloning assay in the current version. Our present data sufficiently validate the oncogenic function of STAT4, and the main conclusions remain reliable.

      Reviewer #3 (Recommendations for the authors):

      I recommend publishing the revised manuscript in eLife.

      We sincerely thank you for your positive evaluation and endorsement for the publication of our revised manuscript. We greatly appreciate your rigorous review and insightful comments that have substantially improved the quality and readability of this work.

      We sincerely appreciate all reviewers and editors for your thorough reviewing work and thoughtful feedback. Your suggestions have helped us greatly improve this manuscript. Thank you very much.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary

      In this study, the authors have performed tissue-specific ribosome pulldown to identify gene expression (translatome) differences in the anterior vs posterior cells of the C. elegans intestine. They have performed this analysis in fed and fasted states of the animal. The data generated will be very useful to the C. elegans community, and the role of pyruvate shown in this study will result in interesting follow-up investigations.

      However, several strong claims made in the study are solely based on in silico predictions and are not supported by experimental evidence.

      Strengths:

      Several studies in the past have predicted different functions of the anterior (INT1) vs posterior (INT2-9) epithelial cells of the C. elegans intestine based on their anatomy and ultrastructure, but detailed characterization of differences in gene expression between these cell types (and whether indeed these are different 'cell types') was lacking prior to this study. The genes and drivers identified to be exclusively expressed in the anterior vs posterior segments of the intestine will be very helpful to selectively modulate different parts of the C. elegans intestine in future studies.

      Another strength of this study is the careful experimental design to test how the anterior vs posterior cell types of the intestine respond differently to food deprivation and recovery after return to food. These comparisons between 'states' of a cell in different physiological conditions are difficult to pick up in single-cell analyses due to low sequencing depth, which can fail to identify subtle modulation of gene expression.

      The TRAP-associated bulk RNA-seq approach used in this study is more suitable for such comparisons and provides additional information on post-transcriptional regulation during metabolic stress.

      A key finding of this study is that pyruvate levels modulate the translation state of anterior intestinal cells during fasting. Characterization of pyruvate metabolism genes, especially of the enzymes involved in its mitochondrial breakdown, provides novel insights into how gut epithelial cells respond to the acute absence of food.

      Weaknesses:

      Unlike previous TRAP-seq studies (PMID: 30580965, 36044259, 36977417) that reported sequencing data for both input and IP samples, this study only reports the sequencing data for IP samples. Since biochemical pulldowns are variable across replicates, it is difficult to know if the observed differences between different conditions are due to biological factors or differences in IP efficiency. More importantly, since two different TRAP lines were utilized in this study and a large proportion of the results focus on the differences between the translational profiles of INT1 vs INT2-9 cells, it is essential to know if the IP worked with similar efficiency for both TRAP strains that likely have different expression levels of the HA-tagged ribosomal protein. One way to estimate this would be to perform qRT-PCR of genes that are known to be enriched in all intestinal cells and determine whether their fold-enrichment over housekeeping genes (normalized to input) is similar in INT1 vs INT2-9 TRAP strains and across the fed vs fasted conditions. The authors, in fact, mention variability across biological replicates, due to which certain replicates were excluded from their WGCNA analysis.

      We appreciate the reviewer's comments. We agree that the lack of matched input sequencing libraries limits our ability to directly assess IP efficiency across replicates, conditions, and TRAP strains. However, several features of the dataset support the conclusion that the major differences reported here reflect biological rather than purely technical variation. First, the RPL-22-3xHA construct was integrated into each line to improve consistency across experiments. Second, although the INT2-9 TRAP strain yielded more RNA than the INT1 strain, as expected given the larger number of labeled cells, downstream analyses were performed on normalized count data rather than raw counts. Third, principal component analysis showed robust separation by promoter identity across all conditions, and expected INT1-enriched genes such as ins-7 were recovered in the INT1 dataset. Together, these observations support the interpretation that the TRAP datasets capture reproducible, cell-type-specific differences in ribosome-associated transcripts. Nonetheless, we agree that direct input-normalized measurements would further strengthen the study, and we will explicitly note this as an important limitation.

      It appears that GFP expression is also detectable in INT2 (in addition to strong expression in INT1 in Fig.1A). Compared to INT3-9, which looks red, INT2 cells appear yellow, suggesting that the expression patterns of the two TRAP drivers are not mutually exclusive, which changes the interpretation of many of the results described in the study.

      We agree that the Pges-1ΔB promoter is not absolutely restricted to INT1 and that weak GFP expression can also be detected in INT2. Because Pges-1ΔB is an engineered promoter derived from the intestine-specific Pges-11 promoter, this low-level INT2 expression is not unexpected. However, we note that the expression level in INT1 is substantially higher than in INT2. Thus, although the expression patterns of the two TRAP drivers are not completely mutually exclusive, Pges-1ΔB still provides the most selective available tool for enriching the INT1 translatome in the context of the current study.

      Some parts of the study overemphasize the differences between the INT1 vs INT2-9 cell types, which is a biased representation of the results. For example, the authors specifically point out that 270 genes are differentially expressed in opposite directions in INT1 vs INT2-9 cell types during acute (30 min) fasting without mentioning the 1,268 genes that are differentially expressed in the same direction. They also do not mention here that 96% of the genes are differentially expressed in the same direction in INT1 and INT2-9 cell types after prolonged (180 min) fasting, suggesting that the divergent translational responses of these cell types are only observed in the first 30 minutes of food deprivation. Similar results have also been reported for the effect of fasting on locomotory and feeding behaviors, where 30 min of fasting produces more variable effects, which become more consistent after longer periods of fasting (PMID: 36083280). Hence, the effects of brief food deprivation should be interpreted with caution.

      The intestine functions as a discrete and cohesive organ, so the expected result is that there would be no differences across the different cell types. For us, the surprise was that, in fact, there are differences between these cells at all. However, the point is well taken, and we have added a statement in the text to reflect that many genes change similarly in INT1 and INT2-9, while the differences reflect important functional divergence between these cell types.

      Many of the interpretations of this study primarily rely on pathway enrichment analyses, which are based on the known function of genes. The function of uncharacterized genes that were found to be differentially expressed in INT1 vs INT2-9 cell types, e.g., the ShKT proteins, was not explored in this study. In addition, overreliance on pathway enrichment tools (instead of functional validation) has resulted in several conflicting findings. For example, one of the main messages of this study is that INT1 cells specialize in immune and stress response in response to fasting, which relies on pathway analysis in Figs 5E and 5F. However, pathway analysis at a different time point (shown in Figure S5A) indicates that INT2-9 cells show a much stronger increase in translation of stress and pathogen-responsive genes compared to INT1 cells. Hence, some of the results should be interpreted as different translational effects in INT1 vs INT2-9 cells after different lengths of food deprivation, without making broad claims about selective pathways being affected only in specific cell types.

      We agree that some interpretations in the manuscript relied heavily on pathway enrichment analyses and should be stated more cautiously. In particular, we agree that the current data are most consistent with state-dependent differences in translational responses between INT1 and INT2-9 cells across different durations of food deprivation, rather than with the strongest version of a claim that specific pathways are selectively engaged only in one intestinal subset. We also agree that uncharacterized genes, including the ShKT family, were not mechanistically explored in the present study and should be presented as important candidates for future investigation.

      The authors have compared their TRAP-seq results with genes enriched in the anterior and posterior intestine clusters from a previously published whole-animal adult scRNA dataset (PMID: 37352352). They claim that their TRAP-seq results are in agreement with the findings of the scRNA study. However, among the 10 genes from the 'posterior intestine' scRNA cluster in Fig.S1E, six are downregulated in the INT1 vs INT2-9 comparison, while four are upregulated. Hence, there is no clear agreement between the two studies in terms of the top enriched genes in the anterior vs posterior intestine, which should be considered for cross-study comparisons in the future.

      We have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. We instead assessed the expression levels of our INT1 up-regulated genes in their intestinal cluster and found that they have higher expression in the anterior intestine cluster (new Figure S1C). These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9.

      The authors describe in the manuscript that they have performed INT1-specific RNAi for two C-type lectin genes that are upregulated during fasting. Due to a recent expansion of C-type lectin genes in C. elegans, there is a high chance of off-target effects of RNAi that is designed for members of this gene family. More trustworthy results could have been obtained using CRISPR-based loss-of-function alleles for these genes, one of which is publicly available. Also, the authors do not provide any explanation for why knockdown of these stress-response genes, which are activated in INT1 cells in response to food deprivation, results in improved resistance to pathogens. This, in fact, suggests a role of INT1 cells in increasing pathogen susceptibility, and not pathogen resistance, during food deprivation.

      We agree that RNAi targeting C-type lectin family members may be susceptible to off-target effects, and that validation with CRISPR null alleles, where available, would strengthen these findings. In the current study, we used INT1-specific RNAi as a cell-specific first-pass approach to test candidate gene function. We also agree that the pathogen phenotype requires cautious interpretation. Specifically, the finding that knockdown of fasting-induced INT1 lectin genes improves pathogen resistance does not support a simple protective model for these genes. Instead, it suggests that INT1-expressed stress-response genes modulate host susceptibility or host-pathogen interactions.

      Many of the studies in this field (e.g., references 2-4 in this article) have investigated the effects of food deprivation ranging from 4 hr to 24 hr, which results in activation of starvation responses in C. elegans. In contrast, the authors have used shorter time periods of fasting (30 min and 180 min), and most of their follow-up experiments have used 30 min of food deprivation. Previous work has shown that the effects of food deprivation can either accumulate over time (i.e., the effect gets stronger with longer food deprivation) or can be transient (i.e., only observed briefly after removal of food and not observed during long-term food deprivation). Starvation-induced transcription factors such as DAF-16/FoxO and HLH-30 show strong translocation to the nucleus only after 30 min of fasting. Though gene expression changes in all stages of food deprivation are of biological relevance, the authors have missed the opportunity to explore whether increased INS-7 secretion from the anterior intestine is dependent on these starvation-induced transcription factors (which can be easily tested using loss-of-function alleles) or is due to other fast-acting regulatory mechanisms induced due to the absence of food contents in the gut lumen. A previous study (PMID: 40991693) has shown that DAF-16 activation during prolonged starvation shuts down insulin peptide secretion from the intestinal epithelial cells. Hence, it is not clear if increased INS-7 secretion is only a feature of short-term food deprivation or is also a signature of long-term starvation (e.g., at 8 hr or 16 hr timepoints). Since most of the INS-7 secretion data in this study are for 30 min of fasting, it remains unknown whether the discovered regulators of INS-7 secretion can be generalized for extended food deprivation that triggers major metabolic changes, such as fat loss (e.g., conditions shown in Figure 1D).

      We agree that short-term food deprivation and prolonged starvation likely engage distinct regulatory mechanisms, and that our study primarily addresses an early phase of food deprivation rather than the full spectrum of starvation responses described in prior work. We selected the 30 min fasting condition because our previous study showed that INS-7 secretion is induced within this interval and returns to baseline upon refeeding, even before detectable intestinal fat loss. We also included a 180 min fasting condition to capture a later state associated with metabolic changes. However, we agree that the present study does not determine whether the regulators of INS-7 secretion identified here also govern secretion during more prolonged starvation (for example, 8 hr or 16 hr), nor does it test whether starvation-responsive transcription factors such as DAF-16 or HLH-30 contribute to this regulation. We appreciate that determining how this response transitions during prolonged starvation will be an important direction for future work.

      Two previous studies (PMID: 18025456, 40991693) have shown a strong reduction in the expression of ins-7 in the anterior intestine using GFP-based reporters (both promoter fusions and endogenous CRISPR-generated) and in whole-animal RNA-seq data from starved animals. These results are in contrast to the increased INS-7 secretion from INT1 cells during fasting that is reported in this study. The authors here have reported that INS-7 translation is higher in INT1 compared to INT2-9 during fed, acute fasted, and chronic fasted conditions, but they have not shown whether INS-7 translation is upregulated during acute and chronic fasting in INT1 cells in their TRAP-seq analysis. Knowing whether increased INS-7 secretion during acute fasting is due to increased transcription, translation, or secretion of INS-7 is crucial to resolve the discrepancy between these studies.

      In our dataset, INS-7 translation in INT1 tended to increase during acute fasting relative to the fed state (log<sub>2</sub>FC = 0.69), although this effect did not reach statistical significance after adjustment for multiple comparisons. Consistent with this trend, our secretion assay showed that INS-7 release from INT1 increases during fasting. However, we agree that the current data do not distinguish whether this increase in secretion is driven by enhanced synthesis, regulated release of pre-existing peptide stores, or a combination of both.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors set out to understand whether the discrete segments of the C.elegans intestine were specialized to carry out distinct functions during an animal's exposure and adaptation to a fast-changing nutrient environment. To achieve this, the authors used a method called Translating ribosome affinity purification (TRAP), which provides a snapshot of what genes are being translated into proteins (and therefore functionally prioritized by the animal) under different fasting and re-feeding conditions. By expressing the TRAP constructs in two distinct segments of the intestine (INT1) and (INT2-9), the authors were able to identify how these segments responded to changing nutrient availability.

      Already under steady state nutrient conditions, the authors found that INT1 and INT2-9 appeared to have different 'tasks', with INT1 expressing more immune- and stress-response related genes. Exposing animals to different regimens of starvation and refeeding also showed marked differences between the intestinal segments, and the gene expression patterns in INT1 were consistent with INT1 cells playing an integrative role in linking nutrient cues to the secretion of insulin molecules that regulate fat metabolism with food intake. In summary, the data presented catalogue, for the first time, gene expression differences between two areas of the intestine, suspected to play different roles, and through clever experiments, links these gene expression changes to responses to nutrient availability.

      Strengths:

      The data presented catalogue - for the first time and in a careful manner - gene expression differences between two areas of the intestine. They strongly support the presence of intriguing differences between two areas of the intestine in immune, metabolic, and stress-response regulation, and link these gene expression changes to the responses of these regions to nutrient availability.

      Weaknesses:

      The conclusions of this paper are mostly well-supported by data, but the relevance of the changing gene expression patterns could be better clarified and extended in the discussion.

      We thank the reviewer for this constructive comment. In the revised manuscript, we have now expanded the discussion to more clearly interpret these dynamic translatomic changes in the context of intestinal subset specialization. The most pronounced difference between INT1 and INT2-9 cells is the enrichment of stress-response genes in INT1. Based on the present findings, together with our previous work identifying INS-7 as an INT1-secreted signal (PMID: 39127676), we propose that INT1 cells are sentinel enteroendocrine cells that integrate information from the luminal environment and the metabolic state of intestinal cells.

      Reviewer #3 (Public review):

      Summary:

      In this study, Liu and colleagues utilize TRAP-seq to profile the repertoire of actively translated mRNAs in different intestinal cell types (anterior INT1 vs. posterior INT2-9 cells) in C. elegans. A key goal of this study was to identify transcripts differentially expressed/translated between these intestinal cell subtypes in the context of animals being well fed or subjected to acute (30 minutes) or chronic (3 hours) starvation, followed by refeeding.

      The authors identify a number of differentially expressed genes across all of the conditions tested. They then provide an initial survey of the landscape of translatome changes through Weighted Gene Network Correlation Analysis (WGNA), and some high-level functional surveys via Gene Ontology (GO) term analysis and protein domain analysis. The authors validate the enriched expression patterns of some of their identified candidate genes using fluorescent promoter fusion reporters, confirming INT1-specific expression. The authors further implicate the role of several other candidate genes in pathogen avoidance and in response to nutritional cues by knocking them down specifically in INT1 cells by RNAi. Finally, the authors identify pyruvate as a major nutrient signal coming from the bacterial diet that suppresses the release of a key insulin peptide (INS-7), and identify some of the genes expressed in INT1 that are required for this response.

      Strengths:

      (1) Good use of and justification for TRAP-seq, because scRNA-seq would be difficult under the varied conditions used (starvation, refeeding).

      (2) The manuscript is generally clear to read, and the data are generally well-presented with good supporting data that includes replicates, sample sizes, error measurements, and associated statistics.

      (3) The dataset will be an interesting resource to mine for future studies focusing on mechanisms of how particular intestinal cell types respond to different environmental signals.

      Weaknesses:

      (1) A limitation of TRAP-seq, although powerful, is that only relative comparisons can be made between genotypes/conditions to identify differentially-expressed genes, rather than assessing whether a given gene is expressed at a certain level in a cell type under a certain condition. This limitation is due to the non-specific association of sticky RNA species with the beads during the immunoprecipitation step. This is a minor point, however, and the authors do a nice job of focusing their analysis on differentially expressed transcripts in the current study.

      We agree that a limitation of TRAP-seq is that it is best suited for relative comparisons across cell types or conditions, rather than for determining the absolute expression level of a given transcript in a specific cell type. As the reviewer notes, this limitation arises in part from nonspecific recovery of background or sticky RNAs during the immunoprecipitation step, complicating the interpretation of absolute expression levels. For this reason, our analysis was designed to focus primarily on differentially enriched transcripts between INT1 and INT2-9 cells and across feeding states, rather than on assigning absolute expression levels to individual genes. We appreciate the reviewer’s recognition of this point. Our study uses TRAP-seq specifically to define relative translatomic differences between intestinal subsets and physiological states, which is well aligned with the strengths of this approach.

      (2) Another limitation of the current study is that the experiments testing the role of candidate genes identified by their profiling experiments do not delve a bit deeper into providing a mechanistic understanding of the phenotypes being studied. At present, the results are thus viewed more as a genomics-based screen with some limited follow-up on interesting hits. However, this reviewer appreciates that when placed in the context of the work presented, a presentation of the profiling data along with some validation is an excellent starting point for future mechanistic studies elaborating on these interesting candidates.

      We agree that the current study does not fully resolve the molecular mechanisms by which the candidate genes identified by TRAP-seq regulate the phenotypes examined here. Our primary goal was to generate a spatially resolved translatomic framework for INT1 and INT2-9 cells across feeding states, and to perform focused validation of selected candidates to establish the physiological relevance of the profiling results. We therefore view the current functional analyses as an initial validation and proof of principle, rather than a comprehensive mechanistic dissection of the molecular pathways for each candidate. We appreciate the reviewer’s recognition that these findings provide an excellent starting point for future studies.

      Appraisal of whether the authors achieved their aims, and whether the results support their conclusions:

      The main goal of the study was to survey the dynamic responses at the level of actively translated mRNAs of the INT1 vs INT2-9 cells in response to metabolic challenge.

      Overall, the authors use established methods to perform their genome-wide analysis, and the set of differentially regulated genes is enriched for expected molecular functions and forms coherent networks in anticipated pathways.

      The validation experiments (promoter::GFP fusion reporters, INT1-specific knockdowns of highly regulated genes) further corroborate the quality of the TRAP-seq datasets generated.

      I have a few points for the authors that would further strengthen this work:

      (1) The authors rightfully focus on the top differentially-regulated candidates, but it's unclear at present how far down their fold change list would lead to expression pattern validations. It would be useful to test a few more promoter::GFP fusion reporters at different enrichment/fold-change/statistical cutoffs.

      Testing additional promoter::mNeonGreen reporters across a wider range of fold-change and statistical thresholds could be somewhat useful for calibrating ranked TRAP-seq candidate genes. However, given the variation in strains bearing extrachromosomal arrays, we did not consider this a stringent enough test, given that the sensitivity and dynamic range of RNA-seq far outpaces genetic fluorescence-based reporters. For these reasons, we focused on the top differentially enriched candidates to provide not only an initial validation of the dataset, but also to determine whether these candidates regulate biological functions in INT1 cells and thus serve as potentially useful biological readouts in future efforts.

      (2) Although the INT1-specific RNAi provides a convenient strategy for rapidly perturbing and testing genes of interest for phenotypes, independently validating the knockdowns with genetic mutants, or alternatively (if genes are essential), degron alleles.

      We agree that validating the INT1-specific RNAi phenotypes with independent genetic approaches, including null-allele or degron-based alleles for essential genes, would further strengthen the conclusions. In the current study, we used INT1-specific RNAi as a rapid and spatially restricted strategy to functionally test candidates identified by TRAP-seq and to determine whether these genes contribute to the specialized physiological functions of INT1 cells. We consider these experiments an initial validation of candidate function rather than a complete genetic dissection, which could be conducted in future efforts to study other aspects of INT1 function.

      Impact:

      The TRAP-seq data and list of differentially-expressed candidate genes will form an interesting set of high-priority candidates to study for their role in the reception and transduction of nutritional cues in response to food status and pathogens. This data will thus benefit the C. elegans community of researchers studying the mechanisms governing these phenomena.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major comments:

      (1) The authors need to describe the fasting method used in detail. Was fasting performed on unseeded NGM plates or in liquid (M9 buffer)? Were the animals washed with buffer prior to starvation? If yes, how many times? These details are critical for any researcher to follow up on their results.

      We have clarified the fasting/refeeding procedure in the revised Methods section. Briefly, worms were washed off OP50-seeded NGM plates with M9 buffer, washed three times in M9 buffer, and then transferred to unseeded NGM plates for fasting. For refeeding, worms were collected from the unseeded NGM plates with M9 buffer and transferred back to OP50-seeded NGM plates.

      (2) The authors claim that "INT1 and INT2-9 cells maintain fundamentally different molecular identities independent of any and all acute or chronic conditions", which they primarily based on Principal Component Analysis (PCA). The circles shown in Figure 2A are arbitrary, and many such circles can be drawn in the 2D space to separate the samples in different ways. The authors should show this comparison in a translatome-wide similarity heatmap with hierarchical clustering (similar to Fig.2C, but with all the experimental conditions and their replicates on both x- and y-axes).

      The ellipses shown in Figure 2A were generated using the stat_ellipse() function in ggplot2, which calculates the mean and covariance of the PC1 and PC2 for each line and draws ellipses corresponding to the 95% confidence level. The separation between lines is primarily driven by PC2, which accounts for 14% of the variance in the translatomic dataset. Although this difference is less pronounced when considering the full translatome, samples from the same line nevertheless cluster together, supporting line-specific differences in translatomic profile.

      (3) Figure 4 of the study shows 18 Venn diagrams for genes that are differentially expressed between INT1 and INT2-9 cell types in fed, fasted, and refed conditions. In the absence of any statistical comparisons, it is difficult to interpret whether the extents of overlap (higher or lower than expected) are significant. Ideally, P values for hypergeometric tests should be provided for the overlap regions.

      We appreciate the reviewer’s suggestion. We explored the use of hypergeometric testing, implemented through the SuperExactTest package in R, to assess the statistical significance of the overlaps shown in the Venn diagrams. However, this analysis yielded significant P values for essentially all overlap regions, including cases in which the degree of overlap was not especially informative and did not align with the interpretation presented in the text. This outcome likely reflects the dependence of the test on the size of the input gene sets and background universe, which can make statistical significance difficult to interpret meaningfully in this context.

      (4) The claims made in lines 242-244 (Figure 5D) need to be supported by P values from hypergeometric tests.

      Similar to the previous point.

      (5) The interpretation of Figures 6E and 6F described in lines 296-297 needs to be supported by statistical analyses. The authors claim that the undulating pattern of expression of the turquoise module genes is stronger in INT1 compared to INT2-9. However, based on Figures 2C and 6F, it appears that the expression change is not necessarily weaker in INT2-9, but instead is different, i.e., the expression of turquoise module genes goes up during fasting in INT1 and goes down after refeeding, while their expression goes up during fasting and stays up after refeeding in INT2-9 cells.

      We appreciate the reviewer’s point and agree that the turquoise module shows dynamic regulation in both cell populations. The key difference is not the presence versus absence of an undulating pattern, but rather the magnitude of that change, which is greater in INT1. Because of the limited number of biological replicates in some conditions, particularly the fasting group, we interpreted these results cautiously and used a nonparametric approach to assess differences in average module expression between states. This analysis indicated that the turquoise module changes significantly in both lines, but with a larger effect size in INT1. We have included the corresponding statistical analysis and effect size in the revised manuscript.

      (6) Since the INS-7 coelomocyte uptake assay was used extensively in this study, some representative microscopy images should be included to complement the quantification.

      We have added a new Figure 6G showing representative images corresponding to the quantification presented in Figure 6H.

      (7) The authors claim that INT1-specific fmo-2 RNAi results in reduced basal INS-7 secretion, but they do not have the direct statistical comparison for this. Were experiments shown in Figures 6G and 6K done on the same day?

      In the original Figure 6K (now Figure 6L), the data are presented as the percentage of normalized INS-7::mCherry fluorescence intensity relative to fed animals treated with vector RNAi. A statistical comparison between fed animals treated with INT1-specific fmo-2 RNAi and fed vector RNAi controls was performed and was significant. We have also clarified that the experiments shown in the original Figures 6G and 6I (now Figures 6H and 6J) were performed on the same day.

      (8) It is not clear why blocking the mitochondrial breakdown of pyruvate (Figures 7E and 7F) does not mimic the fasted state in terms of increased INS-7 secretion from INT1 cells. Doesn't this contradict the proposed model in this study? Can the authors speculate why this is the case?

      We do not interpret inhibition of pyruvate dehydrogenase or pyruvate carboxylase as equivalent to the fasted state. Rather, our model is that fasting induces INS-7 secretion by lowering intracellular pyruvate in INT1 cells. Under this framework, blocking mitochondrial pyruvate breakdown would be expected to reduce pyruvate utilization and thus maintain intracellular pyruvate, preventing the drop in pyruvate that normally occurs during fasting. This would explain why these manipulations suppress fasting-induced INS-7 secretion. To directly examine this possibility, we performed the experiment in Figure 7G, which tests whether maintaining pyruvate levels in INT1 cells during fasting is sufficient to suppress INS-7 secretion. The results are consistent with this interpretation and further support a model in which decreased intracellular pyruvate is a key determinant of fasting-induced INS-7 secretion.

      Minor comments:

      (1) Figures 2E and 2G are very similar and represent the same result in two different ways (unbiased vs guided comparison). One of these should be moved to the supplementary figures.

      Although these figures show similar patterns, they were derived from two independent analytical approaches, WGCNA and differential expression analysis. We therefore interpret the concordance between these independent methods as strengthening the robustness of the association and increasing confidence in the biological relevance of the observed pattern.

      (2) In Figures S1C, S1D, and S1E, a more significant P-value is shown with a smaller circle, and a less significant P-value is shown with a larger circle. This is confusing to the reader and should be inverted.

      We appreciate the reviewer’s comment and have removed the original Figure S1C–E, replacing it with a more informative analysis. The genes used in the original panel were drawn from the top markers reported for intestinal clusters in Ghaddar et al. (PMID: 37352352). However, these markers were defined by comparison with all C. elegans cell types, rather than by comparisons among anterior, middle, and posterior intestinal populations, and are therefore not optimal for resolving differences between intestinal subregions. Our further examination of marker expression across the intestinal clusters in the Ghaddar et al. (PMID: 37352352). dataset confirmed this limitation. These results underscore the strength of our dataset for identifying genes that distinguish INT1 from INT2–9. We also note that the spatial identities of the intestinal clusters in Ghaddar et al. (PMID: 37352352) were not clearly established in the text or by spatial transcriptomic evidence, making it difficult to assign the annotated anterior, middle, and posterior clusters to specific intestinal cells. We have revised the manuscript accordingly and replaced the original figure panels.

      (3) Line 131 mentions the comprehensive characterization of the translatomic differences between INT1 and INT2-9 cells under each acute and chronic condition. However, the paragraph only discusses the differences in the fed condition. This is confusing, and the authors should mention the comparison between these cell types under acute and chronic conditions in subsequent sections where it is described.

      In this paragraph, we indeed discuss the ‘fed’ condition, but in subsequent sections we follow with details analyses of acute versus chronic, as well as regional differences across the intestine. We have clarified this in the opening sentence of the referenced paragraph.

      (4) The Venn diagrams in Figure 4 look very similar, and it is hard to differentiate between how 4A is different from 4I, how 4B is different from 4J, etc. The authors should include the labels for 'acute' or 'chronic' above each Venn diagram to guide the reader through these panels.

      We have added labels indicating the acute and chronic conditions to the left side of each Venn diagram in the revised Figure 4.

      (5) It is not clear in the figure legends how Figure 5E is different from Figure S6A, and how Figure 5F is different from Figure S7A. This should be better described in the figure legends.

      We have added a sentence to better describe this in the figure legend.

      (6) Figure 6C: Survival parameters such as median lifespan, number of animals for each condition, etc., should be reported for the different conditions.

      We have revised Figure 6C to indicate the number of animals analyzed in each condition, and the median survival for each group is now reported in the corresponding figure legend.

      (7) The colors used for control RNAi and clec-160 RNAi are very similar in Fig.6C. Easily distinguishable colors should be used.

      We have changed the colors as suggested.

      (8) The INT1-specific RNAi strain should be first described in line 285.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (9) Line 304: 'REF' should be replaced with the reference.

      We have replaced the “REF” with the reference (PMID: 39127676)

      (10) The P value for statistical comparison between the fed and 30 min refed states should be shown in Figures 6G, 6I, and 6K.

      We have now included the p value for the comparison as suggested. Figures 6G, 6I, and 6K are now labeled as 6H, 6J, and 6L, respectively.

      (11) In Figure 7, the authors should consider replacing the 'refed' label with 'recovery' because the pyruvate treatment was done in the absence of 'feeding' (= bacteria consumption).

      We appreciate the reviewer’s point. However, we chose to retain the label “refed” in Figure 7 to maintain consistency across the set of conditions examined, including 2% glucose and OP50 supernatant, which likewise do not involve bacterial consumption despite not showing effect on the refeeding response of INS-7 secretion.

      (12) The full form of DISN should be mentioned in the figure legend of Figure 7.

      We have included the full form of D1SN in the figure legend of Figure 7A.

      (13) Line 367: 'normalization' should be replaced with 'return to basal levels'. 'Normalization of INS-7 secretion' might also mean normalization of INS-7::mCherry signal to CLM::GFP signal.

      We have revised the wording per the reviewer's suggestion.

      (14) The methods section has a quantitative RT-PCR section, but it is not clear if RT-PCR data are reported in any of the figures. Also, no qPCR primers are listed in Table S3.

      We have removed the quantitative RT-PCR part from the methods section.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors describe that the RPL-22-3xHA constructs are not integrated, at the very end, in the section "Limitations of the data". An earlier mention of this caveat would have been useful. In addition, it would help if the authors could provide their defense (which I think is very valid) of using non-integrated strains in the results section, as they describe the experimental setup. Also, some details were missing, which left me wanting to know: Were there expression differences? How were they accounted for? Was expression normalized between these two constructs, and if so, how?

      The RPL-22-3xHA construct was integrated into each line to ensure more consistent transgene expression across experiments. Because the INT2-9 construct is expressed in a larger number of cells than the INT1 construct, the INT2–9 samples yielded greater amounts of pulled-down nascent RNA, as reflected in the supplemental table and in the higher aligned RNA counts observed for the INT2-9 samples. To account for these differences, differential expression analysis was performed using DESeq2, which corrects for library size by estimating sample-specific size factors with the median-of-ratios method. Raw counts are then normalized using these size factors, thereby accounting for differences in sequencing depth and minimizing confounding effects due to variation in library size. Such differences are common in RNA-seq experiments, particularly when comparing samples derived from distinct input populations.

      (2) The 'acute' and 'chronic' exposures are thought through and carefully defined. The question I do have is whether the 3-hour fasting can be considered chronic fasting, given how surprisingly fast the animals lose their fat content. Could these kinetics indicate that the 30-minute fasting is reflective of mechanisms during which senses change in food availability, whereas the 30 minutes represents acute fasting (with chronic fasting - meaning fasting, during which the animal activated alternative pathways - occurring later)? While this may appear to be pure semantics, it could influence how the authors interpret their results. One method to more objectively separate an 'acute' from a 'chronic' stage may be to conduct a time course of fat loss-does fat loss plateau after 3 hours? The timing when the rate of decrease levels off could be more indicative of the beginning of a chronic phase.

      We appreciate this important point and agree that it should be more clearly discussed. We interpret acute fasting as a pre-fat-loss state, since it is 30 minutes off food and no difference in fat levels are detectable at this stage (Fig 1B). The translatomic changes observed under acute fasting therefore likely reflect food-sensing mechanisms and early preparatory responses that promote subsequent fat mobilization. In contrast, chronic fasting (180 minutes off food – see Fig 1D) appears to represent a post-fat-loss state, in which fat stores have already been depleted, and the corresponding translatomic changes likely reflect the effects of sustained metabolic stress.

      (3) The age of the animals used has to be more explicitly stated. Were these animals egg-laying? Or L4/young adults? This is likely to impact the changes that the animals undergo.

      Day 1 young adults were subjected to the fasting. Great care was taken to ensure consistency across biological replicates.

      (4) What is the rationale, in the authors' view, that stress response genes are apparently more enriched than metabolic or mitochondrial enzymes, and membrane receptor changes? Are the latter mostly regulated by PTMs/localization changes, etc?

      Based on our current data, we cannot exclude the possibility that metabolic or mitochondrial enzymes, as well as membrane receptors, are regulated in INT1 and INT2–9 cells through mechanisms not captured at the translatome level, including post-translational modification or changes in subcellular localization under different fasting and refeeding conditions.

      (5) The refeeding experiment with latex beads and killed OP50 is very clever. Details on when INS-7 was evaluated in the caoelomocytes would help the reader better understand and interpret these results.

      INS-7mCherry signal was evaluated in the coelomocytes immediately after refeeding; we included this information in the methods section and referenced our previous paper.

      (6) In the Discussion, I was looking for a more detailed context for how to think about the differences and similarities in the RNA-seq data between the two segments, and perhaps a discussion of whether there were any indications that the two segments communicated with each other.

      The data show that the most pronounced difference between INT1 and the rest of the intestine at the RNAseq level, is the expression of stress response genes in INT1. Although there are some nuanced differences, the prevalence of stress response genes persists across feeding and fasting conditions. This difference, combined with the evidence that INT1 cells secrete the enteroendocrine peptide INS-7 (Fig 6 and PMID: 39127676) is strongly reminiscent of the mammalian enteroendocrine cells, which also secrete peptides and show strong expression of stress response genes (PMID: 37626258 and 27148273). We suggest that this category term reflects not only a canonical stress response, but also a broader response to shifts in the luminal environment, which INT1 cells are anatomically poised to detect well before the absorption of nutrients has begun further down the intestine (INT2-9). Thus, we believe INT1 cells are a newly defined enteroendocrine cell type within the C. elegans intestine.

      Regarding communication between INT1 and INT2-9 – this is an intriguing possibility that we have considered, given that peptide genes and receptors are found in the RNAseq datasets. The extent to which the expression of these genes leads to functional effects is the subject of future investigation.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1A - It would be better to also show single fluorescent protein channels to assess the specificity of the expression patterns. A schematic or labels of where the INT1 vs. INT2-9 boundaries are located would be helpful to non-experts.

      (2) Figure 4 - At present, the Venn Diagrams are a very complicated way to visualize all of the comparisons/conditions. I would recommend that the authors consider using UpSet plots to better summarize the relevant comparisons they would like to make. The same consideration applies to Figure 5D.

      (3) Line 301 - Description of the INT1-specific RNAi strategy. I think it would be better to bring this information earlier, close to line 285, where the authors first mention performing INT1-specific RNAi experiments.

      We have added the description for the INT1-specific RNAi strain in line 285.

      (4) Line 304 - I think the authors meant to cite a reference where the REF placeholder text is found

      We have replaced the “REF” with the reference.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study investigated mitochondrial dysfunction and the impairment of the ciliary Sonic Hedgehog signaling in Lowe syndrome (LS), a timely topic given the limited research in this area. The data from patient iPSC-derived neurons and a mouse model were collected using solid methods, but the evidence supporting key claims is incomplete, and some technical aspects fall short of expectations. Despite these limitations, the study provides a useful foundation for exploring the relationship between mitochondrial defects and primary cilia in neural development.We appreciate the editorial assessment highlighting the importance of studying mitochondrial dysfunction and ciliary signaling in Lowe syndrome. We acknowledge that our study is largely associative, and we have revised the manuscript to clearly state this limitation, toned down causal claims, and emphasized that our work provides a foundation for future mechanistic studies.

      We appreciate the editorial assessment highlighting the importance of studying mitochondrial dysfunction and ciliary signaling in Lowe syndrome. We acknowledge that our study is largely associative, and we have revised the manuscript to clearly state this limitation, toned down causal claims, and emphasized that our work provides a foundation for future mechanistic studies.

      We have also:

      - Improved figure clarity and consistency

      - Corrected errors in gene annotations and normalization

      - Refined the mechanistic framework linking OCRL, mitochondria, and cilia

      New Experimental Data:

      Figure 5, Supplementary Figure 4. We confirmed mitochondrial defects by generating ocrl-KO zebrafish (Supplementary Figure 4). We first assessed mitochondrial reactive oxygen species (mitoROS) using MitoSOX staining. Next, we evaluated mitochondrial membrane potential (ΔΨm) using MitoTracker CMXRos. Finally, we assessed mitochondrial content via TOM20 staining. For all analyses, we focused on the ocular and cranial regions of the zebrafish to maintain consistency (see Author response image 1). Notably, previous studies have reported that ocrl-KO zebrafish exhibit seizures and brain developmental abnormalities, supporting their relevance as a model for Lowe syndrome-like phenotypes [1].

      Author response image 1.

      Figure 6 c. In addition to quantifying the proportion of ciliated cells in the IOB mouse brain, we measured cilia length and compared it between IOB and WT brain sections. Our results show that IOB mice exhibit elongated cilia compared to WT controls, suggesting that OCRL deficiency is associated with stress-related alterations in ciliary structure. These findings are consistent with previous studies reporting that cilia elongation can be associated with increased ROS levels and mitochondrial dysfunction [2,3].

      Public Reviews:

      Reviewer #1 (Public review):

      The preparation of the manuscript requires improvement. There are many errors in the presentation of data.

      We thank the reviewer for this important comment. We have carefully revised the manuscript to improve the clarity, accuracy, and consistency of data presentation.

      Specifically, we have corrected inconsistencies in gene nomenclature (e.g., CO2 vs COX2, DLOOP) across the text, figures, and legends. We standardized normalization methods and ensured consistency between figures and descriptions. We revised figure labels, legends, and annotations for clarity and accuracy. We corrected referencing errors and ensured appropriate citation of prior work. We improved overall figure quality and readability. In addition, we performed a thorough review of the entire manuscript to eliminate typographical errors and ensure consistency in terminology and data interpretation.

      The use of references needs to be re-considered. Sometimes a reference is used when in fact the results included in that paper are the opposite of what the authors intend.

      We thank the reviewer for this important comment. We have carefully re-evaluated all references throughout the manuscript to ensure that they accurately reflect the findings they are cited to support. In cases where the cited studies did not fully align with our interpretation or could be misleading, we have either revised the text to more accurately represent the original findings or replaced the references with more appropriate sources. We have also clarified instances where prior studies report differing or context-dependent results to avoid overinterpretation.

      The authors conclude the paper by claiming that mitochondrial dysfunction and impairments of the ciliary SHH contribute to abnormal neuronal differentiation in LS, but the mechanism by which this sequence of events might happen hasn't been shown.

      We thank the reviewer for this important comment. We agree that the current study does not establish a direct causal mechanism linking mitochondrial dysfunction, ciliary SHH signaling, and altered neuronal differentiation in Lowe syndrome. Our data demonstrate that these processes co-occur consistently across multiple model systems, supporting a potential functional relationship. However, we acknowledge that the precise sequence of events and mechanistic connections remains to be defined. To address this, we have revised the manuscript to clarify that our conclusions are based on associative findings rather than direct mechanistic evidence. We have also updated the Discussion to explicitly acknowledge this limitation and to frame our model (Figure 7) as a proposed working hypothesis. Future studies will be required to determine whether mitochondrial dysfunction directly impacts ciliary SHH signaling and how these pathways influence neuronal differentiation.

      Phenotype of increased astrocytes in both the IOB mouse brain or iPSC-derived cultures iN cells requires clarification as one of the markers used as an astrocyte marker, BRN2, is commonly used as a neuronal marker. As LS is a neurodevelopmental disorder, and the phenotype in question is related to differentiation, it is crucial to shed light on the developmental timeline in which this phenotype is seen in the mouse brain.

      We thank the reviewer for this important comment. We agree that the use of BRN2 as an astrocytic marker was inappropriate, as it is primarily recognized as a neuronal marker. Accordingly, we have revised the manuscript to remove BRN2 from the interpretation of astrocytic identity and now rely on GFAP expression as the primary astrocytic marker. We have also clarified this point in both the Results and figure legends to avoid misinterpretation. In addition, we have revised the text to more accurately describe our findings as an altered balance in neuronal versus astrocytic marker expression, rather than a definitive increase in astrocyte numbers.

      Regarding the developmental context, we acknowledge that Lowe syndrome is a neurodevelopmental disorder and that temporal aspects are highly relevant. In our study, the in vivo analyses were performed on adult 2-month-old IOB mouse brains, which we have now explicitly stated in the manuscript. We recognize that this limits our ability to directly assess developmental dynamics of lineage specification. We have therefore added this as a limitation in the Discussion and clarified that future studies examining earlier developmental stages will be necessary to determine when these alterations arise.

      Mitochondrial dysfunction in astrocytes has been shown to induce a ciliogenic program. However, almost the opposite is shown in this paper, with regards to ciliation. Morphology of the cilia was not assessed either, which is an important feature of ciliary homeostasis. The improper ciliary homeostasis here appears to be the improper Shh signalling, which has not been shown to be related to mitochondrial dysfunction. This leaves one wondering how exactly the different phenotypes shown in this paper are connected.

      We thank the reviewer for this important comment. We agree that the relationship between mitochondrial dysfunction, ciliogenesis, and Shh signaling is complex and not fully resolved in the current study.

      As noted by the reviewer, prior studies have reported that mitochondrial dysfunction can promote a ciliogenic program [4]. In contrast, our data show a reduced proportion of ciliated cells together with increased cilia length, indicating altered ciliary homeostasis rather than a straightforward increase in ciliogenesis. To address this point, we have revised the manuscript to describe our findings as context-dependent alterations in ciliary parameters more clearly, and we now explicitly discuss this apparent discrepancy with the literature in the Discussion. We also acknowledge the reviewer’s point regarding ciliary morphology. In the revised manuscript, we have included quantification of cilia length in addition to the proportion of ciliated cells, and we have expanded the Methods section to detail how these measurements were performed. We agree that additional ultrastructural and functional analyses would further strengthen the characterization of ciliary homeostasis, and we now include this as a limitation and future direction.

      Regarding the link between mitochondrial dysfunction, ciliary alterations, and Shh signaling, we agree that our study does not establish a direct mechanistic connection. Our data demonstrate that these phenotypes co-occur consistently across multiple models, but do not define causality. To address this concern, we have revised the manuscript to clarify that our conclusions are associative, and we now present our integrated model (Figure 7) as a working hypothesis rather than a demonstrated mechanism. We also explicitly state in the Discussion that future studies will be required to determine whether mitochondrial dysfunction directly impacts ciliary signaling and Shh pathway activity.

      This paper lacks a clear mechanistic approach. While the data validates the 3 broad phenotypes mentioned, there is a lack of connection between these phenotypes or an answer to why these phenotypes appear. While the discussion attempts to shed light on this by referencing previous studies, some of the referenced studies show contradicting results. Hence, it would be beneficial to clarify these gaps with further experiments and address the larger question of the connection between the mitochondria, Shh signalling, and astrocyte formation.

      We thank the reviewer for this important and insightful comment. We agree that the current study does not establish a direct mechanistic link connecting mitochondrial dysfunction, altered Shh signaling, and changes in neuronal versus astrocytic differentiation.

      Our primary goal in this work was to identify and validate phenotypes associated with OCRL deficiency across multiple independent model systems. We demonstrate that mitochondrial dysfunction, oxidative stress, altered ciliary/Shh signaling, and changes in neural lineage-associated markers co-occur consistently in these models. However, we acknowledge that the causal relationships between these processes remain to be defined.

      To address this concern, we have revised the manuscript to more clearly state that our conclusions are associative rather than mechanistic, and we now present our integrated model (Figure 7) as a working hypothesis that links these phenotypes through a potential mitochondria-ROS-signaling axis. We have also expanded the Discussion to explicitly acknowledge this limitation and to avoid overinterpretation of causality.

      In addition, we have carefully re-evaluated and revised the cited literature to ensure accuracy, particularly in cases where prior studies report context-dependent or seemingly contradictory effects of mitochondrial dysfunction on ciliogenesis and signaling pathways. These points are now discussed more explicitly to better position our findings within the existing literature.

      We agree that further experiments, such as targeted rescue of mitochondrial function or modulation of Shh signaling, will be necessary to establish causal relationships between these pathways. These directions are now clearly outlined in the revised Discussion as important next steps.

      Most importantly, there is no mention of how the loss of OCRL, a 5-phosphatase enzyme, results in the appearance of the mentioned phenotypes. Since there are multiple studies in the field of Lowe Syndrome that shed light on the various functions of OCRL, both catalytic and non-catalytic, it is important to address the role of OCRL in resulting in these phenotypes.

      We thank the reviewer for this important comment. We agree that the link between OCRL function and the observed phenotypes was not sufficiently developed in the original version of the manuscript.

      In the revised manuscript, we have expanded the Discussion to more clearly outline how loss of OCRL could contribute to the observed mitochondrial, ciliary, and differentiation phenotypes. OCRL encodes a PI(4,5)P₂ 5-phosphatase that regulates phosphoinositide homeostasis and membrane dynamics. Disruption of this activity is known to affect endolysosomal trafficking, actin organization, and membrane remodeling-processes that are critical for organelle maintenance and ciliary function. We now discuss how these alterations could impact mitochondrial homeostasis, for example, through defects in membrane contact sites, vesicular trafficking, or organelle quality control pathways.

      In addition, we have incorporated discussion of potential non-catalytic roles of OCRL, including protein–protein interactions and scaffolding functions, which may contribute to the coordination of intracellular trafficking and cytoskeletal organization. These aspects may provide an additional layer of regulation linking OCRL loss to both mitochondrial dysfunction and ciliary alterations.

      We emphasize that, while these mechanisms are supported by prior studies, our data do not directly test them. Therefore, we have carefully framed this section as a plausible mechanistic framework rather than a demonstrated pathway and have explicitly stated this limitation. We also outline future experiments aimed at dissecting catalytic versus non-catalytic contributions of OCRL to these phenotypes.

      There are numerous errors in the qPCR experiments performed concerning the genes that were assayed. The genes mentioned in the text section do not match those indicated in the graphs or legends. This takes away the confidence of the reader in this data.

      We thank the reviewer for this important observation. We agree that the inconsistencies between the genes described in the text and those shown in the figures and legends could reduce confidence in the data. In the revised manuscript, we have carefully rechecked all qPCR experiments and corrected the gene names across the Results, figures, and figure legends to ensure full consistency. We have also standardized the nomenclature throughout the manuscript (including consistent use of gene symbols and formatting) and verified that all plotted data correspond to the correct targets.

      In addition, we have clarified the qPCR methodology, including normalization (all data are normalized to GAPDH) and primer information, to improve transparency and reproducibility.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates how neural cell development is affected in Lowe syndrome. Using neural cultures differentiated from human iPSCs carrying either an LS mutation or a genetically engineered mutation in OCRL, the authors show a depletion of mitochondrial DNA and a decrease in mitochondrial activities that correlate with an increased formation of astrocytes at the expense of neurons. Similar effects on mitochondria and on astrocyte development were observed in an LS mouse model. Moreover, these mutant brain cells are less likely to be ciliated and show a reduction in Sonic Hedgehog signalling.

      Strengths/Weaknesses:

      The study derives strength from the analyses of two different models of Lowe syndrome, both reaching similar conclusions. However, the observed changes in mitochondrial defects, neuronal/astrocytic development, and primary cilia are only correlated, with no attempt to investigate a causal relationship. Moreover, the mouse model is only analysed at the adult stage providing no insights into the development of the defects. Different brain regions are analysed with immunostainings and qRT-PCR making it challenging to draw clear correlations between these findings. The quality of the corresponding figures is often poor and the selection of markers is frequently inappropriate. Taken together, these limitations complicate the interpretations of the data and significantly limit the conclusions that can be drawn from the study.

      We have carefully revised the manuscript to address the concerns raised, and we have revised the manuscript with additional supporting data.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The authors have checked the expression of neuronal markers NeuN and FoxG1, and apart from GFAP, they categorise Brn2 as one of the astrocytes markers that they have also checked. But Brn2 is not an astrocyte marker. It is a neuronal marker that is expressed in layer 2/3 of the cortex. In fact, Brn2 is reported to be a key driver of neurogenesis in primate telencephalon development1 and for reprogramming of astrocytes to neurons2. Hence, the only glial marker they have used here is GFAP. BRN2 is a neuronal marker. It has been used as a neuronal marker even in the reference (Zhang et al, 2013), from which the protocol for inducing iPSCs to induced neurons (iNs) was taken. Hence, the qPCR results in 1g of overexpression of BRN2 indicate an increase in expression of a neuronal marker, not an astrocyte marker.

      We thank the reviewer for this important and well-founded comment. We fully agree that BRN2 is a neuronal marker and not an astrocytic marker, and that its inclusion as an astrocyte marker in our original interpretation was incorrect.

      In the revised manuscript, we have removed BRN2 from the analysis and interpretation of astrocytic identity. We now treat BRN2 exclusively as a neuronal marker and have updated the Results, figure legends, and text accordingly. Specifically, the qPCR data previously presented in Figure 1g are now interpreted as reflecting neuronal marker expression, not astrocytic differentiation. We have also revised our conclusions to avoid overinterpretation of astrocyte abundance. Our findings are now described more accurately as an altered balance in neuronal versus astrocytic marker expression, rather than a definitive increase in astrocyte numbers. In this context, GFAP remains the primary astrocytic marker used in this study.

      We acknowledge the reviewer’s point that reliance on a single astrocytic marker is a limitation. This has now been explicitly stated in the Discussion, and we note that additional astrocyte markers will be required in future studies to more comprehensively define lineage-specific changes.

      Incorrect marker usage (BRN2 as astrocyte marker)

      We thank the reviewer for identifying this critical issue. We corrected the classification of BRN2 as a neuronal marker. Also, we re-analyzed the interpretation accordingly, revised all relevant text and figures. Importantly, Astrocyte conclusions are now based primarily on GFAP expression, and we explicitly acknowledge this limitation in the Discussion

      The graphs for the RT-PCR results indicate that gene expression values are normalized to actin whereas the legend mentions that they are normalized to GAPDH. This needs clarification.

      We thank the reviewer for pointing out this inconsistency. We confirm that all qPCR data were normalized to GAPDH, and the reference to actin was an error. This has now been corrected throughout the figures, legends, and text to ensure consistency.

      The use of wording to refer to the generation of induced neurons (iNs) should ideally be changed from "we developed"; as the protocol from Zhang et al, 2013 seems to have been directly adapted in this paper.

      We thank the reviewer for this helpful suggestion. We agree that the wording was inappropriate. In the revised manuscript, we have replaced “we developed” with language indicating that iNs were generated using an established protocol, and we now explicitly state that the method was adapted from Zhang et al., 2013 [5].

      OCRL KO iPSCs were obtained from Herbert Lachman's lab and not generated in this study. Hence, the Ran et al, 2013 reference is not necessary.

      We thank the reviewer for this clarification. We agree that the OCRL knockout iPSCs were obtained from Herbert Lachman’s laboratory and were not generated in this study. Accordingly, we have removed the Ran et al., 2013 reference and revised the manuscript to clearly state the origin of the OCRL KO iPSC line.

      The reference for Figure 1c is given as Ran et al, 2013 which is wrong. It should be Zhang et al, 2013.

      We thank the reviewer for noting this error. We have corrected the reference for Figure 1c from Ran et al., 2013 to Zhang et al., 2013 in the revised manuscript.Limited in vivo mitochondrial characterization

      Figure 1D, E: GFAP is a cytoskeletal marker but its expression here is very grainy and looks like an artifact. Is it possible to show the astrocyte phenotype using other astrocytes nuclei and cytosolic markers such as NFIA and S100B, respectively?

      We thank the reviewer for this important suggestion. We acknowledge that GFAP is a cytoskeletal marker and that the signal in the current images may appear granular. We have carefully re-evaluated the staining and image processing to ensure that the signal represents true GFAP expression and have improved the image quality and presentation in the revised figures. We agree that inclusion of additional astrocytic markers such as NFIA and S100B would further strengthen the characterization. While we were not able to include these additional markers in the current revision, we now explicitly acknowledge this as a limitation in the Discussion and note that future studies will incorporate a broader panel of astrocyte markers to more comprehensively define astrocytic identity.

      Since the authors have not used enough markers to understand the cell-type composition in WT and OCRL KO/mutant lines, it's not sufficient to conclude that the NPCs preferentially favour astrocytes over neuronal lineage. Any conclusive comments regarding the cell-state/cell-type specification necessitate evidence such as genetic lineage tracing using reporters for neuronal and astrocyte markers, and/or RNA/ATAC/scRNA sequencing.

      We thank the reviewer for this important point. We agree that the current marker panel is not sufficient to definitively determine cell-type composition or to conclude preferential lineage specification.

      In the revised manuscript, we have tempered our conclusions and now describe our findings as changes in neuronal versus astrocytic marker expression, rather than evidence of a shift in lineage fate. We also explicitly acknowledge this limitation in the Discussion. We agree that approaches such as genetic lineage tracing, reporter-based assays, and single-cell transcriptomic or epigenomic analyses (e.g., scRNA-seq or scATAC-seq) would be required to rigorously define cell-state transitions and lineage outcomes. These are important directions for future studies and are now highlighted in the revised manuscript.

      Figure 2:

      The mt-DNA gene CO2 was checked, not COX2. A typographical error in the written section, which does not match with the qPCR graph of the same.

      We thank the reviewer for noting this inconsistency. We confirm that the gene analyzed was CO2, and the reference to COX2 in the text was incorrect. This has now been corrected throughout the manuscript to ensure consistency between the text, figures, and qPCR data.

      The word "neurogenesis" is used very loosely throughout the paper. In the opinion of the reviewer, there is no evidence presented that there is a defect in neurogenesis in either of the models used in this paper.

      We thank the reviewer for this important comment. We agree that the term “neurogenesis” was used too broadly and is not directly supported by our data. In the revised manuscript, we have removed or replaced this term where appropriate and now refer more precisely to changes in neuronal versus astrocytic marker expression.

      Line 150: They say that they have examined the functional properties of mitochondria during neurogenesis but it would have been better to understand OXPHOS at various time points of neurogenesis to actually conclude reduced OXPHOS 'during neurogenesis'. Moreover, genes related to other pathways such as glycolysis could have been checked to understand the bioenergetics of LS patients. Also, oxidative stress could have been checked using more than one marker. Since they are trying to understand the functional role of mitochondrial defects during neurogenesis, they could have performed live imaging of mitochondrial potential during various stages of neurogenesis. Isolation of mitochondria from LS patients and transcriptomics/proteomics might provide further clues about mitochondrial defects.

      We thank the reviewer for these thoughtful suggestions. We agree that our data do not capture mitochondrial function across multiple stages of neurogenesis. In the revised manuscript, we have modified the wording to avoid implying temporal analysis “during neurogenesis” and instead describe mitochondrial parameters in differentiated cells. To strengthen the study, we have included additional in vivo validation in the zebrafish model, where we assessed multiple mitochondrial readouts, including mitochondrial membrane potential (ΔΨm) using MitoTracker CMXRos, oxidative mitochondrial stress ROS (mitoROS) using MitoSOX staining, and mitochondrial content (TOM20), supporting mitochondrial dysfunction across systems.

      Elevated astrocytic reaction during the differentiation of NSPCs in the Lowe syndrome (IOB) mouse model.

      Title: What does astrocyte reaction mean? This term should not be used without clear evidence of reactive astrocytes being present in the model.

      We thank the reviewer for this important comment. We agree that the term “astrocytic reaction” is not appropriate without specific evidence of reactive astrocytes. In the revised manuscript, we have removed this terminology and replaced it with more accurate wording, describing our findings as altered astrocytic marker expression. This change better reflects the data and avoids overinterpretation.

      In 1a, no quantification of the mouse brain size is given. From the given images alone, there appears to be no obvious decrease in brain size between the WT and IOB mice. This contradicts the text which indicates that the IOB mouse brain is smaller.

      We appreciate this important point. We have:

      Removed claims regarding reduced brain size

      Clarified that our analysis was limited to available sections and no definitive conclusion about global brain morphology can be made

      Figure 3E: Why is the astrocyte to neuron ratio measured using a cytoskeletal marker for astrocytes, GFAP but a nuclear marker for neurons, NeuN? Ratios to measure the percentage or proportion of astrocytes to neurons can only be checked by markers of the same nature such as GFAP to MAP2 (neuronal cytoskeletal marker) or NFIA (astrocyte nuclear marker) to NeuN.

      We thank the reviewer for this important point. We agree that comparing a cytoskeletal marker (GFAP) with a nuclear marker (NeuN) is not ideal for deriving cell-type ratios. In the revised manuscript, we have removed the astrocyte-to-neuron ratio analysis and now present these data as relative marker expression/signals rather than proportions. We have also clarified this limitation in the text and Discussion.

      (3) Lack of clarity in the experiments performed on the mice brains. PAX6 is used here as a neuronal marker along with NeuN, a mature neuronal marker. This is misleading as PAX6 is rather a marker for neural stem/progenitor cells and not neurons. The age of the mice has also not been mentioned, which is crucial considering the different markers used to characterize the mouse brain as well as since the authors are indicating that there is an abnormal neurodevelopment in the IOB mouse during development. Again, BRN2 is used here as an astrocyte marker. However, it is a neuronal marker. Hence, the phenotype of increased astrocytes currently is held by GFAP expression alone. Another astrocyte marker should be used.

      We thank the reviewer for these important points. We have revised the manuscript to correct marker interpretation, now describing PAX6 as a progenitor marker rather than neuronal, and BRN2 as a neuronal marker, removing it from astrocyte-related analysis. We have also explicitly stated the age of the mice (2 months) in the Methods and Results. In addition, we have tempered our conclusions, describing the data as changes in marker expression rather than definitive cell-type shifts, and we now acknowledge that reliance on GFAP as a single astrocytic marker is a limitation, which is discussed in the revised manuscript.

      Figure 6: Increase in astrocytes, mitochondrial dysfunction, and ciliary Shh signalling are 3 phenotypes discussed in this study. However, no experiments were done to shed light on the mechanistic connection between these phenotypes. This is reflected in the abstract shown in Figure 6. There is no comment on the mechanism behind these phenotypes.

      We thank the reviewer for this important comment. We agree that the current study does not establish a direct mechanistic link between mitochondrial dysfunction, altered ciliary Shh signaling, and changes in astrocytic markers. Our aim was to identify and validate these phenotypes across multiple models. In the revised manuscript, we have clarified that Figure 7 represents a proposed working model based on associative findings rather than a defined mechanism. We have also revised the Discussion to explicitly acknowledge this limitation and to outline future experiments required to establish causal relationships between these

      Reviewer #2 (Recommendations for the authors):

      The authors report interesting findings in two different experimental models but the manuscript would benefit significantly from an analysis of a potential causal relationship between different findings. They often mention neural stem cells or the neuron/glia switch but their analysis of the mouse mutant is restricted to the adult stage. A more consistent analysis of specific brain regions would also be beneficial.

      We thank the reviewer for this constructive comment. We agree that establishing causal relationships between the observed phenotypes is an important next step. In the revised manuscript, we have clarified that our conclusions are based on associative findings and have expanded the Discussion to outline experimental strategies that could address causality in future studies. We also acknowledge that our in vivo analysis is restricted to adult (2-month-old) IOB mouse brains, which limits our ability to assess developmental dynamics such as neural stem cell behavior or neuron-glial transitions. This limitation is now explicitly stated in the Discussion. Finally, we agree that region-specific analysis would strengthen the study. Due to the availability of samples, our analysis was not systematically performed across defined brain regions. We now acknowledge this limitation and note that future studies focusing on specific regions (e.g., cortex, hippocampus) will be important to better understand the spatial aspects of the phenotype.

      Figure 1: The authors only measured the expression of marker genes by qRT-PCR. This could reflect higher expression levels in individual cells rather than a change in the proportion of neurons and astrocytes. They need to determine the cell proportions of astrocytes and neurons in addition. Moreover, there is a poor marker choice. Loss of FOXG1 expression could indicate a loss of telencephalic identity. BRN2 is expressed by cortical neurons.

      We thank the reviewer for this important comment. We agree that qPCR-based marker analysis does not directly reflect cell-type proportions and may instead represent changes in gene expression per cell. Accordingly, we have revised the manuscript to avoid conclusions about cell proportions and now describe the data as changes in marker expression. We have also corrected marker interpretation, removing BRN2 from astrocyte analysis and clarifying that FOXG1 reflects telencephalic identity. These limitations and the need for more comprehensive cell-type characterization are now acknowledged in the Discussion.

      In Figure 2, the authors determine the properties of mitochondria and claim that functional mitochondrial activities are decreased during neurogenesis in mutant iN cells. They need to take into account that according to Figure 1 the proportion of neurons and astrocytes may be changed. Hence, the decreased mitochondrial activity may reflect a fundamental difference between neurons and astrocytes. The authors need to clearly distinguish between neurons and astrocytes in their analysis. Moreover, the use of the term neurogenesis is confusing. They are analysing the neuron-to-glial switch, not the formation of neurons.

      We thank the reviewer for this important comment. We agree that differences in cell-type composition may influence mitochondrial measurements. In the revised manuscript, we have tempered our interpretation, describing these data as changes in mitochondrial parameters at the population level rather than neuron-specific effects. We also acknowledge this limitation in the Discussion and note that cell-type-specific analyses will be required in future studies. In addition, we have revised the terminology throughout the manuscript, removing the term “neurogenesis” and instead referring to changes in neuronal versus glial marker expression to more accurately reflect the scope of our analysis.

      Figure 3: The authors claim that astrocyte numbers are elevated in the IOB mouse model, however, it seems as if the authors analysed late postnatal, potentially adult brains but no age of the brains is provided. Given the large time lag between the formation of astrocytes and their analysis, the increased number of astrocytes could be due to a number of processes including altered proliferation and cell death. The authors need to investigate the proportion of astrocytes and neurons closer to the neuron-to-glia switch. Cell fate experiments like the long-term application of BrdU would be much better suited and would provide mechanistic insights. Again, markers are not adequate to reach their conclusion. Pax6 is only expressed in a tiny subset of neurons, Brn2 on the other hand is not astrocyte-specific as it is expressed in cortical neurons as well. Moreover, qRT-PCR analyses were done in the cortex and hippocampus whereas the boxes in Figure 3D are located in the basal ganglia. It would be much more informative and provide better comparisons to perform gene expression analysis and cell counts in the same brain regions.

      We thank the reviewer for these important and constructive comments. We agree that our analysis is limited by the use of adult (2-month-old) IOB mouse brains, which do not allow direct assessment of developmental processes such as the neuron-to-glia transition. We have now explicitly stated the age of the animals and clarified this limitation in the Discussion, including the possibility that changes in astrocytic markers may reflect processes such as proliferation or survival rather than lineage specification.

      We also agree that our marker panel was insufficient for definitive conclusions. Accordingly, we have revised the manuscript to remove overinterpretation, corrected marker usage, and now describe the data as changes in marker expression rather than cell-type proportions. The need for more rigorous approaches, such as lineage tracing (e.g., BrdU) and expanded marker panels, is now acknowledged as a future direction. Finally, we thank the reviewer for pointing out the inconsistency in the brain regions analyzed. We have clarified the regions used for qPCR, and we now explicitly acknowledge this limitation, noting that future studies will aim to perform region-matched molecular and histological analyses for more accurate comparisons.

      Experiments in Figure 4 assess "whether changes in mitochondrial activity are involved in the altered differentiation of stem cells and progenitor cells in the LS mouse model" in 3-month-old brain sections. The murine adult brain only contains a few neural stem cells in the SVZ and in the dentate gyrus. Instead, this analysis needed to be done at late embryonic/early postnatal stages to capture the neuronal/glial switch. In addition, RT-PCR and immunostainings should be performed in the same brain region as stated above.

      We thank the reviewer for this important point. We agree that analysis in adult (2-month-old) brains does not capture developmental stages such as the neuron-glia transition. We have revised the manuscript to remove implications of developmental analysis and now describe these data as mitochondrial parameters in adult tissue, explicitly acknowledging this limitation in the Discussion. We also clarify the brain regions used for qPCR and immunostaining and note as a limitation that these were not fully matched; future studies will perform region-specific, developmentally timed analyses.

      Figure 5: The authors examine a potential link between mitochondrial defects and primary cilia. Mutant iN cell cultures contain lower levels of SHH mRNA and show concomitantly lower expression of the SHH target genes GLI1 and PTCH1. The authors link this finding with a reduced proportion of ciliated cells but the reduced SHH signalling is most likely explained by the decreased SHH expression. The authors also limit their analysis of primary cilia to one brain region, but they should also include the cortex and hippocampus as these regions were used for their qRT-PCR analysis. SHH signalling acts as a switch to stop the proteolytic processing of GLI3 and to promote the formation of the GLI3 activator form. It is therefore important to determine the ratio of GLI3 repressor and GLI3 activator using western blots. The authors claim that they found defective cilia formation, but cilia are poorly characterised. Are there differences in intraflagellar transport, the formation of the transition zone, etc? Is ciliary length altered? The authors only make a correlative link between mitochondrial defects and cilia but present no experiments to investigate causation. They should at least discuss potential mechanisms which could explain defects in cilia.

      We thank the reviewer for these insightful comments. We agree that reduced SHH pathway activity may be influenced by decreased SHH expression, and we have revised the text to avoid overattributing this effect to ciliary changes. Our conclusions are now framed as associative, not causal.

      We have expanded our cilia analysis to include quantification of both the proportion of ciliated cells and cilia length, and clarified these methods in the manuscript. We also acknowledge that additional characterization (e.g., intraflagellar transport, transition zone structure, GLI3 activator/repressor ratios) would further strengthen the analysis, and we now include this as a limitation and future direction. Regarding regional analysis, we agree that broader brain region coverage would be valuable. Due to sample availability, our analysis was limited, and this is now explicitly acknowledged as a limitation, with future studies aimed at region-matched analyses (e.g., cortex and hippocampus). Finally, we have expanded the Discussion to outline potential mechanisms linking mitochondrial dysfunction and ciliary alterations, while clearly stating that causal relationships remain to be established.

      (1) The methods section does not contain any information on how immunostainings on brain sections were performed.

      We agree that the description of immunostaining on brain sections was missing. We have now added a detailed protocol for brain section immunostaining in the Methods section to improve clarity and reproducibility.

      (2) The abbreviation "RT-PCR" is used for both, real-time PCR and reverse transcription PCR

      We also acknowledge the inconsistent use of the term “RT-PCR.” In the revised manuscript, we have standardized the terminology, using “qPCR” (quantitative real-time PCR) throughout to avoid confusion

      Conclusion

      We believe that these revisions significantly strengthen the manuscript. While the study remains primarily associative, it provides a multi-model, cross-species framework linking mitochondrial dysfunction, ciliary signaling, and altered neural differentiation in Lowe syndrome.

      References:

      (1) Ramirez IB-R, Pietka G, Jones DR, Divecha N, Alia A, Baraban SC, et al. Impaired neural development in a zebrafish model for Lowe syndrome. Hum Mol Genet. 2012;21:1744–59. https://doi.org/10.1093/hmg/ddr608

      (2) Kim JI, Kim J, Jang H-S, Noh MR, Lipschutz JH, Park KM. Reduction of oxidative stress during recovery accelerates normalization of primary cilia length that is altered after ischemic injury in murine kidneys. Am J Physiol Renal Physiol. 2013;304:F1283-1294. https://doi.org/10.1152/ajprenal.00427.2012

      (3) Moruzzi N, Valladolid-Acebes I, Kannabiran SA, Bulgaro S, Burtscher I, Leibiger B, et al. Mitochondrial impairment and intracellular reactive oxygen species alter primary cilia morphology. Life Sci Alliance. 2022;5:e202201505. https://doi.org/10.26508/lsa.202201505

      (4) Ignatenko O, Malinen S, Rybas S, Vihinen H, Nikkanen J, Kononov A, et al. Mitochondrial dysfunction compromises ciliary homeostasis in astrocytes. J Cell Biol. 2022;222:e202203019. https://doi.org/10.1083/jcb.202203019

      (5) Zhang Y, Pak C, Han Y, Ahlenius H, Zhang Z, Chanda S, et al. Rapid Single-Step Induction of Functional Neurons from Human Pluripotent Stem Cells. Neuron. 2013;78:785–98. https://doi.org/10.1016/j.neuron.2013.05.029

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      The authors of this study developed a method to quantify calvarial bone marrow from MRI head scans, enabling the study of its composition in large datasets of adults, usually collected to study the brain. Bone marrow intensity can be semi-quantitatively measured in T1-weighted MRI scans due to the greater signal intensity of fat than watery red marrow. This is an ingenious use of the MRI-produced information for other important phenotypes, such as bone structure and marrow content. Different head types were tested for complying with the model, which is notable.

      The model was also successfully validated using several publicly available MRI resources - real data - in (1) a dataset consisting of 30 individuals that were scanned 10 times each at 3-day intervals, and (2) the monozygotic (MZ) twin data from the Human Connectome Project cohort. Then the authors applied this validated method to head-MRI scans from the UK Biobank (n=33,042) to extract information on the spatial distribution of bone marrow adiposity (BMA) in the calvaria, allowing a GWAS to identify associated genes.

      The authors revealed high heritability and identified 41 genetic loci significantly associated with the BMA trait, including six sex-specific loci. Of note, statistics estimate that 99% of BMA trait-influencing variants are shared with BMD (497 of 500 variants), which may mean these results demonstrate the biological relevance to bone health. Some of the BMA genes were found related to the Wnt pathway, including WNT16, WNT4, NXN; this is a "positive control", since the Wnt/β-catenin signaling pathway was suggested as an important determinant of BMA. Also, associations in genes (BMP4, DLX5, LGR4, LRP4, SFRP4) that are known to specifically influence adiposity, are encouraging. Integrating mapped genes with bone marrow single-cell RNA-seq data revealed patterns of adipogenic lineage differentiation and lipid loading.

      With regards to the reviewer’s comment on the overlap between BMA and BMD trait-influencing variants, we would like to add that the correlation of effect sizes within the overlap is -0.95 as we would expect from bifurcating differentiation of mesenchymal stem cells into osteoblasts or adipocytes: the underlying common biology driving these traits results in shared traits with a negative correlation in the effects.

      The study also investigated the genetic overlap between BMA and twelve (or 13) "brain and body" traits and identified significant genetic correlations with BMI, cognitive ability, and Parkinson's disease.

      In sum, since MRI head scans present a hitherto unexplored opportunity to address unresolved aspects of bone marrow biology, this study is both timely and innovative.

      There are, however, some assumptions, findings, and their interpretation, which require more critical focus.

      Sex-specificity is well described and studied here. Men have higher BMA than women, but post-menopausal women catch up in the BMA values. The authors believe that calvarial marrow has a number of features that make it particularly well-suited to the study of BMA process - which is clinically important in other bone sites. It has a simple "sandwiched" structure that they are able to model. This is true only to some extent: a condition called "Hyperostosis frontalis interna", of unknown etiology (described by Smith & Hemphill in 1956) - is characterized by irregular overgrowth of the inner table of the frontal bone (symmetric/bilateral). Although not of clinical significance, typically benign, studies report a prevalence of 12%; However, it's most common in postmenopausal women - where prevalences up to 49% in women over the age of 65 - have been reported. Thus, sexual dimorphism is obvious and the effect of estrogen is likely shared with whichever bone - and marrow - age-related pathology. So, for women not using HRT, this new layer of the bone might interfere with the calvarial BMA readings and in turn, affect the BMA-related analyses.

      Thank you for bringing the "Hyperostosis frontalis interna" condition to our attention. It is particularly interesting to hear that the etiology is unknown and one may suspect that some kind of calvarial bone marrow dysregulation may be part of the cause. Our model for bone marrow location was trained on simulated data which included variation in the thickness of all anatomical layers (including the inner table), so it will be robust to some thickening of the inner table. It might not be robust to the most extreme cases of inner table thickening (as described in some case reports), but these are rare. Further, it should also be noted that the other calvarial bones, representing a much greater fraction of the calvarial surface, remain largely unaffected by the thickening and would therefore yield correct localisation of the bone marrow layers. In summary, although the severe cases of hyperostosis frontalis interna have the potential to affect our identification of the bone marrow layer, the low frequency of such cases and the restriction of the phenotype to the frontal bone means that the potential for bias is very limited.

      It would be interesting to develop a method for detection of thickened inner bone so that the condition’s prevalence can be quantified in a large sample like the UK Biobank and its genetic architecture be determined. This could help elucidate the etiology.

      The authors suspect that the effect of BMA on BMD may be biased in women; they should comment on those "with low BMD and high BMA" given that hyperostosis frontalis might be an issue. A strong effect of SNPs in the ESR1 chromosomal region might be akin to the above concern.

      Thank you for raising this point, which we have followed up with a new analysis.

      According to ICD-10 data in UK Biobank there are only N=105 individuals with an M85.2-diagnosed disorder. Given the total sample size of N=446,814 individuals with ICD-10 data, this would translate to a prevalence of 0.02%, which speaks for an underdiagnosis in this sample such that we cannot simply remove diagnosed individuals to control for a potential diagnostic confound.

      We have therefore taken a different approach to investigate this potential issue: As you elaborated in your previous comment, the prevalence of hyperostosis frontalis increases with age in females. The literature also suggests that prevalence rates do not differ between males and females in young age / prior to menopause. Therefore, we have studied the association between BMD and BMA for males and females separately, and in two age groups based on a median split of our sample (left plot: younger than 65, right plot: subjects older than 65). In these plots, the relatively large shift in female BMA and BMD is visible with the large yellow cloud at low BMA and high BMD in the left plot disappearing in the right plot. Despite this, we observe:

      (1) Associations in both males (blue) and females (yellow), suggesting that the associations were not driven only by females.

      (2) BMA-BMD association is largely similar across the two age groups.

      If we consider that the old age group is likely to contain more cases of hyperostosis frontalis than the young group, and if we consider that old-aged females are more likely to be in this condition than men of any age, then we would expect an impact of hyperostosis frontalis on our measures to result in observable differences in Author response image 1. This is not the case. We see global age-related shifts in BMA in women, yet the association with BMD remains similar across age groups.

      The technical properties of our neural network (trained on simulated data) makes it unlikely that frontal bone will contaminate the bone marrow detection globally (description above) and these results show that hyperostosis frontalis is not a considerable issue in our analysis.

      Author response image 1.

      Then, there is a perfect overlap of the BMA SNPs that are shared with BMD (497 of 500 variants), which may prove a "face validity" of the MRI-derived BMA. However, the BMD in the study was heel-derived eBMD - which is a good proxy for osteoporosis and is mostly driven by trabecular bone. Thus, there might be a concern that the BMA metrics capture some trabecular BMD.

      The reviewer is correct in pointing out that the BMA causal variants are a near-perfect subset of the BMD causal variants. The reviewer raises the concern that the BMA measurements may capture some trabecular BMD, however it should be noted that the correlation of effect sizes for the BMA/BMD overlapping causal SNPs is negative (-0.95). If our measure of BMA had been erroneously capturing trabecular BMD then we would expect to see a positive correlation of effect sizes for the BMA/BMD overlapping causal SNPs, not a negative one.

      Next, integrating mapped genes with existing bone marrow single-cell RNA-sequencing data revealed patterns of adipogenic lineage differentiation and lipid loading. The problem here is that the scRNAseq studies of the Bone Marrow niche are overwhelmingly mouse. The authors might wish to justify why they are relevant to humans (in the absence of the human-specific scRNAseq).

      We thank the reviewer for pointing this out. We noticed that, although Figure 4 and the Methods do explicitly state that the scRNAseq data is from mouse, it is not stated in the text of the Results. This is now corrected.

      The mouse is commonly used as the model organism for in vivo investigation of human phenotypes and bone marrow adiposity is no exception because, although mice have lower bone marrow adiposity than humans, the timing and sequence in bone marrow adiposity development are similar. BMA research makes extensive use of mouse models literature as exemplified by this review of research within the field (https://www.frontiersin.org/journals/endocrinology/articles/10.3389/fendo.2016.00127/full) and this article recent article (Koh et al. 2024. “Adult skull bone marrow is an expanding and resilient haematopoietic reservoir”. https://www.nature.com/articles/s41586-024-08163-9)

      We updated the results section (line 279):

      “Mesenchymal stem cells of the BM niche commit to either the adipogenic or the osteogenic lineage (Figure 4A) and both the number committing to the adipogenic lineage and their level of lipid-loading influences the total level of BMA. This aspect of BM biology is shared between humans and mice (29), so we made use of an existing mouse scRNAseq dataset of BM mesenchymal lineage cells (30) to study variation in the expression of BMA-associated genes as cells differentiate (Figure 4B).”

      For genetic correlation analysis, the authors selected 7 body and 6 brain traits. The latter traits reflect cognition (general cognitive ability and educational attainment) and brain-related disorders. This selection might seem arbitrary. The interpretation of genetic correlation with cognitive ability, education, and Parkinson's disease was attributed to the recently discovered vascular channels that link calvarial bone marrow to the meninges. This is a fascinating hypothesis, which requires functional proof. However, there might be simpler explanations. Thus, the diploe and the inner table of the calvarium are drained by the same veins as the dura. From the anatomy textbook, we know that diploic veins connect the pericranial and endocranial venous system through the skull.

      Whilst it is true that we did not systematically compare the results of the BMA GWAS to all potentially relevant brain and body phenotypes, we did use criteria to select the phenotypes we compared to. As stated in the manuscript (line 304):

      “We selected body traits (BMD, BMI, waist-to-hip ratio, systolic and diastolic blood pressure, type-2 diabetes, coronary artery disease) that have a logical connection to BMA given the mesenchymal stem cells origin of BM adipocytes and their role in bone, fat, and vasculature (29). For the brain, we selected traits reflecting cognition (general cognitive ability and educational attainment) and disorders that are prevalent in adulthood (insomnia, multiple sclerosis, Parkinson's disease, Alzheimer's disease) since it is primarily in adulthood that the adiposity of calvarial BM experiences a substantial change”

      We entirely agree that the suggestion that the genetic correlation between BMA and cerebral traits may be mediated by the vascular channels linking calvarial bone marrow to the meninges is merely a hypothesis. We have therefore updated the text of the Discussion (line 470):

      “We tentatively speculate that calvarial MALPs may be involved in sensing perivascular flows of CSF from the meninges to the BM and in influencing the BM’s hematopoietic response, and that this might be the basis of the observed genetic overlap between BMA and some cerebral traits. However, more conventional anatomical pathways may also be relevant, as the diploë and inner table communicate with meningeal and dural venous systems through diploic veins.”

      Reviewer #2 (Public review):

      Summary:

      This study develops a new artificial intelligence method for high-throughput analysis of skull bone marrow from MRI data, which may be useful for large-scale biological analyses. Using this method, the authors then attempt to estimate skull bone marrow adiposity (BMA) using T1-weighted signal intensity from MRI scans of ~33,000 people, followed by genome-wide association analysis; however, the approach is inadequate because T1-weighted signal intensity is not validated for measurement of bone marrow adiposity. If it could be validated, the study would be an important advance in understanding of bone marrow adiposity and skeletal biology.

      Strengths:

      This paper is well-written, and the figures are nicely presented. The neural network method used for analysing skull bone marrow is innovative, and the authors validate this through several approaches. Therefore, the authors have achieved the aim of developing a method for large-scale analysis of skull bone marrow from MRI data.

      The GWAS is reasonably well-powered and addresses potential ethnicity differences, with one GWAS done across white males and females, and a separate GWAS in non-white participants. The methodology also conforms to common GWAS standards, including for mapping genetic variants to candidate genes. Moreover, the study further investigates the biological roles of these genes by analysing their expression in single-cell RNA sequencing data.

      Weaknesses:

      The fundamental weakness is that T1-weighted MRI signal intensity (T1W) is used as an estimate of BMA, but it has never been validated for this. The authors show that this T1W parameter measures something that is heritable and can be compared between subjects, but they don't show that it actually measures (or even estimates) calvarial BMA. There is an attempt to do so by comparing the T1W parameter with data from quantitative T1 images: the authors show a reasonable correlation with some of the quantitative T1 image data. However, this still does not show that the parameter is measuring BMA; it could be measuring some other biological characteristic, but this remains unclear. So, there is a need to validate the T1W parameter against an established measure of BMA, such as the bone marrow fat-fraction or proton density fat fraction measured from multi-echo MRI analysis.

      Without validating this BMA measurement method, it is not possible to interpret the GWAS or other findings reported in the study.

      We reject this criticism.

      Although T1-weighted has not been validated as a quantitative measure of fat-fraction, there are several studies showing that it is a semi-quantitative measure of fat content (e.g. Loevner et al 2002, Shen et al 2013, Zhang et al 2020) and we also provide data that support this (figures S9-11).

      Semi-quantitative measures are used in many biomedical GWASes for instance even highly heritable neuropsychiatric disorders (such as schizophrenia and bipolar disorder) involve assessment by clinicians where the test-retest kappas are in the range 0.4-0.6.

      Further, we would suggest that the shortcoming of the imperfect correlation of T1w signal intensity with fat content is more than outweighed by our precision in identifying the calvarial BM cavity and the fact that the flat calvarial bone marrow has a wide range of adiposity in middle-aged and elderly individuals (compared to other bones). This lies at the root of why:

      We clearly recapitulate the known sex and age profiles, as well as the effect of HRT.

      We estimate high BMA heritabilities (43% in males and 23% in females)

      We find clear sex differences (which is a known feature of BMA biology)

      We identify a large number of the genes already known to affect BMA from earlier animal and cell work

      A noisy measurement of an entity with strong biological signal (a well-defined bone marrow cavity with variation in BMA across subjects) will often be more informative than a highly precise measurement of a poorly defined entity with little signal.

      A less critical weakness is that the GWAS has been done only on a single cohort, without replicating the findings in a follow-up cohort. For example, the authors could repeat their analysis on the remaining ~50,000 UK Biobank imaging participants for whom MRI data is now available. However, this would be pointless without knowing what biological characteristic(s) the T1W parameter is actually reflecting.

      We disagree with this comment. We separated the UKB data into discovery and replication sets prior to running the GWAS, so these datasets are independent:

      (1) Further, we ran the discovery (white british individuals) GWAS separately for males and females (prior to combining) and reported in the results section: “We found them to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      (2) We performed our replication GWAS in non-white British males and females. As noted in the results section: “Out of the 168 significant discovery SNPs, 62% replicated at P<.05, and 39% of the 41 lead SNPs replicated at P<.05 (Table S6). One locus replicated at genome-wide significance (P<5e-8). Furthermore, 92.7% of the lead SNPs of the discovery sample showed same effect direction in the replication sample (Table S5).”

      Reviewer #3 (Public review):

      Summary:

      This manuscript, "Estimating bone marrow adiposity from head MRI and identifying its genetic 2 architecture", brings together the groups of Drs. Kaufmann and Hughes in a tour de force work to develop an artificial neural network that localizes calvaria bone marrow in T1-weighted MRI head scans, with the goal of studying its composition in several large MRI datasets, and to model sex-dimorphic age trajectories, including the effect of menopause.

      Strengths:

      Bone marrow adiposity is a very active tissue with far-reaching implications for tissue crosstalk and human health than we had initially recognized. Although MRI has been used to measure BM, studies such as the one by these two groups are still lacking whereas very large datasets are analyzed using advanced AI machine learning tools coupled with genetic studies and a specific pathology. The groups had to develop new methods and new AI machine-learning tools for the imaging analyses.

      Weaknesses:

      Some aspects of the work that authors could add additional clarification.

      (1) Imaging Limitations: The authors provide an excellent overview and references supporting the use of MRI as a method for assessing marrow fat, particularly with some specific modifications. However, MRI images can be affected by various factors, including the presence of other tissues as well as specific MRI settings, which are much harder to precisely control when using different datasets.

      We thank the reviewer for his positive assessment of our review of methods.

      Regarding MRI settings: We agree with the reviewer that differences in scan protocols can create substantial differences in the resulting images between samples. Different tools exist for harmonization of imaging data across sites, but they usually operate on tabulated data and there is no one-size-fits-all approach yet [1]. Here, we took a different approach to prevent confounding bias: We generated a large set of simulated data for training of the neural network. The simulations circumvented potential issues emerging from confound biases in training sets that we might have seen had we had combined multiple samples with different scan protocols. Nevertheless, applied to real data the models may still face confound issues, such as better BMA estimates for some scan protocols over others. We have addressed these issues as follows: (1) Validation analyses (10 repeat scans of 30 individuals and twin pairs, figure 1d and 1e) are based fully on data that was acquired on the same scanner with the same protocol. (2) Analysis in UK Biobank included data from different scan sites albeit harmonized protocols. Here we accounted for scan site in all statistical models (including GWAS).

      (1) Dominik Kraft, Gloria Matte Bon, Édith Breton, Philipp Seidel, Tobias Kaufmann; Removing scanner effects with a multivariate latent approach: A RELIEF for the ABCD imaging data?. Imaging Neuroscience 2024; 2 1–7. doi: https://doi.org/10.1162/imag_a_00157

      Regarding the presence of other tissues: We recognise in the existing text of the results section that sometimes inner or outer table voxels are wrongly identified as bone marrow, but we show that this does not have a major impact on the correct identification of the bone marrow cavity. The existing text reads:

      “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap. However, this did not result in a corresponding fall in the ratio of the predicted intensity of BM to its true intensity, because the typical BM intensity was only marginally higher than the neighbouring bone intensity. This property of the typical relative intensities of these anatomic structures also explains why the intensity ratio at high overlap is not centred on 1: any misidentification of cortical bone as BM, will typically result in an underestimate of true BM intensity (Figure S2). The neural network performed well and intensity ratios were in the range 0.9-1.1 for the vast majority of head types (Figure S3).”

      Also note that we implement a number of QC measures to exclude scans where there is evidence that we may have failed to correctly identify the bone marrow cavity. The existing text reads:

      “We used two additional QC metrics to filter out calvaria where BM location was likely to have failed. First, we set an upper limit of 30 on the standard deviation of the intensity of the outer table as scans with higher values were clear outliers and were probably cases where the location of both outer table and BM has failed (Figure S6). Second, for each calvarium, we computed the Mahalonobis distance for all vertices in the two dimensions “first layer of the BM” and “BM intensity” (Figure S7). By manual inspection we found that data points with MD > 25 often had errors in BM layer identification, typically where the network had erroneously predicted a higher and more intense layer to be the BM. We considered a calvarium as failing this QC criterium if more than 0.5% of vertices have MD > 25. This criterium is very strict as errors on only 0.5% of data points in a calvarium would not significantly affect the average BM intensity for a calvarium.”

      (2) The specific density of cranial bones as it relates to the types of bone marrow: Cranial bones are extremely dense structures, which naturally interfere with MRI imaging. While it is thought that cranial bones have mostly "red bone marrow", this is only true for a short time in humans. How sensitive is their system in differentiating between red and yellow BM?

      We implemented several measures to ensure that our method would be robust to anatomical variation between individuals. As noted in the current version of the Methods section: “In order to train the neural network model, we generated a large synthetic dataset of intensity arrays, with known boundaries between anatomical structures, by simulating the thickness and intensity of the different structures located between the outer skin and the subarachnoid space. The simulation incorporated the following real-world complexities:

      Different anatomical architectures (skin, subcutaneous fat, aponeurosis, outer table, BM, inner table, dura mater, arachnoid space), including when a structure is not present throughout the calvarium

      Variation in thickness and intensity between vertices (on the same calvarium)

      A wide variety of different calvarium types with different combinations of levels of BM adiposity, bone thickness, and subcutaneous adiposity.”

      Further, as noted in the Results section:

      “We evaluated the performance of the neural network on simulated data using two metrics (Figure 1B): 1. the overlap between the predicted and the true BM location, and 2. the ratio between the predicted intensity of the BM and the true intensity of the BM”. The accuracy in localising the bone marrow layers was good, with the only exception being: “Poor overlap (below 0.7) was almost only observed in the thinnest bone and is explained by the fact that when the BM part of the bone is only a few layers thick (1 layer = 0.5 mm), an error by one layer will inevitably lead to a substantial fall in overlap”

      We also validated our procedure on real data (see Results section, subsection “Procedure validation on real data and heritability estimate”). Briefly, we checked the accuracy of our method using a dataset from the Consortium for Reliability and Reproducibility, a twin dataset from the Human Connectome Project and by manually checking many hundreds of UKBiobank scans.

      We are thus confident that we accurately identify the bone marrow cavity irrespective of whether the bone marrow is red (low adiposity) or yellow (high adiposity).

      (3) Both items above are further complicated by aging, but aging is not a linear event as we have learned. There are specific bursts of aging in humans around the age of 45 and early 60s. How do the system and model predict or incorporate these peaks of aging? It seems from the data shown that aging is reflected more as a linear phenomenon. Is this because additional aging datasets are needed?

      We agree with the reviewer that ageing probably occurs in bursts rather than being a linear process. We do see a non-linear relationship between age and BMA in our data (see figure 2B), with a more rapid rise in BMA between the ages of 45 and 65, than later in life (in women). As a result of this, when we model BMA using regression, we use orthogonal polynomials of degree 2 which allows for a non-linear relationship. However, we cannot observe bursts of BMA increase in our data because it is cross-sectional. Longitudinal data would be required to obtain information on the nature and timing of any bursts in bone marrow adiposity.

      (4) The authors describe in richness of detail their AI learning programming and how it extracted the data from datasets. The authors also show some important correlations with specific genes, SNPs. What is not clear is how conditions such as anemia for example. An expected finding would be that patients with chronic anemia have lower bone marrow (BM) signal intensity on MRI scans than healthy people. This is because the signal intensity of BM depends on the fat-to-cell ratio in the tissue.

      We agree with the reviewer that conditions affecting the bone marrow niche have a potential to affect and be affected by bone marrow adiposity, with leukemia being a known example. This is why we believe that a method, such as the one we present here, has the potential to be useful in several biomedical fields (hematology and osteology).

      Furthermore, patients with a host of musculoskeletal disorders ranging from osteopenia to osteoporosis, sarcopenia, and osteosarcopenia will also have altered MRI scans. When using such large datasets how did the authors control or exclude these pathological conditions, or were all these conditions likely present?

      We did not exclude specific pathologies. We were careful to train our NN model on a wide variety of skull thicknesses, bone marrow adiposity, and subcutaneous adiposity and to evaluate the performance of the model on simulated and real datasets (see answer to your point 2).

      Reviewer 1 raised the issue of individuals displaying Hyperostosis frontalis interna (thickening of the inner table) and we recognize that in extreme case of this condition, where there is a major change in the anatomy of the calvarial bone, our method would probably not correctly localise the bone marrow. However, such extreme cases are rare and thus would not have a major impact on our results derived from over thirty thousand individuals. We demonstrate this with an extra analysis performed in response to the point about hyperostosis frontalis interna made by reviewer 1.

      (5) Some of the genes and SNPs although significant showed very small correlations. What is their likely physiological significance?

      Bone marrow adiposity is a polygenic trait and we have identified 41 statistically significant loci. We had a discovery sample of approximately 30k individuals which is modest for a GWAS study, so these 41 loci are a lower bound on the number of genes influencing the BMA trait. When a large number of genes influence a trait, the effect size of an individual gene is typically relatively small. However, the SNP heritability estimates of 31.5% indicates that we are able to explain approximately one third of the phenotypic variation with the effect sizes estimated by our GWAS: this is quite a high fraction relative to many other GWASs of biomedical traits.

      (6) The authors could use this excellent manuscript to expand their discussion to include the need for studies like theirs to be also complemented by multi-OMICS studies that will include proteomics and lipidomics of BM, bones, and muscles.

      We agree with the reviewer and hope that such studies will be undertaken in the future. We attempted to point in this direction in the last sentence of the Discussion (line 517): “Future studies can build on our developments to further validate the proposed measure of bone marrow composition and to study its effect on bone, blood, and brain”. Word count limits prevented us from further expanding on the specific kinds of studies that should be performed.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      More moderate concerns include:

      (1) In the "Genetic correlation and overlap" part, it is unclear why a high effect correlation with high SNP overlap is suggestive of vertical pleiotropy, while "moderate overlap and a low correlation of effect sizes ... are more indicative of horizontal pleiotropy". This is not intuitive.

      In vertical pleiotropy (genetics > phenotype A > phenotype B): genetics drive phenotype A and phenotype A drives phenotype B (B has few direct genetic drivers of its own). If we perform a GWAS of phenotype A and a GWAS of phenotype B, one would expect to see a high overlap in the causal variants (because phenotype B is largely indirectly determined by the genetics of phenotype A) and the correlation should be high (either negative or positive) because there is a cause-effect relationship between A and B (whether phenotype A has a positive or negative effect on phenotype B).

      In horizontal pleiotropy: the A and B phenotypes share the same genetic loci (but have no phenotypic influence on each other). In this case, we would observe high overlap in the associated loci, but one would not expect to see a high correlation in effect sizes (across loci) because there is no a priori reason to expect that genes associated with both phenotype A and phenotype B would have a consistent (negative or positive) effect across loci. For example, gene X may increase A and B, whereas gene Y may increase A but decrease B.

      We have made a small update to the relevant part of the Results section and otherwise rely on the explanation above (which will be publicly available along with the manuscript):

      Line 325: “These patterns of high overlap and high effect correlation within this overlap are suggestive of vertical pleiotropy i.e. a molecular mechanism influencing one trait, that in turn influences a second trait, such that most of the variants driving the first trait either have the same or the opposite direction of effect on the second trait.”

      (2) "possible causal effect of BMA on cognition" asks for a formal analysis, like Mendelian randomization.

      Given the current wording, the reviewer is justified in asking for a formal analysis. Since we did not perform this analysis, we have changed the wording:

      Line 434: “Since these are two highly correlated traits [36], a high overlap and correlation of genetic effects for both traits with BMA may be consistent with the hypothesis that BMA could have a causal effect on cognition (Table 1).”

      (3) ll. 132-134: please reword this sentence for clarity: "Using the network-estimated location of the BM within ... averaged these across all datapoints...". Please define threshold of desirable overlap between the predicted BM and the true BM (=0.7?).

      Background: The model predicts the BM localisation for a datapoint (which interval of layers of the 50 layers is bone marrow). We tested the model on a wide variety of simulated data and aim for the overlap to be as close to 1 as possible, but some error is inevitable. We found that the average overlap between the true and predicted bone marrow was only below 0.7 when the bone layer is only 4 mm thick (meaning that the bone marrow is only 1-2 mm thick). This demonstrates the high accuracy of our method in identifying a very small anatomical feature.

      When applying our method to real data, we do not know the truth and therefore cannot compute the overlap between the predicted and true value. It is therefore not possible to identify datapoints where the overlap is poor (e.g. lower than 0.7) and filter them out.

      Given the above, we struggle to understand in what way an overlap threshold is relevant to how we compute the signal intensity for a datapoint. Nevertheless, we recognize that the sentence pointed to by the reviewer is poorly formulated and have tried to make it clearer:

      Line 130-133: “To obtain the BM signal intensity for an individual datapoint of the calvarium, we used the network model to estimate the location of the BM within the datapoint’s intensity array and averaged these BM intensities to get the BM intensity for that datapoint. Then, we averaged these datapoint intensities across the calvarium to produce the global BMA measure for the scan.”

      In the GWAS Results, please clarify the phrases - what was "significantly genetically correlated (Rg=.94)" (also, l. 402, "genetic correlation between the sexes" - in what?).

      Genetic correlation is a statistical measure that quantifies the extent to which two traits (or the same trait in two different cohorts) are influenced by the same genetic factors. Simply put, it is the effect sizes of the SNPs in the two GWASs of interest that are correlated (after correcting for confounding effects, such as linkage desequilibrium). When comparing two GWASs, the standard formulation is to refer to their “genetic correlation”. We made a modification to the text to clarify this:

      Line 239: “We found the male and female GWASs to have low genomic inflation (Figure 3A and Table S4) and to be significantly genetically correlated (Rg=.94, P=6e-27, Figure 3B)”

      "a more than two-fold difference between the sexes" - in which metric?

      We feel that what is being compared is stated clearly in the original sentence:

      Line 262: “A comparison of the male and female effect sizes of the top lead SNPs of each locus revealed 6 loci in which there is a more than two-fold difference between the sexes (loci 10, 18, 26, 30, 32, 37 in Table S5)”.

      (4) Also In GWAS Results, a locus Dlx5 is called "SHFM" in the Supplementary Table.

      Background:

      We identified 41 genome-wide significant loci and named the locus after the gene closest to the top lead SNP (bold in Figure 3C). Other genes in each locus for which genome-wide significant SNPs were eQTLs, are listed below the closest gene in normal font (Figure 3C).

      In table S5, we report details of the top lead SNP for all 41 loci. We report only the nearest gene to the top lead SNP.

      For locus 14, SHFM1 is the closest gene to the top lead SNP whereas DLX5 and DLX6 are genes in the locus for which genome-wide significant SNPs were eQTLs. This explains why DLX5 appears under SHFM1 in Figure 3C, but does not appear in Table S5.

      (5) Please reword MRI jargon - "Dixon method", vertix - should be introduced, as well as abbreviation "KDE".

      We had recognised that the word “vertex” would be confusing and had replaced it by datapoint, but had unfortunately missed one occurrence in the text. This is now corrected.

      Thank you for pointing out the lack of introduction of the term “KDE”. This was only explained in the supplementary materials, but has now been added to the main text:

      Line 500-508: “To ensure between-subject comparability, we used the intensity normalised nu.mgz volume output by FreeSurfer. We validated this approach through comparison with well-established intensity normalization methods; Kernel Density Estimation (KDE), WhiteStripe (WS), Gaussian Mixture Model (GMM), Fuzzy C-Means (FCM), and Z-score normalization (ZS). We found the highest test-retest reliability with our approach (Figure S9), and, together with KDE (based on reference signal intensity in WM), the highest correlation with quantitative T1 relaxation maps (Figure S10).”

      Reviewer #2 (Recommendations for the authors):

      (1) This would be an extremely useful advance for the bone and BMA fields if only it could be confirmed that the T1W signal intensity is actually measuring BMA in some meaningful way. Or, even if not BMA, to confirm what other biological characteristic(s) it is in fact capturing. This is essential for interpreting the findings.

      (2) I note that you have compared the normalized T1W parameter with quantitative T1 data (e.g. Figure S10). However, this doesn't address the fundamental issue, because even these quantitative T1 data (e.g. from MP2RAGE) may not be measuring calvarial BMA. T1W sequences have been used to estimate BM cellularity (if not BMA directly) but are not nearly as precise as water-fat imaging. For example, one study found a reasonable correlation (0.71) between T1 relaxation times and BM fat (https://www.nature.com/articles/s41598-019-57030-5). So, if your normalized T1W parameter shows a correlation of -0.44 with the T1 MP2RAGE MRI signal (Figure S10), what does this mean in terms of how well your parameter reflects the actual BMA adiposity? We can't know this, because we also don't know if the T1 MP2RAGE signal reflects calvarial BMA.

      (3) I think my recommendations are clear from the public review. Ideally, you would be able to compare the skull BM normalized T1W parameter with PDFF data that have T2* correction (since the skull BM cavity is quite small and so may suffer from T2* effects relating to tissue inhomogeneity). But even if you had only dual-echo BMFF data, this would still be much more informative than relying only on T1 data. I hope this can be done so that the findings of the study can be properly interpreted.

      As explained above, we reject this reviewer’s claim that T1-weighted signal intensity cannot be used to perform a GWAS of BMA: other studies have shown that T1-weighted signal intensity is a semi-quantitative measure of fat fraction, we have performed extra analyses that confirm this, and our results further demonstrate this. For further detail on why we reject this criticism, see our response to this reviewer’s comments.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors demonstrate the stereoselective role of D-serine in 1C metabolism, showing that D-serine competes with L-serine and inhibits mitochondrial L-serine transport. They observe expression of 1C metabolites in their metabolomics approach in primary cortical neurons treated with L-serine, D-serine, and a mixture of both. Their conclusions are based on the reduction in levels of glycine, polyamines, and their intermediates and formate. Single-cell RNA sequencing of N2a cells showed that cells treated with D-serine enhanced expression of genes associated with mitochondrial functions, such as respiratory chain complex assembly, and mitochondrial functions, with downregulation of genes related to amino acid transport, cellular growth, and neuron projection extension. Their work demonstrates that D-serine inhibits tumor cell proliferation and induces apoptosis in neural progenitor cells, highlighting the importance of D-serine in neurodevelopment.

      Strengths:

      D-amino acids are a marvel of nature. It is fascinating that nature decided to make two versions of the same molecule, in this case, an amino acid. While the L-stereoisomer plays well-known roles in biology, the D-stereoisomer seems to function in obscurity. Research into these novel signaling molecules is gathering momentum, with newer stereoisomers being discovered. D-serine has been the most well-studied among the different stereoisomers, and we still continue to learn about this novel neurotransmitter. The roles of these molecules in the context of metabolism is not well studied. The authors aim to elucidate the metabolic role of D-serine in the context of neuronal maturation with implications for 1C metabolism and in cell proliferation. The metabolic role of these molecules is just beginning to be uncovered, especially in the context of mammalian biology. This is the strength of the manuscript. The authors have done important work in prior publications elucidating the role of D-amino acids. The advancement of the field of D-amino acids in mammalian biology is significant, as not much is known. The presentation of RNA seq data is a valuable resource to the community, however, with caveats as mentioned below.

      Weaknesses:

      The following are some of the issues that come out in a critical reading of the manuscript. Addressing these would only strengthen and clarify the work.

      (1) Kinetic assessment of D-serine versus L-serine: While the authors mention that D-serine is not a good substrate for SHMT2 compared to L-serine, the kinetic data are presented for only D-serine. In a substrate comparison with an enzyme, data must be presented for L-serine as well to make the conclusion about substrate specificity and affinity. Since the authors talk about one versus another substrate, there needs to be a kinetic comparison of both with Km (affinity). (Ref Figure 2 panel).

      We agree with the reviewer that kinetic parameters for l-serine are important for evaluating the substrate specificity of SHMT2. The hydroxymethyl-transferase activity of SHMT2 toward l-serine has been previously characterized by our co-author Tetsuya Miyamoto (Miyamoto et al., FEBS Journal, 2024; PMID: 37700610), which is appropriately cited in the manuscript (line 143). In that study, the kinetic parameters for l-serine were determined, with a Km of 0.07 ± 0.009 mM and a Kcat of 33.5 ± 0.9 min<sup>-1</sup>. The strong chiral selectivity of SHMT2 for the l-enantiomer in the hydroxymethyl-transferase reaction was also demonstrated in that work. On the other hand, the primary aim of the present study is different from characterizing d-serine as a catalytic substrate for SHMT2. Rather, our goal was to determine whether d-serine interferes with the hydroxymethyl-transferase reaction of SHMT2 by interacting with the l-serine binding site. Accordingly, the analyses shown in Fig. 2B-D were designed to evaluate whether d-serine could structurally occupy or interfere with the l-serine binding pocket of SHMT2. Therefore, our experiments focused on assessing the potential inhibitory effect of d-serine rather than performing a full kinetic comparison of d-serine and l-serine as substrates.

      (2) Molecular Dynamics simulations, while a good first step in modeling interactions at the active site, rely on force fields. These force fields are approximations and do not represent all interactions occurring in the natural world. Setting up the initial conditions in the simulations can impact the final results in non-equilibrium scenarios. The basic question here is this: Is the simulated trajectory long enough so that the system reaches thermodynamic equilibrium and the measured properties converge? Prior studies have shown mixed results with the conclusion that properties of biological systems tend to converge in multi-second trajectories (not nanosecond scales as reported by the authors) and transition rates to low probability conformations require more time. (Ref Figure 2C).

      We thank the reviewer for raising the important point regarding the limitations of molecular dynamics (MD) simulations, including the dependence on force fields and the potential effects of simulation length and initial conditions. We agree that MD simulations represent approximations of molecular behavior and that longer trajectories may be required to fully explore rare conformational states in biological systems.

      In the present study, however, the MD simulations were not intended to provide a comprehensive thermodynamic description of SHMT2 conformational dynamics. Rather, they were used as a structural assessment to evaluate whether d-serine could plausibly occupy the canonical l-serine binding site of SHMT2. As shown in Fig. 2BC and supplementary movie 1, the simulations did not support stable occupation of the l-serine binding pocket by d-serine. Importantly, this structural observation is consistent with our biochemical data showing that d-serine does not inhibit the hydroxymethyl-transferase activity of SHMT2 when l-serine is used as the substrate (Fig. 2D). Together, these results indicate that the inhibitory effect of d-serine on one-carbon metabolism is unlikely to be mediated through direct inhibition of SHMT2.

      As the reviewer correctly notes, it remains possible that d-serine interacts with SHMT2 at sites distinct from the canonical l-serine binding pocket and could exert potential allosteric effects. Indeed, previous work by Miyamoto et al. (FEBS Journal, 2024) demonstrated that SHMT2 exhibits dehydratase activity toward d-serine. However, this reaction is not directly linked to mitochondrial one-carbon metabolism. Therefore, further extensive simulations exploring alternative conformational states or potential allosteric interactions would extend beyond the scope of the present study.

      Importantly, the key conclusions of this study do not rely solely on MD simulations but are supported by multiple independent experimental approaches, including metabolomics, enzymatic assays, and mitochondrial transport analyses.

      (3) The authors use N2a cell line to demonstrate D-serine burden on primary cortical neurons. N2a is an immortalized cell line, and its properties are very different from primary neurons. The authors need to mention a rationale for the use of an immortalized cell line versus primary neurons. The transcriptomic profile of an immortalized cell line is different compared to a primary cell. Hence, the response to D-serine may vary between the two different cell types.

      We thank the reviewer for raising this important point regarding the differences between immortalized cell lines and primary neurons. As the reviewer notes, N2a cells and primary cortical neurons (PCNs) differ in several aspects, including their degree of differentiation and proliferative capacity. We appreciate the opportunity to clarify the rationale for using both systems in this study.

      One-carbon metabolism is known to be particularly active in highly proliferative or relatively undifferentiated cells. Primary cortical neurons are initially obtained as immature neuronal populations and gradually undergo maturation during culture (Fig. S6). In our experiments, we observed that sensitivity to d-serine and dependence on one-carbon metabolism were primarily evident in immature neuronal states rather than in fully mature neurons (Fig. 4EF).

      In this context, immature PCNs share certain metabolic characteristics with proliferative neural cell lines such as N2a cells. Consistent with this idea, the inhibitory effects of d-serine on one-carbon metabolism and cell proliferation were observed in both immature PCNs and N2a cells (Fig. 1E–G, Fig. 2E, Fig. 3AB, and Fig. 4F). Thus, the use of N2a cells provides a complementary experimental model for studying the metabolic vulnerability of immature neural cells that depend on one-carbon metabolism. Importantly, the key findings were consistently reproduced in primary cortical neurons, supporting the physiological relevance of the observations made in N2a cells. For clarity, we added descriptions in lines 107-108 and 167-168 in our revised manuscript.

      (4) In Figure 4D, the authors mention that D-serine activates the cleavage of caspase 3. Figure 4D shows only cleaved caspase 3 as a single band. They need to show the full blot that contains the cleaved fragments along with the major caspase 3 band.

      In our experiments, we used an antibody that specifically recognizes cleaved caspase-3 and does not recognize full-length caspase-3 (Cell Signaling Technology, anti-cleaved caspase-3 antibody, clone 5A1E). Therefore, the Western blot detects only the cleaved caspase-3 fragment at approximately 17 kDa, which appears as a single band in the blot (please see Author response image 1). For clarity, we have also included the antibody information in the Western blot section in the revised Materials and Methods (lines 501-502).

      Author response image 1.

      An original image of western blot for Fig. 4D. An arrow indicates the bands of cleaved caspase-3 (17 kDa).

      (5) In Figure panel 4, the authors use neural progenitor cells (NPCs). They need to demonstrate that the population they are working with is NPCs and not primary neurons. There must be a figure panel staining for NPC markers like SOX2 and PAX6. Also, Figure S5 needs to be properly labeled. It is confusing from the legend what panels B-E refer to? Also, scale bars are not indicated.

      We thank the reviewer for pointing out these issues. First, we have relabeled and rearranged the figure panels in revised Figure S6B-E, and revised the figure legend to provide a clearer description of each panel. We have also added scale bars to the revised images.

      To confirm the identity of neural progenitor cells (NPCs), we performed immunostaining for Nestin, a well-established marker for NPCs, instead of Sox2 and Pax6 (Fig. S6B). Nestin is widely used as a marker for NPCs(Bernal and Arranz, 2018; Bott et al., 2019; Lendahl et al., 1990), and is known to be co-expressed with Sox2 in the mouse embryonic brain(Graham et al., 2003). In addition, Nestin has been identified as a Pax6-bound gene associated with the transcriptional program regulating neural progenitor identity (Thakurela et al., 2016). Consistent with these observations, analysis of published scRNA-seq data from the developing mouse brain (Bella et al., 2021) shows that the RNA expression profile of Nestin (Nes) closely parallels those of Sox2 and Pax6 (new Fig. S9C). Based on these lines of evidence, we consider Nestin-positive cells in our cultures to represent NPCs. Importantly, Nestin-positive cells were co-stained with cleaved caspase-3, suggesting that the apoptotic population corresponds to NPCs (Fig. S6B). Together, these observations support that the apoptotic cells observe in our culture correspond to NPCs rather than differentiated neurons.

      (6) In Supplementary Figure panel 7F, the authors mention phosphatidyl L-serine and phosphatidyl D-serine. A chromatogram of the two species would clarify their presence as they used 2D-HPLC. On an MS platform, these 2 species are not distinguishable. Including a chromatogram of the 2 species would be helpful to the readers.

      We thank the reviewer for this helpful suggestion. As the reviewer correctly noted, mass spectrometry alone cannot distinguish phosphatidyl-d-serine and phosphatidyl-l-serine. To quantify these species separately, lipids were first extracted from cells using the Bligh and Dyer method, followed by phospholipase D treatment to cleave the serine moiety from phosphatidylserine (a schematic of the procedure is shown in Fig. S8D). The released serine was then derivatized with NBD-F and analyzed by 2D-HPLC for enantioselective separation and quantification of d- and l-serine, as described in the Materials and Methods section (“Quantification of glycine and serine enantiomers”). Phosphatidyl-l-serine (Sigma-Aldrich: P0474) was used to generate a standard curve for quantification (Fig. S8E). Because phosphatidyl-d-serine is not commercially available, and because the peak height of free d-serine is equivalent to that of free l-serine in the chromatograms of our 2D-HPLC system, the same standard curve was used to estimate phosphatidyl-d-serine levels.

      As requested by the reviewer, we have now added representative chromatograms of d- and l-serine derived from phosphatidyl-serine in NPC samples (new Fig. S8F).

      (7) The authors mention about enantiomeric shift of serine metabolism during neural development, which appears to be a discussion of prior published data from Hubbard et al, 2013, Burk et al, 2020, and Bella et al, 2021, in Supplementary Figure panels 8 A-E. This should not be presented as a figure panel, as it gives the false impression that the authors have performed the experiment, which is clearly not the case. However, its discussion can well serve as part of the manuscript in the discussion section.

      We appreciate this helpful suggestion. The datasets used in Fig. 4K and Fig. S9 were derived from previously published transcriptomic studies (Hubbard et al, 2013; Burk et al, 2020; Bella et al, 2021). While these studies reported transcriptomic profiles during neuronal development in vitro or in vivo, they did not specifically analyze serine metabolism, one-carbon metabolism, or d-serine biosynthesis, which are the focus of the present study. Therefore, we reanalyzed these publicly available datasets from a metabolic perspective, focusing on genes involved in serine metabolism and one-carbon metabolism. This re-analysis allowed us to examine a developmental shift in serine enantiomer metabolism, we presented the results as figure panels rather than simply citing the datasets in the Discussion. Re-analysis of publicly available transcriptomic datasets to address new biological questions has become a common approach in genomics and transcriptomics studies.

      To avoid the impression that these experiments were performed in this study, we have clearly indicated the original references and clarified this point in the figure legends (Fig. 4K and Fig. S9). These panels are intended to provide supportive evidence for the developmental shift in serine enantiomer metabolism discussed in this study.

      (8) The entire presentation of the section on enantiomeric shift of serine metabolism during neural development (lines 274-312) is a discussion and should be part of the discussion section and not in the results section. This is misleading.

      We thank the reviewer for this comment. This point overlaps with the concern raised in comment 7. Please see our response to comment 7 for a detailed explanation and the revisions made in the manuscript.

      (9) The discussion section is not well written. There is no mention of recent work related to D-serine that has a direct bearing on its metabolic properties. In the discussion section, paragraph 1, the authors mention that their work demonstrates the selective synthesis of D-serine in mature neurons as opposed to neural progenitor cells. This concept has been referred to in prior publications:

      (a) Spatiotemporal relationships among D-serine, serine racemase, and D-amino acid oxidase during mouse postnatal development. PMID:14531937.

      (b) D-cysteine is an endogenous regulator of neural progenitor cell dynamics in the mammalian brain. PMID:34556581.

      We thank the reviewer for this helpful suggestion and for drawing our attention to these studies. We have revised the Discussion to better place our findings in the context of previous work related to d-serine.

      Specifically, we have added references describing the spatiotemporal relationship between d-serine and serine racemase (Srr) during brain development (PMIDs 14531937 and 33592203) (line 333). These studies highlighted the tissue-level (PMID 14531937) and cellular-level (PMID: 33592203) relationship between Srr expression and development. These studies highlighted the role of d-serine in supporting the functional maturation of neurons during postnatal development. In contrast, our study addresses a complementary question of why d-serine is NOT present during embryonic and early postnatal stages of brain development, when proliferative metabolic activity is high. This question is fundamentally different from the previous reports. Our point is that d-serine is not favorable because it interferes with one-carbon metabolism, which is essential for cell proliferation. Therefore, this concept has not been referred to in prior publications. To clarify this point, we added the following sentence to the Discussion (line 325-328): “In addition to the known role of d-serine in the functional maturation of differentiated neurons, our findings highlight a previously unrecognized, stereoselective regulation of cellular metabolism by d-serine, and provide a rationale for its selective synthesis in mature neurons where proliferative metabolic activity is no longer required’.

      We also appreciate the reviewer bringing our attention to the study describing d-cysteine as a regulator of neural progenitor cell dynamics (PMID: 34556581). We have added the description regarding the overlapping and distinct functions of d-cysteine and d-serine to the Discussion (lines 418-436).

      (10) In the abstract, in lines 101 and 102, the authors mention "how d-serine contributes to cellular metabolism beyond neurotransmission remains largely unknown". In 2023, a paper in Stem Cell Reports by Roychaudhuri et al (PMID:37352848) showed that d and l-serine availability impacts lipid metabolism in the subventricular zone in mice, affecting proliferative properties of stem-cell derived neurons using a comprehensive lipidomics approach. There is no mention of this work even in the discussion section, as it bears directly on l and d-serine availability in neurons, which the authors are investigating. In the discussion section in lines 410-411, the authors mention the role of d-serine in neurogenesis, but surprisingly don't refer to the above reference. The role of d-serine in neurogenesis has been demonstrated in the Sultan et al (lines 855-857) and Roychaudhuri et al references.

      We thank the reviewer for highlighting these relevant studies. We have revised the statement in line 99 to avoid overgeneralization and to reflect that the metabolic roles of d-serine are incompletely understood rather than largely unknown. In addition, we have incorporated discussion of previous works, including the work by Roychaudhuri et al. (2023), into the Discussion section (lines 418-436) of our revised manuscript.

      (11) Both D-serine and the structurally similar stereoisomer D-cysteine (sulfur versus oxygen atom) have a bearing on 1C metabolism and the folate cycle. With reference to the folate cycle, Roychaudhuri et al in 2024 (PMID:39368613) have shown in rescue experiments in mice that supplementing a higher methionine diet provides folate cycle precursors to rescue the high insulin phenotype in SR-deficient mice. Since 1C metabolism is being discussed in this manuscript, the authors seem to overlook prior work in the field and not include it in their discussion, even when it is the same enzyme (SR) that synthesizes both serine and cysteine. Since the field of D-amino acid research is in its infancy, the authors must make it a point to include prior work related to D-serine at least, and not claim that it is not known. The known D-stereoisomers are not many, hence any progress in the area must include at least a discussion of the other structurally related stereoisomers.

      We are grateful to the reviewer for drawing our attention to this relevant study.

      We have now incorporated the findings from Roychaudhuri et al. 2024 into the Discussion. In that study, Srr-/- mice exhibited reduced levels of DNMT1 and DNMT3A, resulting in reduced DNA methylation activity. Notably, supplementation with a methyl-donor diet (containing choline, betaine, and methionine) restored the aberrant insulin phenotype in Srr-/- mice. These findings are relevant to our study, as DNA methylation depends on S-adenosyl-methionine (SAM), which is generated through one-carbon metabolism.

      We have expanded the Discussion to include the relationship between d-serine and d-cysteine, both of which are synthesized by Srr in the revised manuscript (lines 418-436).

      (12) Racemases (serine and aspartate) in general are promiscuous enzymes and known to synthesize other stereoisomers in addition to D-serine, D-cysteine, and D-aspartate. A few controls, like D-aspartate, D-cysteine, or even D-alanine must be included in their study to demonstrate the specific actions of D-serine, especially in the N2a cell treatment experiments. Cysteine and Serine are almost identical in structure (sulfur versus oxygen atom), and both are synthesized by serine racemase (published). Cysteine has also been very recently shown to inhibit tumor growth and neural progenitor cell proliferation. (PMIDs: 40797101 and 34556581). How the authors' work relates to the existing findings must be discussed, and this would put things in perspective for the reader.

      We thank the reviewer for this comment regarding the need to demonstrate the specificity of d-serine. As shown in Fig. S7A, d-serine, but not other d-amino acids commonly detected in mammals (Gonda et al., 2023), including d-aspartate, d-alanine, and d-proline, induced the cleavage of caspase-3 under the same experimental conditions, supporting the specific effect of d-serine in our system. We agree that d-cysteine shares structural similarity with d-serine and is also synthesized by Srr, suggesting potential functional overlap with d-serine. To place our findings in this context, we have added a paragraph in the Discussion (lines 418-436) describing the similarities of d-serine and d-cysteine. While both molecules may exert anti-proliferative effects, their underlying mechanisms appear to differ. Notably, supplementation with SAM, methionine, or glutathione did not rescue d-serine-induced growth inhibition in our system (Fig. S5J), suggesting that its effects are not primarily mediated through methylation or sulfur metabolic pathways. Instead, d-serine suppresses cellular proliferation by limiting mitochondrial l-serine availability and one-carbon metabolism. These observations highlight mechanistic divergence between d-serine and d-cysteine.

      Reviewer #2 (Public review):

      Summary:

      This study by Suzuki et al. reports an interesting stereo-selective role of D-serine in regulating one-carbon metabolism during neurodevelopment to adapt the functional transition, probably through the competition with mitochondrial transport of L-serine. The authors provide a multi-layered set of evidence, including metabolomics, enzyme assays, mitochondrial transport competition, and functional assays in immature/neural progenitor cells, to build up a conceptual integration of D-serine as both a neurotransmitter and a metabolic regulator in the central neural system, which raises a broad potential interest to the neuroscience and metabolism communities.

      Strengths:

      This work provides a conceptual advance that D-serine not only serves as a traditional neurotransmitter in the central neural system but also critically contributes to metabolic regulation of neural cells. The authors performed solid metabolomic assays to validate the suppressive effect of D-serine on the one-carbon metabolic pathway, providing some evidence that D-serine competitively inhibits mitochondrial serine transport, but not directly impairs SHMT2 enzymatic activity. All these data indicate a critical role of D-serine synthesis during neural maturation and suggest a potential translational strategy for targeting serine metabolism in neural tumors.

      Weaknesses:

      (1) The detailed mechanism by which D-serine competes with L-serine for its mitochondrial transport is not investigated. For example, although the authors made some discussion, they did not provide direct genetic or biochemical evidence linking these effects to the specific transporters, such as SFXN1.

      We thank the reviewer for this important comment regarding the mitochondrial l-serine transport mechanism. To address this point, we performed additional experiments using N2a cells in which Sfxn1 was knocked down by siRNA. Under semi-permeabilized cell conditions, we newly examined the effect of d-serine on mitochondrial L-serine transport.

      Interestingly, even under conditions where Sfxn1 expression was markedly suppressed, d-serine still inhibited mitochondrial d-serine transport. Given that SFXN1 is known to function redundantly with its paralogs (SFXN2–SFXN5) in mitochondrial serine transport (Kory et al., 2018), these findings suggest that d-serine may interfere with l-serine transport not only through SFXN1 but potentially through multiple members of the SFXN transporter family. These new data have been added as Fig. S3, and the corresponding results and discussion have been incorporated into the revised manuscript (lines 158–165 and 379-385).

      (2) Unlike tumor cells, where SHMT2 usually plays a predominant role in catalyzing serine/THF-derived one-carbon metabolism, normal cells may employ both SHMT1 and SHMT2 to do the work. Even under certain conditions that SHMT2-mediated one-carbon metabolism is suppressed, the activity of SHMT1 could be elevated for compensation. Thus, it is important to investigate whether D-serine affects SHMT1 activity or changes the balance between SHMT1- and SHMT2-mediated one-carbon metabolism. To this aim, the authors are strongly encouraged to perform a metabolic flux assay (MFA) by using 13C-labeled L-serine in the model cells in the presence and absence of D-serine.

      We thank the reviewer for this thoughtful comment regarding the potential contribution of cytosolic SHMT1 to one-carbon metabolism. As the reviewer notes, while mitochondrial SHMT2 is generally considered the predominant enzyme supporting one-carbon metabolism in proliferating or tumor cells, SHMT1 in the cytosol may function in a complementary manner in normal cells. In our experiments using primary cortical neurons, we indeed observed that the sensitivity to d-serine differed between immature and mature neuronal states (Fig. 4E). This observation suggests that the relative contribution of SHMT1- and SHMT2-mediated one-carbon metabolism may vary depending on the differentiation status of the cells.

      However, the primary focus of the present study was to investigate the mechanism underlying the anti-proliferative effect of d-serine in proliferative or undifferentiated neural cells. Our data demonstrate that d-serine inhibits one-carbon metabolism primarily by limiting mitochondrial l-serine availability through inhibition of mitochondrial l-serine transport. Therefore, a detailed analysis of the compensatory balance between SHMT1 and SHMT2 after disruption of mitochondrial serine transport falls beyond the central scope of the present study. Importantly, previous work by Miyamoto et al. (FEBS Journal, 2024) demonstrated that both SHMT1 and SHMT2 exhibit strong stereoselectivity for l-serine in their hydroxymethyl-transferase activity and do not utilize d-serine as a substrate. These findings make a direct effect of d-serine on hydroxymethyl-transferase activity of SHMT1 unlikely.

      Nevertheless, we agree that differences in the anti-proliferative effects of d-serine across cell types or differentiation states could reflect variations in cellular dependence on one-carbon metabolism or potential compensation by SHMT1. To address this point, we have expanded the Discussion section to clarify the possible contribution of SHMT1 and the limitations of the present study (lines 367–376). While isotope tracing analysis would be valuable for further dissecting compartmentalized one-carbon fluxes, such analyses would primarily address the relative contributions of SHMT1 and SHMT2 rather than the mitochondrial serine transport step that constitutes the central mechanism identified in this study.

      (3) A defect in serine-derived one-carbon metabolism may cause multiple cellular stress responses. It is valuable to detect whether cellular NADPH/NADH, GSH, or ROS is altered before and after D-serine treatment.

      We appreciate this insightful comment regarding potential cellular stress responses associated with impaired one-carbon metabolism. Consistent with the reviewer’s suggestion, we examined markers related to redox status. Transcriptomic analysis revealed compensatory changes in genes involved in mitochondrial metabolic function, including components of the NADH dehydrogenase complex (Fig. 2IJ and Fig. S4A), suggesting metabolic adaptation to d-serine treatment. In response to the reviewer’s comment, we measured GSH levels in NPCs and found that d-serine treatment led to a reduction of GSH (new Fig. S7F), indicating altered redox balance. However, supplementation with exogenous GSH did not rescue d-serine-induced cell death (Fig. S7E). These results suggest that while d-serine induces changes in cellular redox status, including GSH depletion, redox imbalance alone is unlikely to be the primary driver of cell death in this context.

      (4) The physiological relevance between D-serine and neural cell maturation/death should be further tested and discussed, since the dosage of D-serine used in the in vitro assay is much higher than that in physiological conditions.

      We thank the reviewer for this comment regarding the physiological relevance of the d-serine concentrations used in our study. We agree that the concentrations of d-serine required to compete with l-serine for mitochondrial transport are higher than those typically observed under physiological conditions, which we acknowledged in the Discussion of our manuscript (lines 386-388). Importantly, however, in vivo, d-serine levels are tightly regulated in a spatiotemporal manner during brain development, with low levels during embryonic and early postnatal stages and increased levels upon neural maturation. This temporal regulation coincides with the transition from proliferative neural progenitor states to differentiated neurons, thereby limiting the potential for d-serine to interfere with one-carbon metabolism during periods of active cell proliferation. Thus, while the concentrations used in vitro may exceed physiological levels, they allow us to uncover a latent metabolic effect of d-serine that may become relevant under specific cellular or developmental contexts. To clarify these points, we have revised the Discussion (lines 388-395) in the revised manuscript to more explicitly address the relationship between d-serine dosage and physiological relevance.

      Reviewer #3 (Public review):

      Summary:

      This manuscript presents a comprehensive and well-executed investigation into the metabolic role of D-serine in the central nervous system. The authors provide solid evidence that D-serine competitively inhibits mitochondrial L-serine transport, thereby impairing one-carbon metabolism. This stereoselective mechanism reduces glycine and formate production, suppresses cellular proliferation, and induces apoptosis in immature neural cells and glioblastoma stem cells. Developmental analyses further reveal a physiological enantiomeric shift in serine metabolism during neurogenesis, aligning with the transition from proliferation to maturation. Overall, the study bridges developmental neurobiology, cancer metabolism, and amino acid transport, uncovering a previously unrecognized metabolic function of D-serine beyond its role in neurotransmission.

      Strengths:

      (1) The discovery that D-serine inhibits one-carbon metabolism by competing for mitochondrial L-serine transport-rather than through enzymatic inhibition or receptor-mediated signaling-represents a significant and previously underappreciated mechanism. This finding has broad implications for understanding metabolic regulation during neurodevelopment and offers potential relevance for targeting metabolic vulnerabilities in cancer.

      (2) The authors integrate metabolomics, mitochondrial transport assays, molecular dynamics simulations, genetic and pharmacologic perturbations, transcriptomics, and both in vitro and ex vivo models. The breadth of experimental approaches, combined with the coherence of the findings across systems, provides strong support for the central conclusions and enhances the overall impact of the study.

      (3) The temporal shift in D-/L-serine levels during neurodevelopment is elegantly linked to the transition from proliferative to mature neuronal states. The selective vulnerability of neural progenitors and tumor cells-contrasted with the resistance of mature neurons-highlights a biologically meaningful and potentially targetable metabolic distinction.

      Weaknesses:

      (1) While the authors attribute D-serine's metabolic effects to competition with mitochondrial L-serine transport, the specific identity of the transporter(s) mediating this process remains undefined. This represents a meaningful mechanistic gap, as the central conclusion depends on D-serine limiting mitochondrial L-serine availability to inhibit one-carbon metabolism.

      We thank the reviewer for this insightful comment regarding the identity of the mitochondrial l-serine transporter. As this concern overlaps with the point raised by Reviewer 2 (Weakness 1), we refer the reviewer to our response there for a detailed description of the additional experiments performed. Briefly, we conducted new experiments using siRNA-mediated knockdown of Sfxn1 in N2a cells and examined mitochondrial l-serine transport under semi-permeabilized conditions. Notably, even with marked suppression of Sfxn1 expression, d-serine continued to inhibit mitochondrial l-serine transport. Given that SFXN family members (SFXN1–SFXN5) are reported to function redundantly in mitochondrial serine transport (Kory et al., 2018), these findings suggest that d-serine may interfere with l-serine transport not only via SFXN1 but potentially across multiple SFXN paralogs. These results have been incorporated into the revised manuscript and are presented in Fig. S3, with the corresponding discussion added to the Results and Discussion sections (lines 158–164 and 379-385).

      (2) The effective concentrations of D-serine used in vitro (IC<sub>50</sub> ≈ 1-2 mM) exceed typical brain levels (~0.3 mM). While the authors acknowledge this, a more focused discussion on whether higher local D-serine concentrations could arise in specific microenvironments - such as synaptic compartments, tumor niches, or pathological states-would help contextualize the in vitro findings and strengthen their physiological relevance. For example, disruptions in D-serine clearance or altered expression of serine racemase and transporters in disease contexts could lead to localized accumulation. Moreover, differences between extracellular and intracellular D-serine pools - and the mechanisms governing their regulation - may further influence its metabolic impact in vivo.

      We appreciate this insightful comment regarding the physiological relevance of the d-serine concentrations used in vitro. We agree that the effective concentrations observed in our assays exceed typical bulk brain levels. To address this point, we have expanded the Discussion (lines 386-399) to consider conditions under which locally elevated d-serine concentrations may arise in vivo. These additions provide a more nuanced interpretation of the relationship between the concentrations used in vitro and the potential physiological contexts in which d-serine may exert metabolic effects.

      (3) While the manuscript focuses on neural stem/progenitor cells and neural tumors, it remains unclear whether the anti-proliferative effects of D-serine are specific to neural lineages or extend to other highly proliferative non-neural cell types. A brief discussion addressing this point would help clarify the scope of D-serine's metabolic impact and whether its mechanism of action reflects a unique vulnerability in neural cells or a more general feature of proliferative metabolism. This distinction is particularly relevant for assessing the broader therapeutic potential of targeting mitochondrial L-serine transport.

      We thank the reviewer for this comment regarding the potential generality of the anti-proliferative effects of d-serine. We agree that it is important to clarify whether the observed effects are specific to neural lineages or reflect a broader vulnerability of proliferative cells. In our study, we focused on neural progenitor cells and neural tumour models. However, the mechanism identified here, namely limitation of mitochondrial L-serine availability, targets a fundamental metabolic pathway that supports cell proliferation. Therefore, it is possible that similar effects may extend to other highly proliferative cell types beyond the neural lineage. To address this point, we have expanded the Discussion (lines 408-410) to clarify that the observed effects of d-serine may reflect a general metabolic vulnerability associated with proliferative states, while also noting that the degree of sensitivity is likely to depend on cell-type-specific reliance on mitochondrial one-carbon metabolism.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Minor issues:

      The authors mention terms like neural tissue (line 109) and neural tumor cells (line 182). Neural tissue can mean anything under the sun. They need to mention the specific tissue being studied or investigated.

      We have changed the terms.

      Reviewer #2 (Recommendations for the authors):

      (1) The figure items were not well organized, and the legends were difficult to read since they apparently lacked key information. For example, in Figure 2D, did the author intend to show SHMT2 activity? It is not clear how this assay was performed (in vitro or in vivo experiment?). Additionally, more detailed information should be provided in the Methods section.

      We have improved the organization of figures and revised figure legends for better readability.

      (2) The writing of the manuscript should be significantly improved, using a professional editing service, in order to increase the readability.

      We appreciate the reviewer’s comment. The manuscript was professionally edited prior to submission. Nevertheless, we have made targeted revisions to the Introduction and Discussion sections to improve overall readability. We hope that these revisions have improved the clarity of the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) The core mechanism centers on competition for mitochondrial L-serine transport, yet the identity of the transporter(s) involved remains speculative. While Kory et al. (2018) identified SFXN1 as a mitochondrial L-serine transporter, this connection is not directly addressed in the current study. It would strengthen the manuscript to clarify whether SFXN1 or related isoforms are expressed in the neural cell models used and whether their expression patterns correspond with the observed D-serine sensitivity. Even if functional validation is beyond the current scope, a more detailed discussion of potential transporter candidates and the limitations of existing data would provide important mechanistic context and help frame future directions.

      We thank the reviewer for this insightful comment regarding the mitochondrial l-serine transporter and the potential involvement of SFXN family members. As this concern overlaps with the point raised by Reviewer 2 (Weakness 1), we refer the reviewer to our response there for a detailed description of the additional experiments performed.

      Briefly, we performed siRNA-mediated knockdown of Sfxn1 in N2a cells and examined mitochondrial l-serine transport under semi-permeabilized conditions. Even under conditions of marked suppression of Sfxn1 expression, d-serine continued to inhibit mitochondrial l-serine transport. Given that members of the SFXN family have been reported to function redundantly in mitochondrial serine transport (e.g., Kory et al., 2018), these findings suggest that d-serine may affect l-serine transport not only through SFXN1 but potentially across multiple SFXN paralogs.

      These new results have been incorporated into the revised manuscript (Fig. S3), and the relevant discussion has been expanded to clarify the potential roles of SFXN family transporters and the current limitations in defining the exact transporter responsible (lines 158–165 and 379-385).

      (2) The authors use ex vivo brain slice cultures with tumor xenografts to demonstrate tissue-level relevance, which is a valuable strength of the study. However, additional context would enhance its translational significance. It would be helpful to discuss whether in vivo D-serine administration (e.g., ICV or systemic) is feasible and safe, especially given the high concentrations required in vitro. Briefly addressing whether genetic models, such as Srr knockout mice, support a role for D-serine in tumor progression or neurodevelopment would also strengthen the interpretation.

      We thank the reviewer for this important comment regarding the translational relevance of d-serine administration in vivo. To address this point, we performed additional exploratory experiments using a subcutaneous tumour xenograft model in nude mice. Because the inhibitory effect of d-serine on one-carbon metabolism becomes evident under l-serine–limited conditions, tumor-bearing mice were fed an l-serine/glycine–deficient diet and administered d-serine in drinking water.

      However, when d-serine was provided at high concentrations (≥ 500 mM) in drinking water, the mice exhibited marked behavioral abnormalities, including increased aggression and other abnormal behaviors. Due to these adverse effects, the experiment was ethically terminated. While our study focuses on the inhibitory effect of d-serine on one-carbon metabolism, d-serine is also a physiological co-agonist of the NMDA receptors, and therefore high systemic concentrations may influence neuronal excitability in vivo. These observations suggest that the concentrations required to achieve antitumor effects may be associated with significant neurological side effects.

      Based on these findings, we consider that direct administration of d-serine itself may have limited therapeutic applicability as an antitumor reagent. Instead, future development of derivatives or strategies that retain the metabolic inhibitory effect while minimizing NMDA receptor–mediated effects may be required.

      Regarding serine racemase (Srr) knockout models, xenograft tumour experiments would require an immunodeficient background, which makes the generation and use of such compound models technically challenging. Therefore, we did not pursue this approach in the present study. We have incorporated these considerations into the revised Discussion (lines 413-417) to clarify the translational implications and current limitations of our findings.

      Bella DJD, Habibi E, Stickels RR, Scalia G, Brown J, Yadollahpour P, Yang SM, Abbate C, Biancalani T, Macosko EZ, Chen F, Regev A, Arlotta P. 2021. Molecular logic of cellular diversification in the mouse cerebral cortex. Nature 595:554–559. DOI: https://doi.org/10.1038/s41586-021-03670-5, PMID: 34163074

      Bernal A, Arranz L. 2018. Nestin-expressing progenitor cells: function, identity and therapeutic implications. Cellular and Molecular Life Sciences 75:2177–2195. DOI: https://doi.org/10.1007/s00018-018-2794-z, PMID: 29541793

      Bott CJ, Johnson CG, Yap CC, Dwyer ND, Litwa KA, Winckler B. 2019. Nestin in immature embryonic neurons affects axon growth cone morphology and Semaphorin3a sensitivity. Molecular Biology of the Cell 30:1214–1229. DOI: https://doi.org/10.1091/mbc.e18-06-0361, PMID: 30840538

      Graham V, Khudyakov J, Ellis P, Pevny L. 2003. SOX2 Functions to Maintain Neural Progenitor Identity. Neuron 39:749–765. DOI: https://doi.org/10.1016/s0896-6273(03)00497-5, PMID: 12948443

      Lendahl U, Zimmerman LB, McKay RDG. 1990. CNS stem cells express a new class of intermediate filament protein. Cell 60:585–595. DOI: https://doi.org/10.1016/0092-8674(90)90662-x, PMID: 1689217

      Thakurela S, Tiwari N, Schick S, Garding A, Ivanek R, Berninger B, Tiwari VK. 2016. Mapping gene regulatory circuitry of Pax6 during neurogenesis. Cell Discovery 2:15045. DOI: https://doi.org/10.1038/celldisc.2015.45, PMID: 27462442

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors present a nanobody-based pulse-labeling system to track yeast NPCs. Transient expression of a nanobody targeting Nup84 (fused to NeonGreen or an affinity tag) permits selective visualization and biochemical capture of NPCs. Short induction effectively labels NPCs, and the resulting purifications match those from conventional Nup84 tagging. Crucially, when induction is repressed, dilution of the labeled pool through successive cell cycles allows the visualization of "old" NPCs (and potentially individual NPCs), providing a powerful view of NPC lifespan and turnover without permanently modifying a core scaffold protein.

      Strengths:

      (1) A brief expression pulse labels NPCs, and subsequent repression allows dilution-based tracking of older (and possibly single) NPCs over multiple cell cycles.

      (2) The affinity-purified complexes closely match known Nup84-associated proteins, indicating specificity and supporting utility for proteomics.

      We thank the reviewer for this evaluation

      Weaknesses:

      (1) Reliance on GAL induction introduces metabolic shifts (raffinose -> galactose -> glucose) that could subtly alter cell physiology or the kinetics of NPC assembly. Alternative induction systems (e.g., β-estradiol-responsive GAL4-ER-VP16) could be discussed as a way to avoid carbon-source changes.

      Indeed, this could be an improvement, and we mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3.

      (2) While proteomics is solid, a comprehensive supplementary table listing all identified proteins (with enrichment and statistics) would enhance transparency.

      Indeed, we now provide source data showing LFQ intensities, fold-enrichment and statistics for all detected proteins.

      (3) Importantly, the authors note that the method is particularly useful "in conditions where direct tagging of Nup84 interferes with its function, while sub-stoichiometric nanobody binding does not." After this sentence, it would be valuable to add concrete examples, such as experiments examining NPC integrity in aging or stress conditions where epitope tags can exacerbate phenotypes. These examples will help readers identify situations in which this approach offers clear advantages.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants, and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks, and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #2 (Public review):

      Summary:

      This preprint describes a practical and useful approach for labeling and tracking NPCs in situ. While useful applications including timelapse imaging, affinity purification, or proximity labeling are envisioned, addressing some outstanding technical questions would give a clearer picture of the sensitivity and temporal resolution of this approach.

      Strengths:

      Clever use of a fluorescently conjugated nanobody that binds directly to the core scaffold nucleoporin Nup84 with nanomolar affinity.

      We thank the reviewer for this evaluation

      Weaknesses:

      The decrease in nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over during this time. However, it is also possible that the nanobody: Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      We thank the reviewer again for the constructive feedback and thoughts.

      Reviewer #3 (Public review):

      Summary:

      Submitted to the Tools and Resources series, this study reports on the use of a single-domain antibody targeting the nucleoporin Nup84 to probe and track NPCs in budding yeast. The authors demonstrate their ability to rapidly label or pull down NPCs by inducing the expression of a tagged version of the nanobody (Figure 1).

      Strengths:

      This tool's main strength is its versatility as an inexpensive, easy-to-set-up alternative to metabolic labelling or optical switching. This same rationale could, in principle, be applied to the study of other multiprotein complexes using similar strategies, provided that single-chain antibodies are available.

      We thank the reviewer for this evaluation

      Weaknesses:

      This approach has no inherent weaknesses, but it would be useful for the authors to verify that their pulse labelling strategy can also be used to detect assembly intermediates, structural variants, or damaged NPCs.

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure but consider such studies to be beyond the scope of this study.

      Overall, the data clearly show that Nup84 nanobodies are a valuable tool for imaging NPC dynamics and investigating their interactomes through affinity purification.

      We thank the reviewer again for the constructive feedback and thoughts.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 1A, and although it is partially mentioned in the legend, it would be helpful to indicate precisely when cells are grown in raffinose, when galactose is added for induction, and when glucose is used to terminate expression.

      We included “galactose” and “glucose” to Panel A to indicate induction and termination of expression, respectively.

      (2) Related to the previous point, consider mentioning the GAL4-ER-VP16 (ADGEV) estradiol-inducible system as an optional strategy to avoid carbon shifts and potentially reduce cell-to-cell variability.

      We mention the benefits of an inducible system that does not alter the cell’s metabolic state in the discussion on p.3

      (3) Add a brief sentence explaining that the ZZ tag is derived from Protein A and binds IgG Fc.

      This information is now added on p.2

      (4) The statement "all Nups significantly coenriched with VHH[Nup84]-ZZ..." is likely inaccurate, since not all Nups are labeled in panels F-H, and some basket components are missing in panel I (particularly basket components such as Nup60, Nup1). Consider revising to "most Nups significantly coenriched...". In panel I, please include a clearly non-enriched protein as a visual reference for the color scale.

      We are very grateful to the reviewer for pointing this out. We accidentally used a faulty filtering on the dataset to generate figure panel I, omitting several Nups that were reproducibly found in all replicas. All Nups, except for Gle1 and Pom33, were detected reproducibly.

      We have made the following adjustments to the figure panel and accompanying text:

      In Fig. 1I, we included the missing Nups and 5 proteins that co-purified with VHH[Nup84] but not specifically enriched, as the reviewer suggested. They cluster in a separate group and their abundance is not going up in time. We randomly selected these 5 proteins from the list of genes that were reproducibly found in all four timepoints.

      For clarity, we removed the NTRs

      We changed the text to “we found that all Nups, except Gle1 and Pom33, significantly coenriched with VHH[Nup84]-ZZ” on p.2.

      We updated the methods section, describing the clustering method and how we selected the 5 random proteins

      (5) Provide a supplementary spreadsheet with LFQ intensities, fold-enrichment, and statistics for all detected proteins. This will address questions about missing Nups and support transparency.

      This information is now added as Source data Figure 1.

      (6) Directly after the statement "Amongst others this is useful in conditions where direct tagging of Nup84 interferes with its function, while sub stoichiometric nanobody binding does not," it would be useful to include concrete instances, such as stress or aging conditions, where Nup84 tagging may sensitize NPC integrity.

      Indeed, we agree this would be useful. For example, in Nup1Δct and Nup60Δ mutants, GFP-tagging of Nup84 leads to slower growth and increased cell size (Ollivaud et al., BioRxiv). We have however not extensively tested nanobody expression in these mutants and cannot conclude that it has no interfering effects. We therefore rephrased to “while sub-stoichiometric nanobody binding does may not, …”. Another situation where we find the nanobody-based labeling useful is when we want to assess the structural integrity (IPs) and localization (imaging) of NPCs in mutant strains, but prefer not to use tagged Nups in the actual experiments. In these cases, we transiently express the Nup84 nanobody to perform these checks and then carry out the experiments without the nanobody to avoid any tag-related interference. We hence also added “,…or when the temporary introduction of a ZZ- or mNG-tagged nanobody allows assessment of the integrity or localization of mutant NPCs prior to performing experiments without the nanobody.

      (7) In panels K and L, since individual points correspond to biological replicates, overlaying a box plot obscures much of the data. Consider overlaying the means per replicate instead of box plots: see the "SuperPlots" approach for a clear explanation of how to present this (PMID: 32346721).

      We thank the reviewer for the “SuperPlots” suggestion, and we agree that representing the data in this way improves the visualization of individual biological replicates. We have updated the summarizing overlay in figures in panel K and L to represent the means per replicate instead of boxplots.

      (8) I spotted a few typos ("Lasty" ? "Lastly"; "in maintained" vs. "is maintained").

      Thank you, these are corrected

      Overall, this is a neat, well-executed methodological advance with clear value to the NPC field and potentially other complex assemblies. I look forward to seeing a revised version.

      Thank you!

      Reviewer #2 (Recommendations for the authors):

      Based on the recent structural analyses and NPC modeling using this nanobody, how accessible is the Nup84 epitope expected to be within the fully assembled NPC? While the data shown indicate that nanobody labeling of NPCs is readily detectable, stating this clearly would help motivate the approach and interpret the resulting data.

      We now included such a statement in the introduction on p.1.

      The decrease of nanobody labeling over 8 hours of chase period is interpreted to indicate that NPCs turn over due to cell division during this time window. However, it is also possible that nanobody:Nup84 association is disrupted during mitosis by phosphorylation, other PTMs, or structural remodeling.

      We thank the reviewer for this thought. It is actually not turnover that we propose to underly the decrease in nanobody labeling, but rather the dilution of labelled NPC to the daughter cell. The current data do not support the interpretation that the nanobody: Nup84 association is disrupted as proposed by the reviewer. The exchange of individual Nups, including Nup84, is slow with half-times in the order of hours (Hakhverdyan et al. 2021; Rabut, Doye, and Ellenberg 2004), and the nanobody: Nup84 association is very stable, namely in the nanomolar range (Nordeen et al. 2020). The association of nanobody with NPCs is thus expected to be very stable. Instead, dilution of labelled NPCs to the daughter - approximately 40% of the existing NPCs are transmitted to the daughter cell in each division (Zsok et al. 2024; Khmelinskii et al. 2010) – will lead to significant decreases in nanobody labelling over time. As the reviewer is likely aware, baker’s yeast NPCs – in contrast to mammalian NPCs - remain largely intact during cell division as there is no nuclear envelope breakdown.

      Reviewer #3 (Recommendations for the authors):

      (1) As mentioned above, to assess the general relevance of this tool, it would be informative to verify whether the VHH[Nup84] nanobody can access and detect NPC species under conditions that challenge their structural organization or biogenesis, for example, in nucleoporin mutants or under stress. The authors could, for instance, analyze the localization of VHH[Nup84] in yeast strains harboring clustered NPCs (nup133Δ), or following stresses known to impact NPC organization (e.g., osmotic stress or energy depletion; PMID: 34762489).

      We agree with the reviewer that it would be informative to see if VHH[Nup84] can bind its epitope in the context of an altered NPC structure and tried to include such data. Unfortunately, this was not successful, and further efforts are beyond the scope of his study. Following the reviewer’s suggestion, we expressed VHH[Nup84] in nup133∆N (nup133∆2-300) (Doye, Wepf, and Hurt 1994) following the experimental set-up in panel A and examined its localization. However, at t=2hrs hardly any nanobody signal was detectable in nup133∆N (see Author response image 1, upper panel A) and only after overnight expression nanobody-labelled NPC clusters are detectable (bottom panel A). Considering that expression levels of free mNG are also lower at t=2hrs in nup133∆N cells compared to WT cells (Author response image 1, panel B), it appears that protein expression under the Gal system is generally reduced in a nup133∆N background. These expression level differences between nup133∆N and WT preclude statements about the accessibility of the Nup84 epitope in nup133∆N. We note that nup133∆N cells do not have general mRNA export defects (Doye, Wepf, and Hurt 1994), so other inducible systems may be better suited for such analysis.

      Author response image 1.

      Expression level differences in WT and Nup133∆N cells. Left: localization of VHH[Nup84]-mNG in Nup133∆N cells at t=2hr following a 20-minute induction pulse and after overnight 0.5% galactose (ON) induction. Right: mNG levels in WT and Nup133∆N cells at t=2hr following a 20-minute induction pulse. Brightness/contrast settings are identical between the two panels. All panels are sum slices projections from 30 z-slices of 0.1µm. Scale bar = 5 µm.

      (2) Since outer rings are found on both sides of NPCs (i.e., the cytoplasmic and nuclear faces), could the authors indicate whether the VHH[Nup84] nanobody can enter the nucleus and probe the nuclear outer rings? Along these lines, it would be useful to provide a summary of the structural organization of NPCs in the introduction.

      Thank you, we have added a sentence on the localization of Nup84 in NPCs in the introduction. Based on what is known about influx (nuclear transport receptor-independent nuclear entry) of proteins with similar size and surface properties (Popken et al. 2015; Timney et al. 2016), the nanobody can rapidly enter the nucleus and hence bind Nup84 on both the nuclear and cytoplasmic side. We have no data to answer if binding might initially be biased towards cytosolic VHH[Nup84] binding the cytoplasmic outer rings.

      (3) The authors state that VHH[Nup84] and direct Nup84 detection are indistinguishable (p. 2). Could they provide images of the endogenously tagged Nup84-GFP strain for comparison?

      We have now included a pairwise comparison in a Figure 1 – supplement 1.

      Minor corrections:

      (1) There are a few typos that need correcting: 'Nup84Δ' (p. 1; should read 'nup84Δ') and 'promotor' (p. 2; should read 'promoter').

      Thank you, these are corrected

      (2) The reference 'Veldsink et al. 2025' (quoted in the PunctaFinder analysis description on page 8) does not appear in the References section.

      Thank you, these are corrected.

      We thank the reviewer again for the constructive feedback and thoughts.

      References

      Doye, V., R. Wepf, and E. C. Hurt. 1994. 'A novel nuclear pore protein Nup133p with distinct roles in poly(A)+ RNA transport and nuclear pore distribution', EMBO J, 13: 6062-75.

      Khmelinskii, Anton, Philipp J. Keller, Holger Lorenz, Elmar Schiebel, and Michael Knop. 2010. 'Segregation of yeast nuclear pores', Nature, 466: E1-E1.

      Popken, Petra, Ali Ghavami, Patrick R. Onck, Bert Poolman, and Liesbeth M. Veenhoff. 2015. 'Size-dependent leak of soluble and membrane proteins through the yeast nuclear pore complex', Molecular Biology of the Cell, 26: 1386-94.

      Timney, Benjamin L., Barak Raveh, Roxana Mironska, Jill M. Trivedi, Seung Joong Kim, Daniel Russel, Susan R. Wente, Andrej Sali, and Michael P. Rout. 2016. 'Simple rules for passive diffusion through the nuclear pore complex', Journal of Cell Biology, 215: 57-76.

      Zsok, J., F. Simon, G. Bayrak, L. Isaki, N. Kerff, Y. Kicheva, A. Wolstenholme, L. E. Weiss, and E. Dultz. 2024. 'Nuclear basket proteins regulate the distribution and mobility of nuclear pore complexes in budding yeast', Mol Biol Cell, 35: ar143.

    1. Author response:

      Reviewer #1:

      Major comments

      (1) Although I myself believe that the datasets in this study should be more consistent and comprehensive, the authors should perform a data mining analysis of previously reported transcriptomic changes of these mutants or similar mutants in the same longevity pathway and compare the reported changes with their findings to highlight the necessity and advances of this study.

      According to this suggestion, we have compared the differentially expressed genes identified in this study to previous gene expression studies involving these long-lived mutant strains. To our knowledge no previous studies have examined gene expression in sod-2 or ife-2 mutants, and at the time that we performed the RNA sequencing gene expression in osm-5 worms had not been examined (it took us a long time to complete this paper). We have included weighted Venn diagrams to illustrate the overlap and supplemental tables to list the overlapping di erentially expressed genes. For our current study, we felt it was important to compare RNA-seq data generated under exactly the same experimental and analysis paradigms in order to best compare across the nine long-lived mutants. These new analyses are included in Figures S19 – S25 and Table S2 . Please see lines 111-114, Figure S19-25, and Table S2.  

      (2) This manuscript does not perform any regulon or transcription factor (TF) analyses. TFs are the drivers of the transcriptomic changes and multiple conserved TFs (e.g., daf-16) have already been identified in these pathways. Therefore, it is necessary to examine and compare the regulons/TFs in these new datasets by bioinformatics. Such analyses can: a) provide more information of the driving force of these transcriptomic changes; b) show the role of these known longevity TFs; c) propose new TFs driving longevity; d) support the findings of 'longevity strategies' and 'longevity groups' from the perspective of TFs.

      According to this suggestion, we have now performed transcription factor analysis on the RNA-seq data to determine which transcription factors might be driving the longevity-associated transcriptional changes. To do this we used two complementary approaches: (1) transcription factor inference, which is based on the coordinated expression changes of known transcription factors; and (2) motif enrichment analysis, which is based on identifying transcription factor binding motifs in the promoters of di erentially expressed genes. After identifying which transcription factors were identified for each individual mutant, we then compared the identified transcription factors across all nine mutants. Interestingly, while 33 of the same transcription factors were implicated in group 1 and group 2 longevity mutants, 25 are modulated in different directions (activated in group 1, repressed in group 2 or vice versa) while only 5 are modulated in the same direction. This indicates that although group 1 and group 2 longevity mutants may modulate overlapping pathways to achieve long lifespan, in most cases these pathways are modulated in opposite directions. These new analyses are included in Figure S31 and Table S5. Please see lines 194-208, Figure S31, and Table S5.  

      (3) osm-5 and daf-2 are categorized into two different groups in this study. Since the longevity of cilia (-) mutants is through daf-16, the same master TF driving daf-2 longevity, please perform further analyses or discussion to clarify this issue.

      Loss of daf-16 is generally detrimental to lifespan. Disruption of daf-16 decreases the lifespan of all nine long-lived mutants that we examined (see supplemental table in our review paper PMID:37127095). However, loss of daf-16 also decreases wild-type lifespan. Thus, without further evidence it is hard to distinguish between the loss of daf-16 non-specifically decreasing lifespan verse activation of DAF-16 actually contributing to lifespan extension. In daf-2 mutants and the long-lived mitochondrial mutants there is increased nuclear localization of DAF-16 and upregulation of DAF-16 target genes. The differentially expressed genes in the long-lived mitochondrial mutants exhibit about a 50% overlap with the differentially expressed genes in daf-2 mutants (see Author response image 1). In contrast, osm-5 mutants show upregulation of some DAF-16 upregulated genes, no change in some DAF-16 upregulated genes and downregulation of other DAF-16 upregulated genes (see Author response image 1). Only about 10% of the differentially expressed genes in osm-5 mutants overlap with differentially expressed genes in daf-2 mutants. We believe that these results are consistent with loss of DAF-16 causing a general decrease in lifespan and not specifically contributing to osm-5 longevity. These comparisons will be included in a manuscript that we are currently preparing on osm-5 mutant longevity.

      Author response image 1.

      (4) This manuscript focused on genes whose RNAi suppressed the mutants longevity. Please also use bioinformatics to analyze the functions of those whose RNAi extends the mutants longevity, because these genes could tell the health price these mutants pay and help improve ageing interventions by reducing side effects.

      We perform enrichment analysis for both genes upregulated and downregulated in the long-lived mutant strains. The downregulated genes are involved in translation, ribosome biogenesis and gene expression. For the RNAi screen, we aimed to identify genes that are contributing to longevity and so we looked for a decrease in the lifespan of long-lived mutants when treated with RNAi. We did not screen for genes that extend the long-lived mutants longevity. While we did, nonetheless, identify multiple RNAi clones that increased either daf-2 or nuo-6 lifespan, there were not enough genes to identify any patterns of enrichment.

      (5) (OPTIONAL) I strongly suggest a comprehensive comparison of these transcriptomic changes in long-lived mutants with published age-related transcriptomic changes in wild type worms.

      According to this suggestion, we have now compared the differentially expressed genes that we identified in the nine long-lived mutants with genes that were found to be differentially expressed with aging. Interestingly, the group 2 long-lived mutants show a larger overlap for genes modulated in the opposite direction as aging (genes downregulated during aging are upregulated in eat-2 and osm-5 mutants). We have added this new analysis to our manuscript. Please see lines 210-223, Figure S32 and Table S6.

      Minor comments

      (1) Please further clarify the analysis of DEGs correlated with lifespan extension in Fig. 2 by a depiction. In Fig. 2C and D, please label data dots from different strains with different colors.

      According to this suggestion, each strain has been labelled a different colour.

      (2) In Fig. 3 and S20, please label the percentage of overlapping genes on top of each bars.

      We have now labelled the percentage of overlapping genes in Figure 3 and S20 (now S27).

      Reviewer #2:

      Major comments

      (1) While the authors identified a set of 196 upregulated genes, the rationale for narrowing these down to the three final candidates (C08F11.7, ugt-62, and K05C4.9) is not clearly described. The authors show that genetic inhibition of several genes, including DC2.5, C05B5.5, T07C4.5, and W03B1.7, decreases lifespan in both nuo-6 mutants and wild-type animals. However, the authors did not describe why these additional validated candidates, which also showed significant effects on longevity, were not pursued for further

      characterization. The authors should explicitly state the criteria used to prioritize these three genes over the other validated genes.

      Due to the costs and time involved in generating and characterizing new strains, we decided that we would select three strains to study further as a proof-of-principle. When deciding which genes to study further, we considered several approaches. In the end, we chose to use the strength/reproducibility of the increase in weighted mortality to identify genes with a clear, consistent impact. C05B5.5 and T07C4.5 were ruled out because they had an inconsistent impact on weighted mortality (Figure S28). W03B1.7 was ruled out because it did not have a strong enough e ect on weighted mortality (Figure S28). That narrowed it down to C08F11.7, ugt-62, DC2.5, and K05C4.9. Of those 4, C08F11.7, ugt-62, and K05C4.9 have the greatest consistent impact on weighted mortality (Figure S28) and so these genes were chosen. We have updated the manuscript to include this justification for focussing on C08F11.7, ugt-62, and K05C4.9. Please see lines 273-278.

      (2) The authors conclude that longevity can be mediated by multiple molecular pathways. However, it remains unclear whether these distinct strategies can operate simultaneously or are mutually exclusive. The authors need to test whether lifespan extension in a Group 1 mutant is further enhanced or suppressed by the knockdown of a key Group 2-specific genes. These experiments would help determine these pathways act additively, antagonistically, or as partially redundant survival programs.

      This is an excellent suggestion. While our data identify several genes that are regulated in opposite directions in group 1 and group 2 longevity mutants, we do not yet know the extent to which each of these genes contribute to the longevity of group 1 and group 2 mutants. The three genes that we focused on for further characterization (C08F11.7, ugt-62 and K05C4.9) are upregulated in group 1 longevity mutants but not group 2 mutants. Contrary to what might be expected, RNAi knockdown of these genes does not decrease the lifespan of the group 1 longevity mutant daf-2 but does decrease the lifespan of the group 2 longevity mutant eat-2. We recently reviewed the e ect of di erent resilience pathways on the lifespan of long-lived genetic mutants. Disruption of daf-16, sek-1, skn-1, hsf-1, ire-1 and trx-1 can decrease lifespan in both group 1 and group 2 longevity mutants, but also decreases lifespan in wild-type worms suggesting that at least in some mutants the e ect on longevity may be non-specific. Disruption of hif-1 does not a ect the longevity of group 2 mutants, but does a ect the lifespan of some group 1 mutants (clk-1, isp-1, nuo-6) but not others (daf-2, glp-1). To more definitively answer the question, it would be interesting to cross different combinations of group 1 and group 2 longevity mutants to see the extent to which different longevity groups synergize. This is something we are currently working on for a separate manuscript. We have added these points to the revised manuscript. Please see lines 363-381.

      (3) The authors provide interesting data on overexpression of the three candidate genes. However, whereas C08F11.7 clearly demonstrates both necessity and sufficiency for lifespan extension, overexpression of ugt-62 and K05C4.9 does not independently extend lifespan. To strengthen the manuscript, the authors should expand the discussion of these divergent results and clarify possible explanations.

      According to this suggestion, we have expanded our discussion to discuss possibilities of why these genes might be having different effects on lifespan. Please see lines 411-423.

      (4) Key citations are missing and the authors should add multiple citations including the following ones. Please cite the following paper and discuss the authors' finding with respect to the related work (Lee et al PMID: 40814218). Add citations in the sentence describing changes in the transcriptome of C. elegans associated with age (Lee et al., PMID: 38508494). Furthermore, please cite papers describing the overviews of survival assay using C. elegans (Kwon et al., PMID: 40436148, Hwang et al., PMID: 40436147).

      We have added the suggested citations to the revised manuscript. Please see lines 211 (Ref #39), 307 (Ref #41), 423 (Ref #54) and 436 (Ref #55).

      Minor comments

      (1) To improve readability, please provide the full names for all abbreviations at their first appearance in the manuscript.

      We have added the full names for each abbreviation on first appearance.

      (2) Please ensure that the labels in the figures match the text exactly. For instance, if different promoters are used for generating overexpression animals, it may be helpful to indicate the specific promoter in the figure panel or legend for clarity.

      We have ensured that the nomenclature in the text and the figures is the same. We have noted the promoter used for the overexpression strains in the figure legend.

      (3) For all lifespan and stress resistance assays, please include the total number of animals (n) and the number of independent biological replicates (N) in the figure legends to confirm statistical reliability.

      We have added the number of animals and independent biological replicates to the figures and figure legends.

      (4) Please clearly specify the exact developmental stage of the animals used for the survival assays in the Materials and Methods section.

      We have updated the methods to describe the developmental stages used for the survival assays.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study aims to clarify MATR3's function and molecular mechanism in oocyte growth and maturation, explore its association with OMA, and its potential as a diagnostic and therapeutic target using specific knockout mouse models, human OMA samples, and multi-omics technologies. And it has fully achieved preset objectives with results strongly supporting conclusions. Specifically, it addresses the gap in the synergistic mechanism of epigenetic and secretory signals regulated by RNA-binding proteins (RBPs) in oocyte growth and enriches the molecular etiological spectrum of oocyte maturation disorders. It is the first time the conservative function of MATR3 has been revealed in multiple species, providing a paradigm for cross-species research on RBPs in the field of reproductive biology. It also provides a new candidate target for OMA, a clinically refractory infertility disease, and is expected to promote the optimization of assisted reproductive technology and the development of precision medicine.

      Strengths:

      The strengths of this study are significant and prominent. First, the research system is comprehensive, integrating knockout mouse models, in vitro knockdown models, multi-species (mouse, porcine, and human) verification, combined with scRNA-seq, LACE-seq, CO-IP, and other multi-omics and molecular biology technologies, forming a complete and progressive evidence chain. Second, the mechanism analysis is in-depth, clarifying the dual molecular mechanisms of MATR3 regulating the transcriptional synthesis and secretion of GDF9 through "recruiting KDM3B to regulate H3K9me2 demethylation" and "directly binding to Rdx mRNA", with a clear logical closed loop. Third, the clinical correlation is close. It is the first time to find abnormal nuclear localization of MATR3 in oocytes of OMA patients, providing new clues for clinical disease mechanism research, and verifying the downstream function of GDF9 through rescue experiments, effectively enhancing the translational value of the results.

      Weaknesses:

      This study included only one OMA patient's oocyte sample. Without clinical screening for MATR3 mutations or abnormal expression, establishing a causal relationship between MATR3 and OMA remains difficult.

      We greatly appreciate positive comments and constructive feedback on our manuscript.

      We are encouraged that you recognize the novelty, rigour, and clinical relevance of our study on MATR3 in oocyte development and OMA. We have carefully considered your comments and revised the manuscript accordingly. We will further expand the OMA patient cohort in future studies to verify the causal relationship between MATR3 and OMA.

      Reviewer #2 (Public review):

      Summary:

      This study investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout mouse models together with in vitro follicle culture and molecular analyses. The authors aim to determine whether MATR3 regulates oocyte maturation and follicle development and to explore potential mechanisms linking MATR3 function to transcriptional and epigenetic regulation in growing oocytes.

      Strengths:

      A major strength of the work is the use of a conditional knockout mouse model combined with complementary in vitro follicle culture approaches, which together provide a useful framework for examining gene function during oocyte development. The study also attempts to integrate cellular phenotypes with molecular analyses of transcriptional activity and epigenetic markers.

      Weaknesses:

      Several weaknesses limit the strength of the conclusions. These include insufficient validation of key experimental manipulations (such as the efficiency of MATR3 knockdown in siRNA experiments), limited quantification or statistical analysis for some datasets, inconsistencies between the text and presented data in certain figures, and incomplete methodological descriptions that make it difficult to fully evaluate reproducibility.

      We greatly appreciate your constructive comments and suggestions. We are grateful for the recognition of our conditional knockout mouse model and experimental design. We have carefully addressed all the weaknesses mentioned by the reviewer, including the validation of key experiments, quantitative and statistical analysis, consistency between text and figures, and detailed methodological descriptions. Details are described point-by-point below.

      Reviewer #3 (Public review):

      Summary:

      The study aims to elucidate the dual molecular mechanisms of the RNA-binding protein MATR3 in oocyte growth and maturation. The authors propose that MATR3, highly expressed in growing oocytes (GOs), regulates oocyte quality through two pathways: epigenetically, by recruiting KDM3B to remove the repressive H3K9me2 mark at the Gdf9 locus to activate transcription; and post-transcriptionally, by binding Rdx mRNA to maintain microvillus structure for GDF9 secretion. This mechanism ensures oocyte-granulosa cell communication and female fertility. The study also explores the link between MATR3 and human oocyte maturation arrest (OMA).

      Strengths:

      The study proposes an innovative dual-mechanism model encompassing "epigenetic transcriptional activation and cytoskeletal regulation," which not only expands the functional understanding of RNA-binding proteins in chromatin regulation but also reveals the coordination between nuclear transcription and organelle structure. By integrating scRNA-seq and LACE-seq, the authors constructed a comprehensive regulatory network for MATR3, identifying both key targets and numerous potential molecules, thereby providing rich resources for future mechanistic studies. Furthermore, the inclusion of oocyte samples from human OMA patients directly links the basic findings to clinical reproductive disorders. Despite the limited sample size, this approach demonstrates strong translational potential.

      Weaknesses:

      The partial phenotypic improvement achieved by exogenous GDF9 supplementation suggests that the downstream effector pathways may involve a more complex network regulation, implying that the current interpretation of GDF9's central role could be further explored. Regarding the developmental abnormalities of granulosa cells in the conditional knockout model, their pathological origins require in-depth analysis to determine whether they represent primary alterations or secondary adaptive responses resulting from the loss of oocyte signaling.

      We greatly appreciate your positive and insightful comments on our study. We are grateful for the recognition of our novel dual-mechanism model, comprehensive multi-omics analysis, and translational potential from basic research to clinical OMA. We have carefully addressed the weaknesses raised by the reviewer, including in-depth discussion of the GDF9-centered regulatory network and clarification of the origin of granulosa cell abnormalities. More details point-by-point responses are provided below.

      Recommendations for the authors:

      Point-by-point responses to reviewers’ comments

      We thank the reviewer very much for his/her reviewing of our work, and we appreciate the constructive comments and suggestions that have helped us to prepare an improved revision. Based on the comments of the reviewer, we have carefully revised the manuscript by performing some new experiments.

      Reviewer #1 (Recommendations for the authors):

      (1) Did most of the follicles cultured in vitro reach the antral follicle stage after 6 days?

      We greatly appreciate your insightful question. We statistically analyzed the survival rate and antral follicle ratio of in vitro-cultured follicles after 6 days of culture. Due to differences in culture systems and protocols, the follicle survival rate in our study (57.43 ± 3.11%) was different from that reported in previous literature (92 ± 10%). However, the proportion of antral follicles among surviving follicles was highly consistent between our results and published data (83 ± 13% vs 80.87 ± 3.27%) (Cortvrindt and Smitz 2002).

      Author response image 1.

      Ratio and survival rate of antral follicles after 6 days of culture. n = 3. Data are represented as mean ± SD.

      (2) In Figure 2F, at which stage did MATR3 begin to affect oocyte diameter?

      Thank you for your careful observation. Our morphological analysis of oocytes collected from PD14 and PD23 mice showed no significant difference in oocyte diameter between the cKO and Ctrl groups at the GO stage (Fig. S3D, E). However, oocytes in the cKO group became significantly smaller than those in the Ctrl group once they reached the FGO stage (Fig. 2E, F). Taken together, these results indicate that the growth defect caused by MATR3 deletion begins to manifest during the transition from the GO to FGO stage, with significant reduction in oocyte diameter clearly observed at the FGO stage as shown in Figure 2F.

      (3) What was the developmental potential of oocytes in Matr3-knockout mice?

      Thank you for this important question. Compared with the Ctrl group, oocytes derived from cKO mice showed a drastically reduced fertilization rate (91.55 ± 1.96% vs 10.55 ± 4.78%) and almost completely failed to develop to the blastocyst stage (76.62 ± 7.56% vs 3.67 ± 3.38%). These results clearly demonstrate that maternal deletion of Matr3 severely compromises the developmental potential of mouse oocytes, including fertilization capacity and subsequent early embryonic development.

      Author response image 2.

      Results of in vitro fertilization of oocytes. 2-cell: 2 days after fertilization; blastocyst: 4 days after fertilization. Data are represented as mean ± SD. ***P < 0.001.

      (4) The legend labels in the figures should not be bold.

      Thank you for your valuable suggestion. We have revised all the figures accordingly.

      Reviewer #2 (Recommendations for the authors):

      This manuscript investigates the role of MATR3 in oocyte development and folliculogenesis using conditional knockout (cKO) mouse models combined with in vitro follicle culture approaches. The topic is relevant to the field of reproductive biology and provides potentially important insights into the molecular mechanisms regulating oocyte maturation and follicle development.

      While the study presents interesting observations and utilizes both in vivo and in vitro experimental systems, several issues need to be addressed before the manuscript can meet the expected standards. These include concerns related to data interpretation, validation of experimental approaches, completeness of methodological descriptions, and clarity in data presentation. In addition, the manuscript requires substantial language editing to improve clarity and readability.

      The comments below outline major issues that should be addressed to strengthen the manuscript, as well as specific minor points regarding presentation and clarity.

      (1) The manuscript requires substantial revision to improve the written language and grammar. Numerous sentences are unclear or awkwardly phrased, which makes interpretation of the results difficult in several sections. The authors are strongly encouraged to have the manuscript professionally edited or thoroughly revised for language and clarity before making a resubmission.

      We sincerely appreciate the careful and constructive comments on the language quality and clarity of the manuscript. We fully agree that the written language, grammar, and sentence structure need substantial improvement to ensure the results are presented clearly and accurately.

      To address these concerns thoroughly, we have carefully revised the entire manuscript, including correcting grammatical errors, refining awkward phrasing, and restructuring unclear sentences to enhance readability and logical flow. In addition, we have sought professional language editing support to further polish the English expression and ensure the manuscript meets the linguistic standards of the journal.

      All revisions related to language and clarity have been completed, and we believe the revised version is significantly improved in terms of readability and precision.

      (2) Interpretation of oocyte maturation results (Line 140; Figure 2E, H). The manuscript states: "During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig. 2E, H)." However, Figure 2H appears to show that a small proportion of knockout oocytes do reach the MII stage. Therefore, the description in the text seems inconsistent with the data presented. The authors should clarify the exact maturation rates in both groups, revise the text to accurately reflect the data, and provide statistical analysis to support the stated conclusions.

      Thank you for your valuable comment. A small proportion of knockout oocytes from PD23 cKO mice can indeed develop to the MII stage. We have revised the corresponding description and supplemented the statistical analysis of maturation rates to support our conclusion.

      Line 140: “During in vitro maturation, oocytes isolated from PD23 cKO mice could not develop to metaphase II (Fig.2E, H).” have been replaced by “During in vitro maturation, the proportion of oocytes from PD23 cKO mice developing to metaphase II stage was significantly reduced (Fig.2E, H, 54.9±2.08% vs 9.57±1.11%).”

      (3) Human oocyte sample size: In Figure 1D, it is unclear how many human oocytes were analyzed. It is important to specify the sample size (n) for all experiments. The authors should clearly indicate the number of oocytes analyzed in this experiment. Provide this information either in the figure legend or in the main text.

      Thank you for this important comment. We agree that the altered subcellular localization of MATR3 in human OMA oocytes is of great physiological significance for understanding the functional role of MATR3 during oocyte development.

      Unfortunately, during a 3‑month period of sample collection, we examined MATR3 localization in immature oocytes that failed to reach the MII stage, obtained from 11 women undergoing IVF treatment. Among these samples, only one donor’s oocytes exhibited the NSN chromatin configuration. Excitingly, these NSN‑stage oocytes from this donor clearly showed the loss of MATR3 nuclear localization, which strongly supports the critical role of MATR3 during oocyte growth and maturation. We have now clearly stated the sample size (n = 11) in the figure legend and main text as suggested. In future studies, we will continue to collect more human oocyte samples to further validate these observations with an expanded sample size.

      (4) Figure annotation issue: The figure legend for Figure 1F refers to an arrow, but no arrow is visible in the figure panel. Please correct this inconsistency by either adding the appropriate arrow to the figure or revising the legend accordingly.

      Thank you for pointing out this error. We have revised the figure legend for Figure 1F accordingly to correct this inconsistency.

      (5) Description of follicle analysis (Line 146): The sentence: "This was reinforced by the data of available follicles within the follicles of mice on PD35 (Fig. 2I, J)." is incorrect or poorly phrased. It should likely read: "...available follicles within the ovaries of mice at PD35...".

      We really appreciate your constructive suggestion on the phrasing. We have revised this sentence in the revised manuscript accordingly.

      (6) Quantification of proliferating cells: Figure 2K shows Ki-positive cells, but quantitative analysis is not provided. The authors should quantify the number or proportion of Ki-positive cells in both control and cKO groups and include statistical analysis to support any claims regarding differences in proliferation.

      Thank you for your valuable suggestion. We have quantified the number of Ki‑67‑positive granulosa cells in both control and cKO groups and performed the corresponding statistical analysis.

      The quantitative results have been added to Fig. S3G, and the relevant description has been supplemented in the main text at Line 150 to support our conclusion regarding cell proliferation differences.

      Line 150: “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K) in cKO mice were lower than those found in the Ctrl.” have been replaced by “Consistently, immunofluorescence staining showed that the numbers of Ki67-positive (Fig. 2K, Fig. S3G) in cKO mice were lower than those found in the Ctrl (73.55±13.29% vs 24.65±7.80%).”

      (7) Validation of findings in the in vivo cKO model (Figure 3): The development of an in vitro follicle culture system is an interesting and valuable component of the study. However, several key analyses performed in vitro (e.g., transcription assays and analysis of epigenetic markers) should ideally also be validated in oocytes derived from the in vivo cKO model.

      Thanks for the valuable concern. We fully agree with you that the in vivo cKO model should be used to validate several key analyses performed in vitro. We collected growing oocytes from Ctrl and cKO mice and conducted transcription assays as well as analysis of epigenetic markers. The results showed that Matr3 knockout significantly downregulated transcriptional activity in GO and increased H3K9me2 levels (Author response image 3), which is consistent with our in vitro findings (Fig 3D E I J). These in vivo results confirm that MATR3 plays a critical role in regulating GO transcriptional activity and H3K9me2 levels.

      Author response image 3.

      Matr3 knockout results in the reduction of transcriptional activity. A EU staining (green) in GO collected from Ctrl and cKO. n = 15. B Quantification of the mean fluorescence intensity of EU in oocytes. C H3K9me2 staining (red) in GO collected from Ctrl and cKO. n = 15. D Quantification of the mean fluorescence intensity of H3K9me2 in oocytes. Scale bar: 20 μm. Data are represented as mean ± S.D. ***P < 0.001.

      (8) To strengthen the conclusions, the authors should consider repeating key experiments using oocytes directly isolated from the cKO mice. This would help confirm that the observed effects are not artifacts of the in vitro culture system.

      Thank you for this valuable and constructive suggestion. We fully agree that the conditional knockout mouse model is essential for verifying the physiological significance of MATR3 in vivo.

      To address this point, we have validated multiple key in vitro findings using oocytes directly isolated from cKO mice. For instances, the changes in oocyte transcriptional activity (EU staining) (in Comments 7), H3K9me2 levels (in Comments 7), GDF9 levels (Fig 4.B C E F), and OO-Mvi (Fig 6.A B, Author response image 4) all showed consistent trends with our in vitro knockdown results. In addition, the complete infertility phenotype of cKO female mice further demonstrates that MATR3 is indispensable for oocyte growth and meiotic maturation. We have also provided supplemental data from GDF9 rescue experiments and sequencing analysis performed in the mouse model.

      Author response image 4.

      Matr3 knockdown impairs the structural integrity of oocyte OO-MVi. A p-ERM staining (green) showing the OO-Mvi in oocyte from NC and si-Matr3. B Quantification of the number of Oo-Mvi vesicles (n = 6). Scale bar: 20 μm. Data are represented as mean ± SD. ***P < 0.001.

      In conclusion, the core conclusions of this study are supported by the mutual validation of key experimental results from MATR3-specific knockdown in vitro and Matr3 conditional knockout mouse models in vivo.

      (9) (1) Validation of MATR3 knockdown: The in vitro MATR3 knockdown experiment presented in Figure 4G raises an important concern: it is unclear whether Matr3 knockdown was effectively achieved in the oocytes analyzed. The authors should provide direct evidence of knockdown efficiency, for example, immunostaining for MATR3 protein on the oocytes. Without such validation, it is difficult to interpret the functional outcomes observed.

      As requested, we have provided direct evidences of the knockdown efficiency via immunostaining, which is now presented in Fig.S2D.

      In this experiment, oocytes from early growing follicles (approximately 150 μm in diameter) were microinjected with Matr3 siRNA. Following 5 days of continuous in vitro culture, oocytes from both the NC and si-Matr3 groups were isolated and subjected to immunofluorescence staining to assess protein levels. As shown in the figure, the oocytes at this stage exhibited the characteristic non-surrounded nucleolus (NSN) chromatin configuration. We observed robust MATR3 protein expression within the nucleus of NC oocytes, whereas the MATR3 protein levels were markedly reduced in the si-Matr3 group. These results confirm the successful construction of the Matr3 knockdown model in early growing follicle oocytes.

      (9) (2) Furthermore, it would be more convincing if the authors could perform the Gdf9 supplementation experiments using follicles isolated from the cKO mice, rather than relying solely on siRNA knockdown in vitro. Such experiments would provide clearer and more physiologically relevant evidence. If these experiments were attempted but did not produce similar results, this should be discussed.

      Thank you for your valuable and insightful suggestion. We fully agree that performing GDF9 supplementation experiments using follicles isolated from cKO mice would provide more direct and physiologically relevant evidence to strengthen our conclusions.

      Unfortunately, when we attempted to conduct GDF9 rescue experiments on follicles from cKO mice, neither the control nor cKO follicles were able to develop to the antral follicle stage (n=3). We speculate that this was caused by insufficient bioactivity of the veterinary-grade FSH been used, as compared to the imported FSH been provided by NHPP. Unfortunately, this particular FSH product has been discontinued. We are currently actively seeking and attempting to purchase new, qualified FSH reagents to repeat these experiments and further validate our findings in future work.

      (10) Figure citation order: Figures are not cited sequentially in the text. For example, Figure 6J is described first (line 250), followed by Figure 6A. Figures should be discussed in logical order, typically starting from panel A. Please revise the text to ensure that figure panels are introduced sequentially.

      Thank you for this careful and important comment.

      We have carefully revised the citation order of all figure panels in the main text, especially for Figure 6, to ensure they are introduced sequentially from panel A to the last panel in logical and numerical order, rather than being cited out of sequence.

      The corresponding adjustments have been made in the revised version of the manuscript.

      (11) Figure 6J interpretation: The purpose of the images shown in this figure is unclear. The authors should provide higher magnification images to clearly visualize the Oo-Mvi structures and include quantification of the observed phenotype to support the interpretation. It is important because the main findings of the paper heavily rely on these results.

      Thank you for this valuable and constructive suggestion. We fully agree that higher‑magnification images and quantitative analysis are essential to clearly demonstrate the Oo‑Mvi structures and reliably support our conclusions, especially given the importance of these results to the main findings of this study.

      Accordingly, we have replaced the original panels in Figure 6A with higher‑magnification images to better visualize Oo‑Mvi structures. In addition, we have supplemented the corresponding quantitative analysis of the observed phenotype to strengthen the interpretation of this figure (Fig 6B). All revisions have been incorporated into the revised manuscript.

      (12) Incomplete Materials and Methods section: The Materials and Methods section lacks important experimental details required for reproducibility. Specifically, the Matr3 flox mouse model. Either provide the appropriate reference describing the Matr3 floxed mice or include details on how the floxed allele was generated.

      Thank you for your valuable and careful comment. We fully agree that detailed experimental information in the Materials and Methods section is crucial for ensuring the reproducibility of the study. And we apologize for the omission of key details regarding the Matr3 flox mouse model.

      In response to your suggestion, we have thoroughly supplemented the relevant experimental details in the Materials and Methods section of the revised manuscript, including the specific construction strategy of the Matr3 floxed allele. These detailed descriptions will enable other researchers to reproduce our mouse model and verify the experimental results.

      All supplementary information has been integrated into the revised manuscript to meet the requirements of experimental reproducibility. We greatly appreciate your guidance in helping us improve the completeness and rigour of our study.

      (13) Follicle isolation: The manuscript does not describe how growing follicles were isolated. Please specify whether follicles were isolated using enzymatic digestion or mechanical dissection and provide sufficient methodological detail so that other researchers can reproduce the experiments.

      Thank you for this valuable comment. We agree that adding this information is essential for ensuring the reproducibility of our experiments. We have supplemented the corresponding description in the Materials and Methods section. Briefly, growing follicles were isolated by mechanical dissection using insulin syringes under a stereomicroscope, without any enzymatic digestion.

      Reviewer #3 (Recommendations for the authors):

      (1) Since KDM3B and MATR3 interact in cell lines, does this relationship affect the functional localization of KDM3B within oocytes? Specifically, does the localization of KDM3B change in cKO mice (e.g., nuclear export or aggregation)?

      Thank you for this insightful and constructive question. To address whether the interaction between KDM3B and MATR3 influences the functional localization of KDM3B in oocytes, we performed immunofluorescence staining to examine the subcellular distribution of KDM3B in cKO oocytes. Our results demonstrated that the nuclear localization of KDM3B remained unaltered; no obvious nuclear export or abnormal aggregation was observed in MATR3-deficient oocytes (Author response image 5).

      Based on these observations combined with our other experimental data, we propose that MATR3 regulates oocyte transcriptional activity through its physical interaction with KDM3B, rather than by controlling the nuclear targeting of KDM3B. Notably, despite unchanged nuclear localization of KDM3B in MATR3 cKO oocytes, we detected significantly elevated global levels of H3K9me2 and markedly reduced transcriptional activity (in Comments 7). These findings indicate that KDM3B loses its physiological function of demethylating H3K9me2 and promoting transcription in the absence of MATR3.

      In line with this mechanism, previous results showed that KDM3B knockout in female mice leads to follicle arrest at the secondary follicle stage and consequent infertility (Liu et al. 2015). Collectively, we conclude that in MATR3 cKO oocytes, although KDM3B is properly localized in the nucleus, it fails to execute its H3K9me2 demethylase activity, thereby impairing normal transcriptional regulation during oocyte development.

      Author response image 5.

      Matr3 knockout has no effect on KDM3B localization. KDM3B staining (red) in GO collected from Ctrl and cKO. n   = 50. Scale bar: 40 μm.

      (2) As the GDF9 rescue experiment only partially restores the phenotype, it is suggested to select 2-3 novel targets with high binding intensity and significant expression changes from the LACE-seq data (e.g., Igf2bp2 or Ccnb1 mentioned in the text) to further illustrate that MATR3 regulates a network.

      Thanks for this meaningful and constructive suggestion.

      We have supplemented the relevant data in Fig. S9 and further elaborated on these findings in the Discussion section in our revised manuscript, following your advice. Briefly, we collected growing oocytes from Ctrl and cKO mice and performed RT‑qPCR analysis to verify the expression of Igf2bp2 and Ccnb1 - two representative novel targets with strong binding intensity and significant expression changes identified from our LACE‑seq bioinformatics analysis. The results showed that both genes were significantly downregulated in cKO oocytes compared with controls, supporting the notion that MATR3 regulates a functional RNA network during oocyte development.

      (3) Are the granulosa cell defects primary or secondary? It is recommended to collect ovaries from earlier-stage cKO mice (e.g., PD7 or PD10) to examine the levels of FOXL2 and PCNA in granulosa cells.

      Thank you for your valuable comment. To clarify whether the granulosa cell defects are primary or secondary, we further investigated the temporal effect of MATR3 deficiency on granulosa cells by collecting ovaries from PD7, which is a critical period for primordial follicle activation and early follicular development.

      To evaluate the status of granulosa cells, we performed immunofluorescence staining on PD7 ovarian sections using FOXL2 (a specific marker for granulosa cells) to quantify the number of granulosa cells, and Ki67 (a proliferation-related marker) to assess the proliferative capacity of granulosa cells. The results, as shown in Author response image 6, demonstrated that there were no significant differences in either the number of granulosa cells or their proliferation levels in primary follicles between cKO and Ctrl.

      These findings are consistent with the data in our supplementary Fig S3F, where we observed no significant differences in the number of primordial follicles and growing follicles at PD7 between the two groups. Collectively, these results indicate that the activation of primordial follicles is not affected by MATR3 deficiency in oocytes, and the impairment of granulosa cells caused by oocyte-specific Matr3 knockout occurs at the secondary follicle stage rather than the early stages of follicular development. Therefore, we conclude that the granulosa cell defects in cKO mice are secondary to the oocyte dysfunction induced by MATR3 deficiency.

      Author response image 6.

      Loss of MATR3 in oocytes does not affect the number and proliferation of granulosa cells in primordial follicles. A Immunohistochemistry results showing granulosa cells in PD7 ovaries from Ctrl and cKO. B Quantification of granulosa cell number in the largest cross-section of primary follicles. C Quantification of the proliferation rate of granulosa cells in the largest cross-section of primary follicles. n = 15. Data are represented as mean ± SD. n.s., not significant.

      References:

      (1) Cortvrindt RG, Smitz JE. 2002. Follicle culture in reproductive toxicology: a tool for in-vitro testing of ovarian function? Human reproduction update 8: 243-254.

      (2) Liu Z, Chen X, Zhou S, Liao L, Jiang R, Xu J. 2015. The histone H3K9 demethylase Kdm3b is required for somatic growth and female reproductive function. International journal of biological sciences 11: 494-507.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study identifies three redundant pathways-glycine cleavage system (GCS), serine hydroxymethyltransferase (GlyA), and formate-tetrahydrofolate ligase/FolD-that feed the one-carbon tetrahydrofolate (1C-THF) pool essential for Listeria monocytogenes growth and virulence. Reactivation of the normally inactive fhs gene rescues 1C-THF deficiency, revealing metabolic plasticity and vulnerability for potential antimicrobial targeting

      Strengths:

      (1) Novel evolutionary insight - reversible reactivation of a pseudogene (fhs) shows adaptive metabolic plasticity, relevant for pathogen evolution.

      (2) They systematically combine targeted gene deletions with suppressor screening to dissect the folate/one-carbon network (GCS, GlyA, Fhs/FolD).

      Weaknesses:

      (1) The study infers 1C-THF depletion mostly genetically and indirectly (growth rescue with adenine) without direct quantification of folate intermediates or fluxes. Biochemical confirmation, LC-MS-based metabolomics of folates/1C donors, or isotopic tracing would strengthen mechanistic claims.

      We agree with the reviewer that quantification of 1C-THF intermediates would strengthen our conclusions. However, the chemical methodologies to extract folates from L. monocytogenes are not established in our lab. Moreover, quantification of C1-substituted folates requires comprehensive biochemical and analytical expertise that we also do not have and which we cannot cover though co-operations. However, to further strengthen our arguments, we have introduced an experiment in the updated manuscript that demonstrates synthetic lethality of a ΔgcvPAB ΔglyA mutant with a deletion of fold (Fig. 7A). This gene encodes 5,10-methylene-tetrahydrofolate dehydrogenase/ 5,10-methylene-tetrahydrofolate cyclohydrolase, which is the third enzyme involved in N5, N10-methylene-THF generation next to GlyA and GcvPBA. Synthetic lethality of a ΔgcvPAB ΔglyA double mutant with a fold deletion is best explained by GcvPAB and GlyA also feeding the N5,N10-methylene-THF pool.

      (2) In multiple result sections, the authors report data from technical triplicates but do not mention independent biological replicates (e.g., Figure 2C, Figure 4A-B, Figure 6D). In addition, some results mention statistical significance but without a detailed description of the specific statistical tests used or replicates, such as Figure 2A-C, Figure 2E, and Figure 2G-I.

      We thank the reviewer for this helpful comment. Experiments were usually repeated three independent times, with each repetition including three technical replicates. Mean values and standard deviations were usually calculated from the technical replicates of a representative run. Statistical significance was calculated using t-tests for pairwise comparisons or t-tests using Bonferroni-Holm correction for multiple comparisons. We made sure that this is explicitly explained for each experiment in the figure legends. Wherever other calculations were used, we also clarified this in the legends.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Freier et al examines the impact of deletion of the glycine cleavage system (GCS) GcvPAB enzyme complex in the facultative intracellular bacterial pathogen Listeria monocytogenes. GcvPAB mediates the oxidative decarboxylation of glycine as a first step in a pathway that leads to the generation of N5, N10-methylene-Tetrahydrofolate (THF) to replenish the 1-carbon THF (1C-THF) pool. 1C-THF species are important for the biosynthesis of purines and pyrimidines as well as for the formation of serine, methionine, and N-formylmethionine, and the authors have previously demonstrated that gcvPAB is important for bacterial replication within macrophages. A significant defect for growth is observed for the gcvPAB deletion mutant in defined media, and this growth defect appears to stem from the sensitivity of the mutant strain to excess glycine, which is hypothesized to further deplete the 1C-THF pool. Selection of suppressor mutations that restored growth of gcvPAB deletion mutants in synthetic media with high glycine yielded mutants that reversed stop codon inactivation of the formatetetrahydrofolate ligase (fhs) gene, supporting the premise that generation of N10-formyl-THF can restore growth. Mutations within the folk, codY, and glyA genes, encoding serine hydroxymethyltransferase, were also identified, although the functional impact of these mutations is somewhat less clear. Overall, the authors report that their work identifies three pathways that feed the 1C-THF pool to support the growth and virulence of L. monocytogenes and that this work represents the first example of the spontaneous reactivation of a L. monocytogenes gene that is inactivated by a premature stop codon.

      Strengths:

      This is an interesting study that takes advantage of a naturally existing fhs mutant Listeria strain to reveal the contributions of different pathways leading to 1C-THF synthesis. The defects observed for the gcvPAB mutant in terms of intracellular growth and virulence are somewhat subtle, indicating that bacteria must be able to access host sources (such as adenine?) to compensate for the loss of purine and fMet synthesis. Overall, the authors do a nice job of assessing the importance of the pathways identified for 1C-THF synthesis.

      Weaknesses:

      (1) Line 114 and Figure 1: The authors indicate that the gcvPAB deletion forms significantly fewer plaques in addition to forming smaller plaques (although this is a bit hard to see in the plaque images). A reduction in the overall number of plaques sounds like a bacterial invasion defect - has this been carefully assessed? The smaller plaque size makes sense with reduced bacterial replication, but I'm not sure I understand the reduction in plaque number.

      The observation that the ΔgcvPAB mutant forms fewer plaques was not our claim, and we have already addressed the possibility of an invasion defect by quantifying bacterial numbers during infection of 3T3 cells. As shown in Fig. 2A, the ΔgcvPAB mutant invades 3T3 cells (the same cells used in the plaque formation assays) as efficiently as the wild type but exhibits reduced intracellular growth. Therefore, the plaquing defect is not due to impaired invasion. Furthermore, we also have analyzed the intracellular dissemination of the ΔgcvPAB mutant in 3T3 fibroblasts compared to a ΔactA mutant by microscopy. This shows that the ΔgcvPAB is evenly distributed throughout the infected host cells as the wild type and unlike the ΔactA mutant, which forms clusters (Fig. S1). Both experiments indicate that the reduced plaque area results from impaired intracellular growth rather than defects in invasion or cell-to-cell spread. The apparent reduction in plaque numbers in ΔgcvPAB-infected 3T3 cells is likely due to a strong reduction in plaque size, with only the largest plaques remaining visible. We have rephrased the relevant section to clarify this point and avoid any potential confusion:

      “In agreement with our previous results, only small plaques were formed in 3T3 cells upon infection with the ΔgcvPAB mutant (plaque area: 15±19% of wild-type level) and small plaques were also formed by the complemented strain in the absence of IPTG (50±19%).”

      (2) Do other Listeria strains contain the stop codon in fhs? How common is this mutation? That would be interesting to know.

      We determined the frequency of inactivated fhs genes among 30,000 publicly available L. monocytogenes genomes. The analysis identified only 10 isolates carrying truncated fhs alleles. These isolates fell into two groups: (i) EGD-e and its descendants, and (ii) a cluster of five ST2 food isolates. These findings have been added as a separate results section.

      (3) Based on the observation that fhs+ ΔgcvPAB ΔglyA mutant is only possible to isolate in complex media, and fhs is responsible for converting formate to 1C-THF with the addition of FolD, have the authors thought of supplementing synthetic media with formate and assessing mutant growth?

      No, we did not test formate supplementation. However, we included additional experiments testing adenine and thymine supplementation (Fig. 6E and 7B). These results show that purine and thymine become limiting in mutants lacking 1C-THF synthesizing pathways.

      Reviewer #3 (Public review):

      Summary:

      In this study, Freier et al. demonstrate that 3 distinct metabolic pathways are critical for the synthesis of 1C-THF, a metabolite that is crucial for the growth and virulence of Listeria monocytogenes. Using an elegant suppressor screen, they also demonstrate the hierarchical importance of these metabolic pathways with respect to the biosynthesis of 1C-THF.

      Strengths:

      This study uses elegant bacterial genetics to confirm that 3 distinct metabolic pathways are critical for 1CTHF synthesis in L. monocytogenes, and the lack of either one of these pathways compromises bacterial growth and virulence. The study uses a combination of in vitro growth assays, macrophage-CFU assays, and murine infection models to demonstrate this.

      Weaknesses:

      (1) The primary finding of the study is that the perturbation of any of the 3 metabolic pathways important for the synthesis of 1C-THF results in reduced growth and virulence of L. monocytogenes. However, there is no evidence demonstrating the levels of 1C-THF in the various knockouts and suppressor mutants used in this study. It is important to measure the levels of this metabolite (ideally using mass spectrometry) in the various knockouts and suppressor mutants, to provide strong causality.

      As already outlined above, we do not have any experimental possibilities to measure 1C-substituted THF in L. monocytogenes extracts directly. However, to provide additional evidence for our interpretation that “Three pathways feed the 1C-THF pool…”, we included additional genetic experiments.

      The first experiment demonstrates that the growth defect of the fhs- ΔglyA igcvPAB strain in synthetic medium lacking IPTG—which reflects the synthetic lethality of fhs with glyA and gcvPAB—can be rescued by the addition of adenine (Fig. 6E). This indicates that the fhs/fold pathway, GlyA, and the glycine cleavage system are essential due to their combined contribution to purine biosynthesis. This result confirms the metabolic model presented in Fig. 1A and thus supports the hypothesis that all three pathways contribute to 1C-THF biosynthesis.

      The second experiment additionally demonstrates synthetic lethality of gcvPAB and glyA with the fold gene. FolD acts downstream of Fhs and is one of the three enzymes synthesizing N5, N10methylene-THF shown in Fig. 1A. The fold gene is essential in EGD-e (PMID: 36114002), most likely explained by fhs inactivation. However, we were able to delete fold in an EGD-e background carrying a reconstituted fhs gene and the resulting fhs<sup>+</sup> Δfold strain was as viable as a fhs<sup>+</sup> ΔglyA ΔgcvPAB strain (Fig. 7A). However, a fhs<sup>+</sup> Δfold ΔglyA igcvPAB strain required IPTG for growth in BHI medium (Fig. 7A), indicating that the simultaneous deletion of glyA and gcvPAB is not tolerated in the absence of fold, similar to what is observed in the absence of fhs. Notably, this growth defect was not rescued by adenine supplementation (Fig. 7B), but was alleviated by thymine addition, which is also consistent with the metabolic model shown in Fig. 1A.

      Even though we are unable to directly demonstrate reduced 1C-THF levels, we hope that these two genetic approaches together with the revised title and heading of the relevant paragraph will, in the reviewers’ eyes, support our hypotheses.

      (2) The story becomes a little hard to follow since macrophage-CFU assays and murine infection model data precede the in vitro growth assays. The manuscript would benefit from a reorganization of Figures 2,3, and 4 for better readability and flow of data.

      We respectfully disagree with the reviewer. The attenuation of the ΔgcvPAB mutant in macrophages and fibroblasts was the primary motivation for further investigating its phenotype. Therefore, we chose to begin the manuscript with results from various virulence studies before presenting the sections that provide mechanistic explanations. In our view, this sequence represents a more logical and coherent way to present the data.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Synthetic medium assumptions: LSM "mimics" intracellular limitation but isn't chemically validated against host cytosolic composition. Nutrient availability conclusions could be biased.

      This is correct. We added this information:

      “LSM broth is a chemically defined medium that contains all components required for growth at defined concentrations, but it has not been chemically validated to reflect host cytosolic conditions…”

      (2) Glycine toxicity: The paper introduces the concept of glycine toxicity in ΔgcvPAB mutants. However, the conditions under which glycine becomes toxic could be further elucidated. Why glycine causes toxicity despite being an essential metabolite in other contexts requires a more in-depth mechanistic explanation.

      The concept of glycine toxicity in GCS mutants has been described previously by other researchers. We have added further details to better explain glycine toxicity and how it accounts for the growth phenotype of the ΔgcvPAB mutant:

      “If glycine cannot be catabolized (and 1C-THF cannot be generated) by the GCS due to deletion of gcvPAB, glycine might be re-routed to the serine hydroxymethyl transferase GlyA for serine formation, even though this would consume 1C-THF and therefore even further deplete the cell for 1C-THF.” and

      “In the complete absence of glycine, growth of the ΔgcvPAB mutant was largely unaffected, presumably because glycine cannot be converted to serine by GlyA anymore, thereby conserving the 1C-THF pool.”

      (3) Figures:

      (a) Figure 1B, scale bar?

      A scale bar was added.

      (b) Figure 1C, the standard error bars for igcvPAB (both with and without IPTG) are relatively wide, indicating high variability in the data. This suggests that the results in the igcvPAB group are not as consistent as the wild-type (wt) or ΔgcvPAB groups. Please show the original data and perform a statistical test (e.g., t-test or ANOVA).

      The original data have been added to Fig. 1B. t-test results (with Bonferroni-Holm correction) are now included for all samples.

      (c) Figure 2D, scale bar?

      These are sections of agar plates. From our point of view, a scale bar does not add relevant information.

      (d) Figure 3, please check the labels of the Figure 3 legend. (D) and (E) or A-B?

      Thanks, corrected.

      (e) Figure 5E, quantification of plaque areas?

      The plaque areas were quantified. A blot showing these quantitative data is now presented in Fig. 5F.

      Reviewer #2 (Recommendations for the authors):

      Line 58: There are published studies that indicate that syncytiotrophoblasts are actually resistant to Listeria infection and that it is extravillous trophoblasts that are likely to serve as entry points for Listeria into the placenta (see, for example, Lowe et al, Infect Immun. 2018 Volume 86 Issue 6 e00801-17).

      We thank the reviewer for this comment and have removed our statement claiming that syncytiothrophoblasts are the entry point as this is not relevant to the understanding of the work presented here.

      Reviewer #3 (Recommendations for the authors):

      (1) Please mention the number of times experiments were performed as independent biological replicates, wherever applicable.

      We added this information to the figure legends wherever this was necessary.

      (2) Please provide the details of the type of statistical analysis used for the various graphs, either in the figure legends or in the materials & methods section.

      This information was also added to the figure legends wherever it still was missing.

      (3) Can the authors comment on how the weight-loss phenotype of animals and the variation in the size of the spleen between animals infected with wild-type and mutants in Figure 2 can be explained without any significant changes in the CFU? Additionally, I did not see details regarding the number of animals used in the murine infection model experiments. Please mention this along with the type of statistical analyses used.

      We do not see a contradiction here, as the apparent differences are explained by the distinct time points at which CFU numbers (day 3 post-infection) and spleen size (day 9 postinfection) were measured. Starting from day 8, the difference in weight between animals infected with the wild-type strain and those infected with the ΔgcvPAB mutant becomes clear for the first time. At day 3, no significant differences are detected in either CFU numbers or weight. By day 9, when the weight difference has become apparent, differences in spleen size are also observed. To improve clarity, the time points at which each analysis was performed have been added to Fig. 2G and Fig. 2I. The number of infected animals and the type of statistical analysis used are now specified in the figure legend.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      A well-designed and preregistered simulation study investigating whether replication-success metrics can be applied to assess animal-to-human translation. The study is comprehensive, uses realistic parameter settings, and provides valuable insights into how different metrics behave under varied conditions.

      Strengths:

      (1) Methodologically rigorous and transparently preregistered.

      (2) Comprehensive simulation design covering a wide range of plausible scenarios.

      (3) Clear description of metrics and decision rules.

      (4) Valuable contribution to understanding the limitations of applying replication metrics to translation questions.

      Weaknesses:

      (1) The conceptual distinction between replication and translation could be more clearly emphasized.

      (2) Interpretation of results is dense and can be challenging to follow without a clear and summarized.

      (3) Some simulation parameters (effect sizes, heterogeneity, and number of animal studies) require more substantial justification.

      (4) Practical recommendations could be more explicit to guide applied researchers.

      We thank Reviewer 1 for the general positive assessment of our study and for the constructive feedback. We have addressed all of the four identified weaknesses in the revised manuscript. Specifically,

      (1) Conceptual distinction between replication and translation. We have reinforced this distinction at multiple points in the manuscript: in Section 2.7 (just before introducing the translation success metrics), in the Discussion, and in a new working definition of translation success added to the Introduction. We further explicitly acknowledge that statistical translation success, as defined here, is narrower than biological translation.

      (2) The dense result section. We have added a summary Table (Table 3) at the end of the Results section that compares all metrics on key properties (strengths and weaknesses, overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch). We also direct readers to this table early in Section 3.2, so that readers less interested in the technical details can obtain the key take-home messages without reading the full section.

      (3) Further justification of simulation parameters. We have substantially extended the rationale for our parameter choices in Section 2.4 and the Limitations section. We explain that our parameters are grounded in an empirical meta-analytic dataset, contextualise the large effect size and heterogeneity value against published benchmarks from preclinical research, and clarify that our main goal was to explore directional trends rather than absolute performance under specific values. We have also added an invitation for others to explore alternative parameter spaces using our openly available code.

      (4) Practical recommendations. We have extended the Recommendations section (pages 21–22) with more explicit scenario-specific guidance, supported by the new summary table.

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to address the issue of high rates of translation failure from animal studies to humans in the literature, where promising results in animal studies fail when conducting human clinical trials. Using parameters from a previous meta-analysis on prenatal amino acid supplementation and the effects it has on maternal blood pressure, the authors assessed the performance of the metrics used and whether they can quantify translation success. Performing a simulation study, the authors compared nine translation success metrics and found that no one method was uniformly optimal. The authors list several limitations of the study, such as comparability of effect sizes between animal and human studies, different goals of animal studies versus human studies, and the focus of the study on one aspect (statistics of translation) is part of a broader, more complex decision-making process before proceeding to human trials. The authors recommend using multiple metrics in combination while taking into consideration their strengths and weaknesses to assess the translation of animal studies to human outcomes. The paper achieves the aim of providing a model with several metrics to evaluate translation success from animal studies to humans.

      Strengths:

      (1) Utilizing 9 different translation success metrics in combination provides strong flexibility in evaluating whether results in animal studies can translate to humans. This would allow researchers to evaluate translation success using multiple different metrics according to the context of the study.

      (2) The authors accommodate for the limited sample size in animal studies, which are typically underpowered, and also caution that special attention should be given to heterogeneity when interpreting translation results.

      (3) Overall, this approach has the potential to be applied to other biomedical studies, provided the limitations for each of the metrics are considered. It would provide a useful tool in assessing translation from animals to humans, in addition to other factors such as safety, pharmacokinetics, etc.

      Weaknesses:

      While the study has several strengths, there are some limitations.

      (1) Preclinical animal study sizes tend to be much smaller than human studies, which results in underpowered results. The authors adjusted for this by pooling animal study data. However, high heterogeneity in the animal studies can affect translation results.

      (2) The study focuses only on evaluating the statistical component of translation, which is only one aspect of the decision-making process to move on to human trials. The study does not take into account safety and toxicological profiles, pharmacokinetics, or genetics, which are important considerations that influence the overall effect in humans.

      We thank Reviewer 2 for the thoughtful summary and for recognising the strengths of our study. We believe that both weaknesses were addressed in the revised version of our manuscript. Specifically,

      (1) Heterogeneity in animal studies. We agree that high heterogeneity in animal studies is an important limitation, and we address it directly in our simulation design by including a wide range of heterogeneity values (including very high levels, as observed in animal studies). Our results show clearly how heterogeneity affects the performance of each metric, and we highlight this in both the new summary Table (Table 3) and the Recommendations section which was extended. We also caution applied researchers to pay special attention to heterogeneity when interpreting translation results.

      (2) Focus on the statistical component of translation. We fully agree that statistical translation success is only one aspect of a broader decision-making process. We have elaborated on this in the revised manuscript, both in a new working definition of translation success in the Introduction (which explicitly distinguishes statistical from biological translation) and in a new paragraph in the Discussion section where we situate our metrics within translational decision-making frameworks such as PATH. They make it clear that progression to human trials depends on a suite of evidence of which statistical translation is only one part.

      Reviewer #3 (Public review):

      Summary:

      This paper focused on how to navigate the complex decision-making process of whether to go into human trials. This is a critical topic considering the well-documented challenges in replicating and translating findings. While these are two distinct topics (i.e., replication and translation), they are related, and the authors simulated many conditions to assess the utility of replication assessment metrics.

      Strengths:

      A major strength of the study is the detailed approach to identifying relevant conditions and metrics, and to providing rich results that outline the strengths and weaknesses of each metric. Any simulation study is challenged by trying to identify the most relevant variables of interest, and this study provided sound justification for its chosen variables of interest. While this study does not make a strong recommendation (which I see as a strength), it does provide a comprehensive overview of the various metrics and conditions that were investigated.

      Weaknesses:

      The weaknesses of the study are the limited focus on specific metrics, the assumptions, particularly in the limited number of human study variables, and the less-than-ideal approachable summary of findings for a non-technical audience.

      Conclusion:

      This paper provides a much-needed investigation and discussion of how decisions are made when assessing whether to go into human trials. This is an important topic that productively challenges the status quo, considering documented challenges in replication and translation in biomedical research.

      We thank Reviewer 3 for the positive assessment and for the constructive suggestions.

      We have addressed the identified weaknesses as follows:

      (1) The assumptions around human study variables. We acknowledge these as inherent constraints of the simulation design. We have added a note in the Limitations section about the fixed human sample size (N = 107 per group), clarifying that while this value is grounded in a power analysis as per regulatory standards, it represents one particular scenario and may not generalise to all contexts. Further, we have contextualised and motivated the other simulation parameters better. We also invite readers to explore alternative conditions using our openly available code.

      (2) Approachability of the summary of findings for a non-technical audience. We have added a summary Table (Table 3) at the end of the Results section, comparing the metrics on key properties including overall type 1 error control, sensitivity to heterogeneity, dependence on animal sample size and number of studies, and behaviour under effect mismatch. We direct readers to this table early in Section 3.2 so that those less interested in the technical details can obtain the main take-home messages without reading the full section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major points:

      (1) Conceptual framing: clearer distinction between replication vs translation

      The Introduction correctly points out the conceptual difference between replication and translation (animal to human), but this distinction needs to be reinforced repeatedly, especially when interpreting metric performance. For instance, several metrics (e.g., meta-analysis, replication BF) inherently assume exchangeability of findings, which is rarely justified in translation because species differ biologically.

      The manuscript should explicitly state why treating animal findings as "original studies" and human findings as "replications" can be misleading. Add a subsection in the Discussion: Why replication metrics behave differently in translation settings. This will help provide a more straightforward interpretation of the results beyond the numerical findings.

      Thank you for your feedback. While the purpose of our study is to assess the applicability of the replication success metrics in the translation context, we agree that the reader should be reminded that these two concepts differ and metrics’ assumptions might not always hold. We have reiterated the difference between replication and translation in Section 2.7, just before we introduce the translation success metrics (see end of page 7). We also reiterate it in the Discussion (see end of page 20). Here, we emphasise that while some of the metrics assume that both studies investigate the same effect, this is unlikely to be the case in translation, leading to some of the metric’s assumptions being violated which impacts the performance of the metrics.

      (2) Stronger justification of the simulation parameters is needed

      The simulation factors are comprehensively presented (Table 1), but certain choices appear arbitrary or oversimplified.

      - Effect sizes: The three levels (0, −4.44, −24.37) are derived from the motivating dataset, but the paper should explain that these represent extremely large effects in many biomedical contexts.

      - Heterogeneity values: τ<sup>2</sup> = 291.1 is enormous; adding context about real-world heterogeneity distributions would help.

      - Number of animal studies (k): Using only 2-5 studies may not reflect reality; many preclinical fields have >30 studies before clinical translation.

      These choices should be more explicitly defended in Section 4 (Limitations), beyond the brief mention already there. Provide a sensitivity analysis, or explain why extrapolation beyond this parameter space is reasonable.

      We agree that the choice of the parameter values might sometimes appear arbitrary. However, instead of arbitrarily choosing parameter values, we base our choice on data from a meta-analysis. This particular meta-analysis might not be representative of all of pre-clinical and clinical research, but because Terstappen included both animal and human studies investigating the same research question it was particularly well suited. They further used an outcome (maternal blood pressure) that is comparable between rats and humans, which is quite rare. We have specified this further in Section 2.4 (Motivating dataset, page 5). In the Limitations section, we acknowledge any possibly unrealistic simulation conditions again, and emphasize that our main goal was to explore trends in the metrics’ behavior as the conditions changed rather than their absolute performance under specific values. Further, the effect sizes (0, −4.44, −24.37 mmHg) span a meaningful range on the unstandardized mean difference scale for blood pressure measurements: from no effect to a modest but clinically relevant reduction to a large effect typical of animal studies. The large heterogeneity value corresponds to a relative heterogeneity of I^2 of 95.33% in the animal meta-analysis. While this appears high, it is frequently observed in preclinical research: Hooijmans et al (2022) showed that 55% of animal study meta-analyses using mean differences as effect size measure have I^2>75%. We also added a footnote reiterating the fact that such high effect sizes (on the raw mean difference scale) are indeed common in animal studies (on page 6). Regarding k, we acknowledge that pooling only 2 to 5 animal studies may not reflect common practice. However, the directional trends in type 1 error and power are clearly visible in our Figures. Larger k decreases the type 1 error of the animal studies, while the power is increased unless there is high heterogeneity between animal studies and there is only a small effect. Extending the range further is unlikely to change the conclusions. Moreover, in practice, the decision to advance to human trials considers evidence well beyond the statistical considerations we simulate. All of the above is now emphasized more explicitly in both the methods, where we have substantially extended the reasoning for choosing the simulation conditions, and the limitations section. Finally, we added an invitation to others to use our open material (i.e., code) and explore the behaviour of the metrics under other conditions (see top of page 21).

      (3) Decision criteria (strict/lenient/no criterion) need a clearer rationale

      The three continuation rules are a strength of the study, but:

      - The lenient criterion (any negative estimate is considered "beneficial") is unrealistic and should be reframed.

      - The strict criterion (p < 0.025) heavily inflates effect sizes (in Figure 1b) and may distort interpretation.

      It would be helpful to provide a table showing, for each criterion, its real-world analogue (e.g., regulatory requirement, exploratory progression, mechanistic plausibility).

      We have followed your suggestion and added a Table (Table 2) with the description of the criterion and a description of its real-world analogue. No criterion represents an important reference scenario used to evaluate metric behaviour independent of progression decisions. The strict criterion is the closest to regulatory-style evidence. It is also highly selective and therefore might induce biases (e.g., inflated effect sizes). We link lenient to an exploratory decision-making where efficacy evidence is considered in addition to other factors (e.g., safety), but not intended to represent a certain regulatory standard.

      (4) Interpretation of simulation results needs more focus

      The Results section is extremely detailed, making it challenging to identify the central take-home messages. The authors should consider adding a concise summary table comparing metrics on key properties:

      - T1E control robustness.

      - Sensitivity to heterogeneity.

      - Dependence on animal sample size.

      - Dependence on k.

      - Bias under asymmetric effects.

      Moving some nested-loop plot descriptions to the Supplement. Right now, descriptions are technically correct but cognitively heavy.

      We agree with your comment and have attempted to implement it in our summary Table 3, at the end of the results section. We also point readers early on to the Table, so that they can skip the more technical and detailed description if they want (see first paragraph section 3.2, page 12). After some trial and error, we agreed that the chosen columns are the most useful for an applied researcher to get a quick overview. Our table now summarises for each metric its main strengths and weaknesses, its behaviour with increasing heterogeneity, its sensitivity to more animal data (i.e., larger k and larger animal sample size), and its behaviour under effect mismatch (i.e., when the true effect in the animal and human study are dissimilar).

      (5) The discussion should provide explicit recommendations.

      The authors provide high-level recommendations, but the recommendations lack specific guidance. When heterogeneity is low, controlled sceptical p-value works well. When effect sizes differ: weighted Edgington is stable. The authors should avoid using replication BF when the animal effect ≠ human effect. Meta-analysis should not be used when human heterogeneity is high, because of inflated T1E.

      We agree that explicit recommendations would be helpful to the applied researcher. As mentioned in the reply to the previous comment, we have added a summary table which lists the strengths and weaknesses of each metric. We also extended the paragraph in the Recommendations section (on page 21 and 22) to give some examples of scenarios in which certain metrics would be recommended over others.

      (6) Recommendations for applied researchers

      The study is missing an explicit definition of "translation success". The manuscript implicitly defines translation success as: "Both animal and human results show a beneficial treatment effect according to metric X". But this is different from biological translation, which concerns underlying mechanisms. The authors briefly mention this conceptual challenge, but this should be elaborated, as it is central to interpretation.

      Thank you for this comment. We agree that “translation success” was not explicitly defined. We have now added a working definition in the Introduction, clarifying that, in this paper, translation success is defined statistically, and depends on the metric. We now explicitly acknowledge that this is a narrower definition than biological translation. We also elaborate on this distinction in the Discussion where we note that the appropriate metric and interpretation of translation success depends on the translation goal and that statistical translation is distinct from biological translation.

      Minor points:

      (1) The abstract could include a direct sentence on the main conclusion. For example, no metric was uniformly optimal; controlled sceptical p-value and weighted Edgington performed most consistently.

      Our abstract already included main conclusions. We added the word “However” to emphasize the sentence “no metric was uniformly optimal” a bit more.

      (2) The figures are informative, but nested loop plots are very dense. Consider providing a guided example in the figure caption explaining how to read them (as partially done in Figure 1a, but repeat for all).

      We agree that the Figures can be very overwhelming at first. We did not want to add specific helping elements as we did in Figure 1 to not make the figures even busier. The goal was to introduce the reader gently to the nested loop plots via Figure 1 before having them look at the remaining figures. We hope that with the added summary Table and the more detailed recommendations, applied researchers less interested in the statistical details will still find the information most relevant for them easily.

      (3) Methods: Section 2.4 could clearly state that effect sizes are in units of mmHg (blood pressure) from the dataset.

      Thank you for pointing this out. This has been added.

      (4) Results: This section is long; consider adding a brief summary paragraph at the end of 3.2.

      We added a summary table, allowing interested readers to skip the long section entirely.

      (5) Limitations: Add a note about publication bias in animal studies (you mention it in the Introduction, but not in Limitations). Add a statement about effect direction consistency (i.e., animal effect negative but human positive), which is not explored in the simulation grid.

      Thank you for pointing out this inconsistency. A note about publication bias in animal studies was added to the Limitations section (that this was not investigated). A note about opposite animal and human effects was added to Section 2.5 (Simulation conditions) under “Animal and human effect sizes”.

      Reviewer #2 (Recommendations for the authors):

      Animal studies are typically highly controlled, using animal models that are either outbred to provide higher genetic variability or inbred with very little genetic variability and with a specific phenotype. Additionally, many rodent models are incomplete models of the overall human phenotype and are typically used to investigate only one aspect of the condition/disease. Some of the rat animal models that the Terstappen et al. (2020) systematic review used as the simulation parameters for the study included outbred (Sprague-Dawley, Wistar) and inbred Spontaneous Hypertensive Rats (SHR), which have different mechanisms in which hypertensive onset can occur, especially if inducing preeclampsia in outbred animals. Is it feasible to reduce heterogeneity in the animal results if only outbred or only SHR are considered instead? I realize this may reduce the sample size even further.

      You raise an important point differentiating biological (rather than statistical) translation. We have added a sentence about differences between rat models and humans to the new paragraph in the Limitations section (bottom page 20 and top page 21) on the distinction between biological and statistical translation. As for reducing heterogeneity in the animal results by focusing on one type of rats, we agree focusing on one type of rats might reduce heterogeneity. We however consider this reduction to be very small (because the results of the study with SHR are actually comparable to the results with Wistar and SD rats). Therefore, rerunning the simulation would not yield results that differ in any meaningful way from those already reported and the substantial computational effort required to do so is not warranted.

      Reviewer #3 (Recommendations for the authors):

      Overall, I found this a very detailed study. However, my recommendation is to provide a more approachable overview of the results to reach a wider audience. Currently, the article is much more technical and statistically focused. I think two additions could help.

      (1) A summary table of each of the metrics and their strengths and weaknesses under the various conditions (e.g., animal and human study characteristics). Currently, this is done via text, but I think a high-level summary via a table could be a compelling way to make the simulations more approachable for a non-technical audience.

      As requested also by reviewer 1, we have added a summary table.

      (2) Contextualize the findings within the decision-making process a little more. The authors have a well-written limitations section that acknowledges this; however, I think the discussion (and maybe the introduction) could be enriched by putting the simulation findings into context. For example, this paper suggests a framework that includes replication as part of the decision-making process for human trials (https://www.cell.com/med/fulltext/S2666-6340(24)00296-4).

      We agree that situating our metrics within existing translational decision-making frameworks adds important context. We have added a paragraph in the Discussion (before the Limitations section on page 21) clarifying that the metrics evaluated here should not be viewed as standalone decision rules for progression from animal studies to human trials. Several frameworks have recently emerged precisely to guide such decisions in a more structured, multidimensional way. We refer to PATH and also to the GALENOS approach [DOI: 10.1186/s12874-026-02891-4]. Within such frameworks, translation success metrics of the kind evaluated here may provide a quantitative assessment of the consistency between animal and human efficacy findings, thereby informing one component of a broader translational evidence assessment. We have also briefly mentioned at the end of the Introduction (page 4) that frameworks for structuring the use of preclinical evidence in translational decisions are being developed, further motivating the need for quantitative tools such as those evaluated here.

      Below are some additional minor comments for the authors to consider:

      (1) In the abstract (4th line), there is an extra 'l' in failure.

      Thank you for the detailed review. We have fixed this.

      (2) I think since the study is completed, the objectives in the introduction should be past tense, not future.

      We have fixed this.

      (3) The limitations section should include the fixed human sample size. N=107 per group is grounded in the literature, but this varies widely based on the effect size of interest. Again, not material to the point of translation under simulated conditions (of which this would have increased the simulations well above the 648 already included), but given the impact this has on insights, this limits this investigation to a degree and should be acknowledged.

      Thank you for your comment. We have added a note about the human sample size to the paragraph about the simulation conditions in the Limitations section. The human sample size was computed via power analysis as per regulations, but we realize this could change depending on the effect size.

      (4) I appreciate how shrinkage was calculated. Though it is worth noting that the Reproducibility Project: Cancer Biology found much higher rates, which are similar to reports from biotech and pharma (e.g., 11% and 20-25% for Amgen and Bayer).

      We already mentioned the high rates of shrinkage in the Replication Project Cancer Biology (see page 10). We have now also emphasised that one could adapt these levels further depending on the situation.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this manuscript, Clausner and colleagues use simultaneous EEG and fMRI recordings to clarify how visual brain rhythms emerge across layers of early visual cortex. They report that gamma activity correlates positively with feature-specific fMRI signals in superficial and deep layers. By contrast, alpha activity generally correlated negatively with fMRI signals, with two higher frequencies within the alpha reflecting feature-specific fMRI signals. This feature-specific alpha code indicates an active role of alpha oscillations in visual feature coding, providing compelling evidence that the functions of alpha oscillations go beyond cortical idling or feature-unspecific suppression.

      The study is very interesting and timely. Methodologically, it is state-of-the-art. The findings on a more active role of alpha activity that goes beyond the classical idling or suppression accounts are in line with recent findings and theories. In sum, this paper makes a very nice contribution. I still have a few comments that I outline below, regarding the data visualization, some methodological aspects, and a couple of theoretical points.

      The authors put a lot of effort into the figure design. For instance, I really like Figure 1, which conveys a lot of information in a nice way. Figures 3 and 4, however, seem over engineered, and it takes a lot of time to distill the contents from them. The fact that they have a supplementary figure explaining the composition of these figures already indicates that the authors realized this is not particularly intuitive. First of all, the ordering of the conditions is not really intuitive. Second, the indication of significance through saturation does not really work; I have a hard time discerning the more and less saturated colors. And finally, the white dots do not really help either. I don't fully understand why they are placed where they are placed (e.g., in Figure 3). My suggestion would be to get rid of one of the factors (I think the voxel selection threshold could go: the authors could run with one of the stricter ones, and the rest could go into the supplement?) and then turn this into a few line plots. That would be so much easier to digest.

      We thank the reviewer for their insightful comments. Below we will address each point separately and highlight the changes made to the manuscript. In agreement with the reviewer we have recompiled Figures 4 and 5 (previously Figures 3 and 4). The new figures only present results for the 10% voxel selection threshold (with 5% and 25% moved to Supplementary Figures, see Figures S1-S9). Instead of the radially arranged layout, we opted for a more traditional figure layout, which significantly improved readability.

      (2) The division between high- and low-frequency alpha in the feature-specific signal correspondence is very interesting. I am wondering whether there is an opposite effect in the feature-unspecific signal correspondence. Would the high-frequency alpha show less of a feature-unspecific correlation with the BOLD?

      Following the reviewer’s interesting suggestion, we added the low/high frequency alpha analysis to the feature-unspecific analysis. Indeed, we have found a significant interaction between the sign of the signal change for selected voxel (positive vs. negative BOLD) and alpha sub-band (low vs high frequency alpha). An analysis of simple effects did not reveal any significant effects, however we found a trend level difference (p=0.097) between low and high-frequency alpha for the positive voxel sub-selection. This indicates a stronger negative relationship between upper alpha and the positive BOLD signal as compared to lower alpha. We interpret this result as partial evidence for a feature-related contribution of the upper alpha band. “Active” cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. The negative relationship between alpha and negative BOLD could thus be interpreted as an indirect effect, resulting from a reduced alpha decrease in cortical patches responding to non-attended receptive field locations. However, the involvement of attention-related processes remains speculative, since attention was not explicitly manipulated as part of the experiment.

      We have added Figure 4 B.

      We have also added this section to the Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub><0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      And the following section of the Discussion was extended:

      “The significant interaction between the sign of the BOLD signal deflection and upper or lower α bands (see Figure 4 B) further indicates that multiple α-related processes contribute differentially to positive or negative BOLD. "Active" cortical patches (positive BOLD) are most likely involved in the processing of visual features (irrespective of the specific feature), and additionally a more general (possibly attention-related) activation. In turn the negative BOLD signal might contain less feature-specific activation and is most likely related to attention-driven deactivation. This hypothesis receives additional support from the trend-level difference in α sub-bands for positive BOLD, indicating that lower α is less related to the active, possibly feature-related processes. The absence of this difference for negative BOLD again indicates a broader, more general process. Future experiments manipulating visual features and attention might reveal a differential upper and lower α response to attended visual features and a more general relationship between α (and possibly superficial layer cortical activity) for suppressed (unattended) receptive fields.”

      (3) In the discussion (line 330 onwards), the authors mention that low-frequency alpha is predominantly related to superficial layers, referencing Figure 4A. I have a hard time appreciating this pattern there. Can the authors provide some more information on where to look?

      We thank the reviewer for pointing out the lack of clarity of this section in the Discussion. We have now rephrased the Discussion, focusing more on the laminar difference and keeping the frequency difference to a separate paragraph. Our main argument for possibly multiple alpha-related processes are twofold: a difference in alpha frequency depending on the underlying analysis (low vs high frequency alpha) and a different layer distribution (superficial layers vs. superficial and deep layers, depending on the analysis). The respective section in the Discussion now focuses on the laminar difference only. We find a negative relationship between alpha and the BOLD signal most prominently in superficial layers (feature-unspecific contrast for the BOLD signal with negative t-values; Figure 4). In addition, we find a superficial and deep layer contribution for the feature-specific contrast (congruent - incongruent; Figure 5A). While the here presented experiment was set out to investigate feature-specific processes, the meaning of the feature-unspecific results are of speculative nature. Future experiments should target the laminar difference between feature-specific and unspecific processes with respect to alpha frequency and layer distribution directly. 

      We have modified the respective sections in the Discussion:

      “Furthermore, we observed that the relationship between the feature-specific BOLD signal and α is predominantly linked to frequencies above 11 Hz (see Figure 5A). An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies. Since individual frequency variations have been included as a random slope in the linear mixed-effects model, these effects cannot be explained by a subset of participants driving lower or upper α separately. Specifically our findings on upper α indicate that α is not exclusively linked to global signal modulations, which has been the traditional perspective [...]”

      “Not only did we find a dissociation in the frequency domain between the relationship of α and the BOLD signal, but furthermore found that the laminar activation patterns provide further evidence for potentially multiple α-related processes. The association between α and the BOLD signal was strongest in superficial layers for negative BOLD activity and feature-specific activity (see Figure 4A and 5A). However, deep layer-related α effects were limited to feature-specific processes only (see Figure 5 A Co-Inco). These findings suggest that superficial layer α reflects are broader, more general process, while deep layer α operates more narrowly, linked to the processing of the visual features themselves. Previous findings using laminar fMRI (which did not include the investigation of oscillatory activity), indicate that superficial layer activity might be more related to the modulation of attention [...]”

      (4) How did the authors deal with the signal-to-noise ratio (SNR) across layers, where the presence of larger drain veins typically increases BOLD (and thereby SNR) in superficial layers? This may explain the pattern of feature-unspecific effects in the alpha (Figure 3). Can the authors perform some type of SNR estimate (e.g., split-half reliability of voxel activations or similar) across layers to check whether SNR plays a role in this general pattern?

      We agree with the reviewer that the vascular draining effect typically leads to increased signal change in superficial layers, the effect on (t)SNR however might be less straightforward. We did not include any counteracting measures, because we were not interested in the amplitude of the signal change, but now include an estimate of tSNR (See Figure S10 in Supplementary Figures). We found that in fact the signal-to-noise ratio is higher in deep layers. Most importantly however, the tSNR layer profiles we identified do not reflect the correlation layer result patterns of the combined EEG-fMRI analysis. This indicates that our results are most likely not the result of tSNR differences. In order to confirm our tSNR pattern we have also conducted a second layer analysis based on the LAYNII toolbox, which assigns voxels between pial and white matter to distinct layers (as compared to our fraction-based approach) and found a similar profile as with our initial analysis. However, absolute tSNR values were found to be higher for our weighted layer analysis. We speculate that while functionally relevant components of the BOLD signal drain towards superficial layers, physiological noise components will drain towards superficial layers as well.

      It is furthermore worth pointing out that for the contrast (congruent - incongruent), the vascular draining effect would cancel out between the conditions. Our findings on superficial and deep layers for those contrasts can hence not be explained by vascular draining at all.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      (5) The GLM used for modelling the fMRI data included lots of regressors, and the scanning was intermittent. How much data was available in the end for sensibly estimating the baseline? This was not really clear to me from the methods (or I might have missed it). This seems relevant here, as the sign of the beta estimates plays a major role in interpreting the results here.

      This is a very important remark and we would like to apologise for the confusion. It was not clear in the manuscript that the GLM was computed on z-transformed fMRI data. We have not specifically collected any “baseline volumes”. A positive beta value would indicate that the sign of the predictor matches the sign of the BOLD signal deflection (and vice versa).

      We have added or modified the following sections in Results and Methods respectively:

      “Before the GLM was computed, the fMRI data was z-transformed across time, separately for each block and voxel.”

      “A general linear model (GLM) has been computed with predictors for each TF bin separately for all voxels in V1 that later have been sub-selected according to the respective condition. Time courses for each voxel have been z-transformed before the GLM was computed for each voxel and experimental block separately. Afterwards, each of the resulting regression coefficients (β values) were multiplied with the voxel-specific layer weights that have been obtained as described above.”

      (6) Some recent research suggests that gamma activity, much in contrast to the prevailing view of the mechanism for feedforward information propagation, relates to the feedback process (e.g., Vinck et al., 2025, TiCS). This view kind of fits with the localization of gamma to the deep layer here?

      (7) Another recent review (Stecher et al., 2025, TiNS) discusses feature-specific codes in visual alpha rhythms quite a bit, and it might be worth discussing how your results align with the results reported there.

      We would like to thank the reviewer for pointing out these papers. Yes, we believe that those could be very related to the effects reported here. At the time of writing the initial manuscript we were not aware of the mentioned publications. 

      We have now included these papers in the Discussion:

      “Recent publications on the information exchange within and between primary visual cortex areas of macaques also reported deep layer γ band activity depending on the stimulus material (Gieselmann et al., 2022; Ferro et al., 2021). Those publications challenge the feed-forward exclusivity of γ altogether by revealing intra-area feedback communication in V1 from layer 5 to layer 6 and layer 6 to supra-granular layers. Possibly, the relationship between γ and deep layer BOLD we observed is also related to similar processes (Vinck et al., 2025).”

      “Similarly, in a recent opinion article, Stecher et al. (2025) promote the idea of "content-aware" α-oscillations. In agreement with our results, the authors argue that α-oscillations are related to content-specific feedback signals, reflected in increased decoding performance based on α power of top-down related processes, even prior to the onset of the stimulus (Hetenyi et al., 2025).. Accordingly, we interpret the lower α effect [...]”

      Reviewer #2 (Public review):

      The authors address a long-standing controversy regarding the functional role of neural oscillations in cortical computations and layer-specific signalling. Several studies have implicated gamma oscillations in bottom-up processing, while lower-frequency oscillations have been associated with top-down signalling. Therefore, the question the authors investigate is both timely and theoretically relevant, contributing to our understanding of feedforward and feedback communication in the brain. This paper presents a novel and complicated data acquisition technique, the application of simultaneous EEG and fMRI, to benefit from both temporal and spatial resolution. A sophisticated data analysis method was executed in order to understand the underlying neural activity during a visual oddball task. Figures are well-designed and appropriately represent the results, which seem to support the overall conclusions. However, some of the claims (particularly those regarding the contribution of gamma oscillations) feel somewhat overstated, as the results offer indeed some significant evidence, but most seem more like a suggestive trend. Nonetheless, the paper is well-written, addresses a relevant and timely research question, introduces a novel and elegant analysis approach, and presents interesting findings. Further investigation will be important to strengthen and expand upon these insights.

      One of the main strengths of the paper lies in the use of a well-established and straightforward experimental paradigm (the visual oddball task). As a result, the behavioural effects reported were largely expected and reassuring to see replicated. The acquisition technique used is very novel, and while this may introduce challenges for data analysis, the authors appear to have addressed these appropriately.

      Later findings are very interesting, and mainly in line with our current understanding of feedback and feedforward signalling. However, the layer weight calculation is lacking in the manuscript. While it is discussed in the methods, it would help to briefly explain in the results how these weights are calculated, so that the reader can better follow what is being interpreted.

      Line 104 states there is one virtual channel per hemisphere for low and high frequencies. It may be helpful to include the number of channels (n=4) in the results section, as specified in the methods. Also, this raises the question of whether a single virtual channel (i.e., voxel) provides sufficient information for reproducibility.

      We thank the reviewer for encouraging us to clarify the virtual channel selection and we agree that the current description could be misleading. Indeed, we selected 4 virtual channels in total: 1 for each frequency band (alpha/gamma), for each hemisphere separately. The main goal of this selection was to find the clearest response of that frequency band to the task. Previous publications used a supervised (ICA-based) approach to extract those responses. To increase reproducibility, we have chosen an unsupervised beamformer-based approach. The reconstruction of time or frequency-resolved sources in the brain typically yields spatially highly correlated results. Publications focusing on this type of analyses report a spatial extent of typically multiple centimetres, which here is the case as well (see Figure 3A of the updated manuscript). As such, the single voxel selection boils down to selecting the peak response within a large patch of very similarly responding voxels. Using this approach we were able to select the frequency response with the highest possible SNR. We do not however claim that the respective single voxel is exclusively carrying this information. In addition we have added a short explanation to the Discussion, since we believe that virtual channel selection with a different objective (e.g. maximising the difference between conditions or maximising cross-frequency coupling, etc.) could indeed profoundly impact the EEG-fMRI correlation, which would open up opportunities for interesting analyses that are however beyond the scope of this project.

      We have added the following section to the Discussion:

      “Future work might also vary the exact virtual channel selection for obtaining EEG-based regressors. Here, we focused on the grid points (voxel locations) with the strongest α or γ response for each frequency band in each hemisphere, derived from the average frequency response to maximise SNR. However, selecting the respective virtual channels based on the response to specific stimulus features or the interaction between high and low frequency bands are possibilities worth exploring in future work.”

      One area that would benefit from further clarification is the interpretation of gamma oscillations. The evidence for gamma involvement in the observed effects appears somewhat limited. For example, no significant gamma-related clusters were found for the feature-unspecific BOLD signal (Figure 2). Significant effects emerged only when the analysis was restricted to positively responding voxels, and even then, only for the contrast between EEG-coherent and EEG-incoherent conditions in the feature-specific BOLD response. It remains unclear how to interpret this selective emergence of gamma-related effects. Given previous literature linking gamma to feedforward processing, one might expect more robust involvement in broader, feature-unspecific contrasts. The current discussion presents the gamma-related findings with some confidence, and the manuscript would benefit from a more nuanced reflection on why these effects may not have appeared more broadly. The explanation provided in line 230, that restricting the analysis to positively responding voxels may have increased the SNR, is reasonable, but it may not fully account for the absence of gamma effects in V1's feature-unspecific response. Including the actual beta values from Figure 4 in the legend or main text would also help readers better assess the strength and specificity of the reported effects.

      We agree with the reviewer that the missing gamma-band response for the feature-unspecific signal, as well as the limitation of the effect solely to the feature-specific contrast for positive voxel selections only was unexpected. In fact, based on previous literature, we were expecting a feature-unspecific effect in the gamma band as well. However, the literature on laminar level EEG-fMRI is sparse and previous experiments used tasks that did not allow for the separation into distinct features (here left or right-oriented gratings). While we cannot fully explain the absence of the gamma effect for the feature-unspecific condition, we reasoned that our stimuli evoked weaker gamma band responses compared to previous literature. 

      The fact that we only see a significant gamma band response for the contrast for positive voxel selections can be interpreted twofold: First, previous experiments limit their analyses to positive BOLD responses only, for which we find an effect as well. Second, the fact that a significant effect could only be obtained for the contrast, might indicate that gamma band activity is related to the actual features themselves. A cortical column responding to left-oriented gratings would then be related to a gamma band response linked to that orientation. If this response to a single orientation could not be fully captured due to SNR-related issues, we would not see this effect in the congruent-only condition and also not in the feature-unspecific condition (because this boils down to both congruent conditions combined). If gamma-band oscillations are actually reflecting the response of a column to a certain orientation, then the lowest possible response would be found for the exact orthogonal orientation (here the incongruent condition). The contrast between most preferred and most not-preferred orientation might have helped to overcome the inherently low SNR, explaining the results for the contrast.

      Lastly, we did not include actual beta values in the main text, because those might be misleading. We compute the relationship between EEG power and the BOLD signal for every voxel separately, then weighted the result with the respective layer weight and lastly aggregated across voxels.This means that the beta values express the strength of the association between EEG and fMRI for an average voxel. For this reason the values are tiny and the values themselves are less meaningful than “typical” beta values.

      We have added or modified the following sections in the Discussion or Methods respectively:

      “Based on previous literature, we expected a γ band effect for the congruent condition of the feature-specific analysis (Scheeringa et al., 2016), which we did not observe. A possible explanation could be the used stimulus material in our experiment as compared to Scheeringa et al., (2016). Muthukumaraswamy et al., (2013) found that stationary gratings evoke a weaker γ band response as compared to moving annular stimuli that have been used by Scheeringa and colleagues. If γ is related to the processing of the actual features themselves (e.g. to a column preferably responding to left-oriented gratings), then contrasting congruent and incongruent voxel selections provides the largest possible contrast-to-noise ratio (CNR). In turn annular stimuli as previously used might have activated all possible orientations and thus might have greatly boosted γ SNR.”

      “The described procedure of computing a GLM based on z-transformed data using z-transformed predictors yields β-coefficients that reflect the average relationship of a single voxel's BOLD response for a given layer (fraction of the single voxel's β) with EEG power changes of a specified frequency.”

      Relating to behavioural findings for underlying neural activity, could the authors test on a trial-by-trial basis how behavioural performance relates to the BOLD signal / oscillatory activity change? Line 305 states that "Since behavioural performance in the present study was consistently high at 94% on average and participants were instructed to respond quickly to potential oddball stimuli, a higher alpha frequency might reflect a more successful stimulus encoding and hence faster and more accurate behavioural performance." Also, this might help to relate the findings to the lower vs upper alpha functionality difference.

      This is a very interesting suggestion. We now include an exploratory analysis of the relationship between frequency and behavioural performance in the Supplementary Figures (see Figure S12). We did not perform a correlation between behavioural performance and alpha over trials because of the low numbers of oddball trials (N=40) and very limited number of false responses (94% response accuracy on average). However, we computed a correlation across participants. After averaging the alpha time-frequency spectrum across non-oddball trials, the individual alpha frequency was determined by the frequency where the alpha decrease (between 0.1 and 0.8 s post-stimulus) was largest. The correlation between alpha frequency and either reaction times and d’ (as a measure for accuracy), yields a significantly positive relationship between d’ and alpha frequency. This indicates that alpha frequency is related to task performance. We interpret those exploratory findings such that high behavioural accuracy is reflected by a stronger modulation of high-frequency alpha power. 

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      In Figure 4, the EEG alpha specificity plot shows relatively large error bars, and there is visible overlap between the lower and upper alpha in both congruent and incongruent conditions. While upper alpha shows a positive slope across conditions and lower alpha remains flat, the interaction appears to be driven by the change from congruent to incongruent in upper alpha. It is worth clarifying whether the simple effects (e.g., lower vs upper within each condition) were tested, given the visual similarity at the incongruent condition. Overall, the significant interaction (p < 0.001, FDR-corrected) is consistent with diverging trends, but a breakdown of simple effects would help interpret the result more clearly. Was there a significant difference between lower and upper alpha in congruent or incongruent conditions?

      We thank the reviewer for this important remark and have added a simple effects analysis (see Figures 4 b and 5 b, e). We found that the main driver for the interaction between congruence condition and alpha frequency is upper alpha. Specifically the negative relationship between upper alpha and the BOLD signal is significantly stronger for the congruent condition and weaker for the incongruent condition. This indicates the upper alpha indeed is related to the processing of visual features.

      We have added a simple effects analysis (See Figures 4 and 5).

      We have added or modified the following in Results, Discussion and Methods respectively:

      In Results:

      “We furthermore found a significant interaction (p<sub>FDR</sub> < 0.05) between positive or negative BOLD signal change and lower or upper α sub-bands (8 - 10 or 11 - 13 Hz respectively) by means of a linear mixed effects model. An analysis of simple effects revealed that upper α frequencies are stronger negatively related to the positive BOLD signal as compared to lower α on a trend level (p<sub>FDR</sub> = 0.097).”

      “After correcting for multiple comparisons, we found a significant interaction (p<sub>FDR</sub> < 0.001). This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis. We found a significantly stronger negative relationship of upper α and the BOLD signal for congruent selections (p<sub>FDR</sub> < 0.01) and the reverse for the incongruent condition (p<sub>FDR</sub> < 0.01), as well as a significantly stronger negative relationship within the upper α sub-band for congruent over incongruent voxel selections (p<sub>FDR</sub> < 0.01).”

      “This interaction is mainly driven by the upper α sub-band, as indicated by the simple effects analysis, which revealed a significantly stronger negative relationship of upper α and the BOLD signal for congruent over incongruent selections (p<sub>FDR</sub> < 0.001).”

      In Discussion:

      “An analysis of upper and lower α sub-bands revealed a significant interaction between congruence condition and α frequency. This interaction was mainly driven by the upper α band (11 to 13 Hz). For congruently selected voxels, the negative relationship was significantly stronger (over lower α), while for incongruent selection it was significantly weaker. No such difference has been observed for the lower α component, which indicates a more feature-specific involvement of upper α and a more general modulatory effect for lower α frequencies.”

      In Methods:

      “Significant interactions were decomposed into simple effects using Wald tests on the model coefficients, ensuring that post-hoc comparisons were derived from the same statistical global variance as the primary interaction.”

      Overall, this study provides a valuable contribution to the literature on oscillatory dynamics and laminar fMRI, though some interpretations would benefit from further clarification or qualification.

      Reviewer #3 (Public review):

      Summary:

      Clausner et al. investigate the relationship between cortical oscillations in the alpha and gamma bands and the feature-specific and feature-unspecific BOLD signals across cortical layers. Using a well-designed stimulus and GLM, they show a method by which different BOLD signals can be differentiated and investigated alongside multiple cortical oscillatory frequencies. In addition to the previously reported positive relationship between gamma and BOLD signals in superficial layers, they show a relationship between gamma and feature-specific BOLD in the deeper layers. Alpha-band power is shown to have a negative relationship with the negative BOLD response for both feature-specific and feature-unspecific contrasts. When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency, and can therefore be interpreted as more feature-specific than lower frequency alpha.

      Strengths:

      The use of interleaved EEG-fMRI has provided a rich dataset that can be used to evaluate the relationship of cortical layer BOLD signals with multiple EEG frequencies. The EEG data were of sufficient quality to see the modulation of both alpha-band and gamma-band oscillations in the group mean VE-channel TFS. The good EEG data quality is backed up with a highly technical analysis pipeline that ultimately enables the interpretation of the cortical layer relationship of the BOLD signal with a range of frequencies in the alpha and gamma bands. The stimulus design allowed for the generation of multiple contrasts for the BOLD signal and the alpha/gamma oscillations in the GLM analysis. Feature-specific and unspecific BOLD contrasts are used with congruently or incongruently selected EEG power regressors to delineate between local and global alpha modulations. A transparent approach is used for the selection of voxels contributing to the final layer profiles, for which statistical analysis is comprehensive but uses an alternative statistical test, which I have not seen in previous layer-fMRI literature.

      A significant negative relationship between alpha-band power and the BOLD signal was seen in congruently (EEGco) selected voxels (predominantly in superficial layers) and in feature-contrast (EEGco-inco) selected (superficial and deep layers). When separated into lower (8-10Hz) and upper (11-13Hz) alpha oscillations, they show that higher frequency alpha showed a significantly stronger negative relationship with congruency than lower frequency alpha. This is interpreted as a frequency dissociation in the alpha-BOLD relationship, with upper frequency alpha being feature-specific and lower frequency alpha corresponding to general modulation. These results are a valuable addition to the current literature and improve our current understanding of the role of cortical alpha oscillations.

      There is not much work in the literature on the relationship between alpha power and the negative BOLD response (NBR), so the data provided here are particularly valuable. The negative relationship between the NBR and alpha power shown here suggests that there is a reduction in alpha power, linked to locally reduced BOLD activity, which is in line with the previously hypothesized inhibitory nature of alpha.

      Weaknesses:

      It is not entirely clear how the draining vein effect seen in GE-BOLD layer-fMRI data has been accounted for in the analysis. For the contrast of congruent-incongruent, it is assumed that the underlying draining effect will be the same for both conditions, and so should be cancelled out. However, for the other contrasts, it is unclear how the final layer profiles aren't confounded by the bias in BOLD signal towards the superficial layers. Many of the profiles in Figure 3 and Figure 4A show an increased negative correlation between alpha power and the BOLD signal towards the superficial layers.

      We thank the reviewer for this important remark. Reviewer 1 raised a similar concern and I would like to refer you to our response to Reviewer 1, point 4. The veinal draining typically results in a higher signal change closer to the cortical surface. We did not take any measures to counteract this effect, but provide an analysis of tSNR in Supplementary Figures (see Figure S10). Possibly due to the drainage of physiological noise towards the surface, we found the highest tSNR in deep, followed by middle and superficial layers. To verify those results we computed the same analysis using a second layering algorithm, which resulted in the same profile, but overall less tSNR. Crucially the tSNR profile is not reflected in our EEG-fMRI results.

      We have added Figure S10 to Supplementary Figures and the following section to the Discussion:

      “A major concern for laminar fMRI is the vascular draining effect (Markuerkiaga et al., 2016), which typically leads to increased signal amplitudes closer to the surface. Here, we did not investigate the signal change per se, but rather the relationship with EEG power changes. To ensure that the results do not stem from differences in tSNR across layers, we conducted a tSNR analysis (see Figure S10 in Supplementary Figures). We found that the highest tSNR was obtained from deep layers, as compared to middle and superficial layers. To verify, we computed the tSNR using a second layering algorithm (LayNii, see Huber et al. 2021), which yielded lower absolute values, but a comparable layer profile. The obtained tSNR is not reflected in any of our result profiles (see Figures 4 and 5), which strengthens the validity of the here presented results. We speculate that tSNR in deep layers is higher, because both functionally relevant components of the BOLD signal and physiological noise components drain towards superficial layers.”

      When investigating if high alpha (8-10 Hz) and low alpha (11-13 Hz) are two different sources of alpha, it would be beneficial to show if this effect is only seen at the group level or can be seen in any single subjects. Inter-subject variability in peak alpha power could result in some subjects having a single low alpha peak and some a single high alpha peak rather than two peaks from different sources.

      We agree with the reviewer that a bias in a subset of participants to generally higher or lower alpha frequencies could potentially skew the presented results. While the initially computed model included a random intercept for the frequencies, we have now added the random slope as well. This ensures that the difference between low and high frequency alpha is indeed only driven by the difference in condition and not the result of individual differences across conditions themselves.

      In order to verify that not a small subset of participants is driving the result pattern, we also computed the fraction of participants that either show the dual alpha pattern (i.e. follow the exact pattern of the group average), contribute to the group average with a single peak or contradict the pattern entirely. Thereby, 40.4% of all participants show a dual alpha pattern, 38.4% a single alpha pattern in the direction of the group average and 21.2% contradict the group average. See Author response image 1:

      Author response image 1.

      Alpha Response Patterns with Example Subjects: V1 Feature Specific Contrast

      We would also like to highlight our added exploratory analysis of the relationship between alpha frequency and behavioural performance, which was requested by Reviewer 2, point 3. We find a significant positive correlation between alpha frequency and task performance on a group level. This indicates that higher alpha frequencies might be related to better discrimination of visual features. We speculate that participants with better task performance are capable of modulating their upper alpha more than participants with worse performance.

      We have added Figure S12 to Supplementary Figures.

      We have also added the following sections to Results and Discussion respectively:

      “An exploratory analysis of the relationship between individual α frequency (IAF) and task performances underlines this finding (see Figure S12 in Supplementary Figures). Thereby the IAF was obtained from the average α power spectrum of each participant. The frequency with the strongest decrease between 0.1 and 0.8 s after stimulus onset served as the IAF. We correlated IAF with average response times to correct oddball trials and d' as a measure for accuracy and found a significant positive correlation between IAF and d' (p < 0.05).”

      “We exploratively correlated the average IAF during non-oddball trials with the average task accuracy (d') across participants and indeed found IAF and task performance to be positively correlated (See Figure S12 in Supplementary Figures).”

      The figure layout used to present the main findings throughout is an innovative way to present so much information, but it is difficult to decipher the main findings described in the text. The readability would be improved if the example (Appendix 0 - Figure 1) in the supplementary material is included as a second panel inside Figure 3, or, if this is not possible, the example (Appendix 0 - Figure 1) should be clearly referred to in the figure caption. 

      Since Reviewer 1 suggested using an entirely different figure layout, we now opted to remove some information from the main text figures (we only show the 10% threshold, but 5% and 25% is in Supplementary Figures) and chose a more common figure layout. See Figures 4 and 5.

      Recommendations for authors:

      Reviewer #2 (Recommendations for the authors):

      The contrasts used in the analysis are not clearly introduced in the main text. While the methods section explains them more thoroughly, some of this explanation would be better placed in the results section, where the contrasts are first used. Specifically, the concepts of "feature-specific" vs. "feature-unspecific" BOLD signals are introduced with a very brief definition, which could be confusing for readers. The same applies to the terms EEG co and EEG inco; it would help to briefly explain these when they are first mentioned in the results. The supplementary figures and legends are helpful, so it is clear that the authors were prioritising clarity overall.

      The respective analyses are now also explained in the Results section:

      “During each trial either a left or a right-oriented grating was presented, from which two types of analyses have been derived: feature-unspecific BOLD activation (i.e. the response to any stimulus orientation), and feature-specific BOLD activation (i.e. the response to a specific stimulus orientation or the contrast between them). Thereby, fMRI data and EEG-based regressors could either be combined congruently (Co) by combining the BOLD signal of orientation-selective voxels with EEG-based regressors built from the same orientation trials, or incongruently (Inco), by combining the orientation-specific BOLD signal with EEG-based regressors built from the other orientation trials. Finally, those two congruency conditions have been contrasted (Co-Inco).”

      Figures are overall clear and illustrative of the results. For Figure 4, however, the use of dotted elements makes it somewhat harder to interpret what's being shown. While the supplementary figure clarifies the findings, rephrasing the figure legend to explain what the dotted lines represent would be helpful.

      Figures 4 and 5 have been replaced with a new layout and legends have been improved.

      The reported ranges overlap (e.g., alpha: 2-32 Hz; gamma: 20-120 Hz). It would be helpful to explain why such overlapping bands were chosen.

      Both frequency bands of interest differ slightly in their later time-frequency analysis (i.e. number of tapers and filter type). The overlap itself is not meaningful per se and results from the selection of a wide band for each respective sub-band. This wide selection was chosen to avoid filter artefacts. For the alpha sub-band, we also wanted to ensure that the beta spectrum is covered which also includes the alpha harmonic and for the gamma band that the full range of high-frequency activity is captured (e.g. EMG activity).

      Only a single time point was used for baseline correction of the low alpha band. Is this typical? The authors note that due to the gradient artefact arising in the pre-stimulus period, the baseline correction is somewhat difficult, although further clarification would be useful here.

      Relatedly, was pilot scanning conducted? If so, was the presence of strong gradient artefacts unexpected? More details about this would strengthen the methodological transparency.

      Indeed only a single time bin was used as the baseline for the alpha sub-band. After the piloting phase a slight adjustment to the final fMRI sequence has been made which was not expected to introduce gradient artefacts so close to the onset of the stimulus. Unexpectedly, those artefacts were visible until 300 ms before the onset of the stimulus. Similarly, a pre-stimulus alpha was observed (starting 250 ms before the onset of the stimulus), which we also aimed to exclude from the baseline period. In the end only the time bin centered at 300 ms prior to stimulus onset was chosen. However, this time bin contains 400 ms of data (the width of the window for the time frequency analysis). Thus, the term time point was misleading, because the actual time window that made up the baseline is 500 ms to 100 ms prior to the onset of the stimulus. 

      We have adjusted our wording in Methods to make this more clear:

      “For this reason, the low frequency baseline period comprised only a single 400 ms time bin centred around -0.3 s, because a pre-stimulus α decrease was expected starting around 0.25 s prior to stimulus onset.”

      Including a one-sentence explanation of the AROS test in the main text for clarity. As line 796 in the methods: "Each significant cluster has been further processed by means of an auto-regressive rank order similarity (aros) test (Clausner and Gentili, 2022). The fundamental idea behind the AROS test is whether group averages (i.e. averages of the signal of cortical layer in the present case), can be ranked such that the rank order is explained significantly better by the data than it would if the average data could not be meaningfully sorted (i.e. is shuffled)."

      An explanation has been added to the Results section:

      “Each significant cluster was then averaged along the frequency dimension at the widest point to enable an auto-regressive rank order similarity (aros) test Clausner & Gentili (2022), testing the laminar activation profile. The aros test transforms the layer averages into a rank order and tests - using a permutation procedure - if the rank order of the layer averages explains the data better than a random rank order (shuffled layer labels) would.”

      Line 223: "In fact, an analysis of the relationship between the EEG signal and the BOLD signal that focused on the feature contrast only (L - R; independent of the comparison to baseline) revealed a trend-level result with an even stronger deep layer contribution as compared to superficial layers." Could you point to which figure represents this finding - Figure 4B?

      This refers to Figure 4A in the old manuscript, for the 25% threshold for the gamma band. Since now the new figures do not include the 25% threshold anymore, it refers to Figure S4i.

      The number of participants is missing from the main text. Including this in the results section would improve clarity.

      The description of our sample has been moved from Methods to Results.

      Given the complexity of the data acquisition and analysis, the well-designed and easy-to-follow analysis pipeline figure (currently in the supplement) would be better placed in the main text.

      The mentioned Figure has been moved to the main text (now Figure 2).

      Also, simply out of curiosity, what do the authors think about the theta blob around 200ms post-stimulus?

      The theta blob most likely reflects the post-stimulus ERP as often observed in response to visual stimuli. We hypothesise that it is stronger in the middle and superficial layers, but we did not want to extend too much the scope of this paper. Additional analyses could be performed in the future on this evoked activity.

      Reviewer #3 (Recommendations for the authors):

      (1) Minor Corrections to the text and figures:

      We would like to thank the reviewer for the very valuable recommendations. Below we shortly describe how each suggestion has been implemented.

      We have made the white box more clear (see Figure 3 B).

      (b) Page 10: Top of 2nd paragraph - 'The full experimental protocol comprised a high resolution anatomical T1 scan lasting for 8 min'. The methods state this scan is 6 min 31 sec.

      The confusion results from the fact that the T1 scan was recorded during a short practice block that the participants performed inside the scanner. This block lasted 8min during which the 6 min 31 sec T1 scan was recorded. We have made this more clear:

      “Once prepared, the participant was placed inside the scanner and performed an 8 min practice block. A T1-weighted scan was acquired during this time in the sagittal orientation using a 3D MPRAGE sequence Brant-Zawadzki et al., (1992) with the following parameters: TR/TI = 2.2/1.1 s, 11° flip angle, FOV 256 x 256 x 180 mm and an 0.8 mm isotropic resolution. Parallel imaging (iPAT = 2) was used to accelerate the acquisition, resulting in an acquisition time of 6 min and 31s.”

      (c) Page 10: 'Stimulus presentation' paragraph - 'Stimuli were projected onto a screen behind the subject's head using'. The use of 'subject' should be replaced with 'participant' throughout.

      We have corrected the phrasing.

      (d) Page 14: Figures 2A and 2B are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (e) Figure 5 caption: 'Regressors are build for each time-frequency bin separately.' should be 'built'

      We have corrected the mistake.

      (f) Page 16, final paragraph: 'Afterwards, each of the resulting regression coefficients (B coefficients) was multiplied with the voxel specific layer weights that have been obtained as described above.' Should be 'were multiplied'

      We have corrected the mistake.

      (g) Page 17: 'Subsequently, separate analyses were done for two frequency of interest (FOI) ranges centerd around' - typo

      We have corrected the mistake.

      (h) Page 17 - 'Within these frequency ranges inferential statistics based a cluster level' - missing word. Should be 'based on a cluster level'

      We have corrected the mistake.

      (i) Page 14 Figure 2B and 5D are referred to incorrectly as being in the supplementary material.

      We have corrected the mistake.

      (2) fMRI data pre-processing:

      Please provide a comment on the EEG-fMRI data quality - e.g. tSNR of EPI data. Perhaps example EPI data could be shown in the supplementary information.

      We included the below Figure S11 in Supplementary Figures showing an example EPI. We have also included an illustration of the result of our layering approach. Furthermore, we included a tSNR analysis (see Response to Reviewer 1, point 4).

      On a practical note - with 14-minute long runs whilst wearing an EEG cap, I would expect participant motion to be a concern. Could you provide some metrics on perhaps the average of the mean and maximum per subject displacement/rotation?

      We ensured that participants receive tactile feedback for their respective head motion from a strip of tape span across their foreheads. This resulted in overall manageable motion during each experimental block. During the main experiment, the average framewise displacement was 0.3 mm, with an average total translation of 1.6 mm and an average total rotation of 1.6 deg within each block. 

      We have added Figure S13 to Supplementary Figures.

      We have added a section to Methods:

      “Subject motion per block was low, with a mean (SD) frame-wise displacement Power et al. (2012) of 0.34 mm (0.24 mm) for the main experiment and 0.23 mm (0.22 mm) for the retinotopy (see also Figure S13 in Supplementary Figures).”

    1. Author response:

      We sincerely thank the editors and reviewers for the positive assessment of our work and for the constructive and insightful feedback.

      We fully agree with the major points raised in the public reviews and outline below our planned revisions to address them.

      Reviewer #1 raised two important concerns regarding our methodology. First, the determination of allele dosage is insufficiently explained, which is central to our ploidy assignment and downstream analyses. Second, the setup and sample sizes of the common garden experiments are unclear, raising questions about the robustness of our conclusions. We accept these criticisms and will address them as follows.

      Regarding allele dosage, we will add a detailed step-by-step description of our calling pipeline in the Methods section, including the criteria for peak height ratios and thresholds used to assign copy numbers. We will also clarify a crucial biological detail: the common reed (Phragmites australis) is an allotetraploid in its origin. As a consequence, many molecular markers, including the widely used SSR markers in previous studies, behave as disomic markers (i.e., two homeologous copies inherited in a Mendelian manner). Therefore, observing more than two alleles at a locus is indeed indicative of higher-level ploidy (hexaploidy or octoploidy) in this system. We will explicitly state this to resolve any confusion about why tetraploids in our dataset are treated as having a maximum of two alleles, while hexaploids and octoploids can carry more.

      Regarding the common garden experiment, we will explicitly report the replication number for each lineage-by-treatment combination and clarify the experimental design. We will also discuss the statistical approaches used given the sample sizes, while acknowledging that the consistency between experimental results and distributional patterns lends additional support to our conclusions.

      Reviewer #2 raised three substantive framing issues. First, ploidy is completely confounded with genetic background, yet our manuscript places undue emphasis on polyploidy as a causal factor. Second, our species distribution models treat each lineage as a homogeneous entity, failing to capture within-lineage variation and thus repeating the oversimplification we criticize. Third, we insufficiently explore the evolutionary significance of asymmetric introgression, gene flow, and the novelty of combining SDM with experiments. We fully agree with these points and will revise accordingly.

      To address the confounding issue, we will substantially reframe the manuscript to de-emphasize claims about polyploidy as a causal driver, and instead focus on the adaptive differentiation among distinct genetic lineages that happen to differ in ploidy. The Discussion will explicitly state that dissecting ploidy effects from background genetic effects will require future experimental approaches.

      To address the simplification in SDMs, we will add a clear acknowledgment of this limitation, discussing how it may affect predictive accuracy and suggesting that future studies incorporating population-level genomic data could more directly assess evolutionary potential.

      To address the insufficient exploration of introgression and the novelty of our approach, we will expand the Introduction to better highlight the value of coupling controlled experiments with SDMs at the intraspecific level. In the Discussion, we will elaborate on the evolutionary significance of asymmetric introgression, including testable hypotheses about how gene flow might mediate the spread of heat-tolerance alleles and influence lineage geographical limits under climate change.<br /> We also thank the reviewer for the suggestion to emphasize the experimental validation of SDM efforts, which we will incorporate into a revised Introduction.

      Looking beyond the present study, we envision three complementary directions that build upon our current findings. Expanding common garden experiments to include admixed individuals would test whether introgressed genomic blocks confer fitness advantages under thermal stress. Leveraging the population genomic framework established here, we will transition to whole-genome resequencing for selection scans and genotype-environment association analyses to pinpoint adaptive loci and reveal whether heat-tolerance alleles are preferentially transferred via asymmetric introgression. We will also integrate transcriptomic profiling with phenotypic measurements to identify candidate genes whose expression correlates with thermal performance and introgressed ancestry, helping to disentangle ploidy effects from genetic background. Together, these directions span expanded phenotyping, whole-genome resequencing, and transcriptome-guided discovery, forming an integrated framework that moves from the correlative patterns reported here toward mechanistic understanding. These perspectives are briefly outlined in our Discussion, and we hope the present study will serve as a foundation for these future investigations, which we plan to pursue in subsequent work.

      We believe these revisions will substantially strengthen the manuscript.

    1. Author response:

      eLife Assessment:

      This study presents a valuable metagenomic analysis of the gut microbiome in sickle cell disease (SCD) patients, revealing associations between bacteriophage, host immunity, and SCD pathophysiology. While these data are interesting and helpful for hypothesis generation, they are deemed incomplete; additional experiments would be needed to test causality and to provide mechanistic insight. Despite these limitations, this work will be of broad interest to researchers studying SCD, immunology, phage biology, and the microbiome, adding to the small but growing literature suggesting a microbial component to SCD.

      The authors would like to thank the reviewers for thorough and constructive comments on our manuscript. We will make major updates to the manuscript addressing the following points and suggestions from the three reviewers: (1) assessing HbAS/AA genotype influence on microbiome composition; (2) conducting the requested beta diversity analysis, (3) conducting the requested sensitivity analysis to assess the impact of disease severity and therapy on microbiome and virome features; (4) modifying our language to clearly state that our results do not indicate causality or mechanism of microbiome interactions with sickle cell disease pathophysiology; (5) improved discussion of the phage results and their strengths and limitations; (6) additional changes throughout for clarity and correction of errors. We will change the title to “Bacterial and viral gut microbiome alterations characterize microbiome-immune-pathophysiology axes in Sickle Cell Disease.” These additions will greatly improve our work and presentation and we are grateful to the reviewers and our editors.

      We have indicated where specific changes were made in response to the public reviews below.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, Flamholz and colleagues use metagenomic sequencing to profile the microbiome of individuals with sickle cell disease (SCD), the most common genetic blood disorder in the world. To build on previous studies that found dysbiosis in SCD, this manuscript aims to examine whether changes in either bacterial species or bacteriophages correlate with inflammatory hallmarks of the disease. The authors claim that sickle cell dysbiosis does not correlate with inflammatory hallmarks of the disease, but instead, aged neutrophil numbers and bacteriophages do. Appropriate control subjects and additional analyses are needed to support that conclusion.

      Strengths:

      The primary strength of this paper is the investigation into disease-associated changes in bacteriophages. This is an entirely novel idea in the sickle cell field, and based on the current results, may be an important, under-recognized disease hallmark. It is unclear, however, if phages are "the chicken or the egg" in terms of sickle cell inflammatory profiles; do these increases in phage number simply result from other disease processes, or are they in any way contributing to disease pathophysiology?

      Weaknesses:

      A primary weakness of the manuscript is the fact that the majority of individuals included in the control group maintain sickle cell trait (HbAS genotype). Although typically asymptomatic, it is unclear if this genotype is associated with microbial changes that would not be observed in a true control group (HbAA genotype). This is a significant limitation that may limit the ability to draw conclusions from the current data set.

      Another key weakness is the lack of beta diversity assessment. Although decreased alpha diversity is observed in individuals with SCD, and specific bacterial taxa are differentially abundant following multivariate analyses, there is no overall comparison of bacterial community composition between individuals with SCD and controls. Prior to drawing conclusions about the relationship (or lack thereof) between the SCD microbiome and inflammatory markers, it is important to know if this study did indeed find disease-associated changes in microbiome composition.

      It is unclear which individuals were used for aged neutrophil (AN) and molecular data assessments. For example, were children who were still receiving penicillin prophylaxis included in these specific assessments? Given the authors' previous work demonstrating that antibiotic treatment decreases AN pathology, it seems critical to limit all AN/molecular analyses to older subjects who are not on daily penicillin treatment (if possible).

      A minor weakness is the continued use of "disease" vs. "healthy" indicators as primary microbiome metrics that are used for molecular correlations. The lack of metric specificity - and lack of discussion regarding which diseases were used to generate these indicators (how similar/different are they to sickle cell?) - could be said to make these metrics essentially meaningless.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will include an analysis evaluating the influence of control genoype (HbAA/HbAS) on our microbiome and virome results. To evaluate whether control genotype influenced major microbiome and virome features, analyses were restricted to control participants only. Controls were stratified by genotype as HbAA or HbAS. Four significant microbiome and virome features were tested: F:B ratio, Shannon diversity, provirus fraction, and virus count. HbAA and HbAS controls were compared using two-sided Mann-Whitney U tests. Benjamini-Hochberg FDR correction was applied across the four tested features. HbAS and HbAA controls did not differ significantly for F:B ratio, Shannon diversity, provirus fraction, or virus count. The inclusion of HbAA/AS will strengthen our results with respect to the observation that sickle cell disease patient microbiomes remain significantly different from sickle trait (HbAS) controls. These results will be reported in a new Supplemental Table.

      We will include a beta diversity analysis using MetaPhlAn species profiles. Beta diversity analyses were performed in Python using pandas and NumPy for data processing, scikit-bio for distance calculations and PERMANOVA, scikit-learn for ordination-related computations, statsmodels for multiple-testing correction where applicable, and matplotlib for visualization.

      For the primary disease/control comparison, samples were grouped as control or SCD. For the genotype control sensitivity analysis, samples were restricted to HbAA and HbAS individuals as described above. Species detected in at least 10% of included samples were retained for beta diversity analysis. To account for the compositional structure of metagenomic relative abundance data, species profiles were transformed using a centered log-ratio transformation after addition of a small pseudocount to accommodate zero values. Aitchison distances were calculated from the CLR-transformed species profiles. Statistical significance of group separation was assessed by PERMANOVA using 999 permutations. For the control versus SCD comparison, PERMANOVA was performed between the two disease-status groups. For the HbAA versus HbAS control comparison, PERMANOVA was performed among controls only.

      In the SCD cohort, beta diversity differed significantly between controls and SCD participants by Aitchison distance after CLR transformation (R<sup>2</sup> = 0.030, p = 0.001). In contrast, HbAA and HbAS controls did not differ significantly in beta diversity (R<sup>2</sup> = 0.024, p = 0.282), supporting the conclusion that the observed SCD/control separation was not driven by control genotype composition. These methods and results will be reported in the revised manuscript.

      The manuscript describing the microbiome health and disease indicators was submitted to eLife jointly with this manuscript as a package; eLife declined to review the indicator manuscript. Briefly, this study conducted a cross-disease meta-analysis of 38 studies comprising 8,204 samples and identified 100 bacterial taxa or “indicators” that are weakly but consistently associated with health or disease across diverse conditions, including, but not limited to, inflammatory bowel disease, colorectal cancer, type 2 diabetes. The indicator taxa were validated in an independent cohort of Graves’ disease patients. We currently cite an older version of this work posted as a preprint. The manuscript is currently under review at another journal and we will update this manuscript with the updated citation when it is available.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #2 (Public review):

      Summary:

      The study analyzes stool metagenomes from 98 SCD patients and 46 controls, with SCD and control groups matched on age, race, sex, and ethnicity. The authors report lower Shannon diversity, lower Firmicutes/Bacteroidetes ratio, loss of health-associated taxa, increased disease-associated indicators, altered butyrate/fatty-acid metabolism pathways, and enrichment of provirus/prophage fractions in SCD. They further correlate aged-like neutrophils and prophage fractions with inflammatory cytokines. The strength is that this is not just another 16S comparison. The use of whole-community metagenomics, immune profiling, neutrophil assays, and clinical metadata makes the study more biologically interesting than prior small SCD microbiome papers. The main weakness is that the causal and mechanistic interpretation is too strong. The data support an association between SCD status and microbiome/virome features, but they do not yet establish a clear "axis of pathophysiology." The provirus findings are intriguing, but require stronger statistical control, better validation, and more cautious interpretation.

      Strengths:

      The major strengths of the study include the clinically relevant disease setting, the use of whole-community sequencing, the integration of microbial, immune-cell, cytokine, and clinical measurements, and the novel attention to bacterial virus-related features. A particularly interesting aspect of the work is the analysis of virus-like elements integrated into bacterial genomes. The authors report that these elements are enriched in the gut microbial communities of patients with sickle cell disease and are associated with several inflammatory signals in blood. This observation is potentially important because it suggests that the microbial contribution to inflammation in sickle cell disease may involve not only bacteria but also bacterial virus-related genetic elements.

      Weaknesses

      The evidence for this proposed immune-related mechanism is incomplete. The study is cross-sectional and largely based on associations, so it cannot determine whether these virus-like elements drive immune activation, reflect immune activation, or are linked indirectly through disease severity, treatment history, or other clinical factors. The main limitations are the single-center design, modest sample size for some immune measurements, limited ability to control for treatment and disease heterogeneity, and the need for clearer multiple-testing correction in the correlation analyses. In particular, stronger adjustment for available clinical factors such as hydroxyurea use, transfusion history, pain admissions, genotype, and other markers of disease burden would help readers judge how specific the microbial and viral findings are to sickle cell disease itself.

      Overall, the authors largely achieve their descriptive aim of identifying gut microbial differences associated with sickle cell disease. The evidence is solid for the presence of broad microbial community differences, but incomplete for the stronger conclusion that virus-like elements form a pathophysiological immune axis. The work will likely be useful to researchers studying the microbiome, inflammation, and sickle cell disease, especially as a hypothesis-generating dataset. Its impact would be strengthened by more cautious interpretation, stronger control of clinical confounders, clearer statistical correction, and future longitudinal or experimental studies to test causality.

      We thank the reviewer for their helpful comments and suggestions. We want to first note that patients on prophylactic penicillin within six months of sample collection were excluded from the study due to the known impact of antibiotics on gut microbiomes, this will be clarified in the main text. We will temper our interpretation of our results, making clear that we are not arguing that either prophages or bacteria are causal or mechanistically associated with SCD biology and pathology. We will strengthen our control of clinical confounders, and add clearer statistical correction in the revision. We look forward to conducting future studies to test causality and understand mechanism.

      We will address the other recommendations from this reviewer in the revision.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Flamholz et al. sought to determine whether consistent and significant interactions exist between the gut microbiome and disease pathology in sickle cell disease (SCD). By sequencing and analysing metagenomes from faecal samples collected from 98 SCD patients and 46 control subjects, they identified community-level shifts in both the bacterial and proviral gut microbiome of SCD patients. They further reported correlations between the proviral microbiome and multiple blood cytokines, whereas similar associations were not observed for the bacterial microbiome. Based on these findings, the authors propose the existence of a viral-immune axis in SCD pathophysiology and targetable functional alterations in the gut microbiome.

      Strengths:

      This work includes the largest SCD cohort analysed to date, enabling analysis with relatively strong statistical power. In addition to profiling the bacterial microbiome, the study also examines the gut proviral microbiome, thereby providing a more comprehensive investigation of the topic. The newly generated metagenomic dataset will also be valuable for further meta-analysis by the wider community. Overall, the authors have largely achieved their aims.

      Weaknesses:

      However, this study represents a single-centre cross-sectional investigation, and most findings remain correlative in nature. In particular, the claim that the study identifies targetable functional alterations in the gut microbiome for disease treatment may be somewhat overstated. Although the reported functional module changes in SCD patients are intriguing, additional mechanistic and/or longitudinal evidence would be required before these features can realistically be considered targetable.

      We thank the reviewer for their helpful comments and suggestions. We will note in the text that additional mechanistic and longitudinal studies are required before we can target the microbiome and virome in SCD and clarified that this is a single center, cross-sectional. We will make further modifications to the manuscript to clarify cohort features (specifically, age and race were matched, other baseline characteristics were balanced), to properly describe the Shannon diversity metric, and to fix several errors that this reviewer caught.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Singh et al. presents an application of MOA-seq to better define transcriptional control underlying the hypoxia response in human endothelial cells. This group's previously described MOA-seq technique allows for precise, identity-agnostic mapping of occupied sites of DNA-binding proteins across the epigenome and over time. Here, they applied MOA-seq to HUVECs under normal oxygen conditions or variable lengths of hypoxia treatment, comparing changes in occupancy over time and associating these changes with corresponding transcriptome alterations. This approach revealed thousands of dynamically occupied sites comprising 10 major kinetic clusters that appear to define distinct subsets and phases of the hypoxia response. Analysis of DNA motifs in these dynamically occupied regions captured the known major roles of HIF1A in the hypoxia response and also implicated new HIF1A-associated regulators. Importantly, they also identified many potential HIF1A-independent candidate TFs that act at HREs, which has been an outstanding question in the field. Additionally, this study identified ~7K additional sites not previously defined as regulatory elements by ENCODE.

      Strengths:

      Overall, this study is well executed and described, providing new biological insights as well as a rich data resource for the field. As MOA-seq was previously developed for use in plants, this work demonstrates the application of this method in mammalian cells and highlights its utility in identifying new potential regulatory sites not captured by DNase-seq or ATAC-seq. The conclusions made by the authors are well supported by the results, with the caveat that extensive use of DNA motif identification and ontology analyses invariably leads to some uncertainty regarding factor identity and gene network properties.

      Weaknesses:

      There are several areas where the clarity of presentation could be improved:

      (1) Given the importance of the methodology, the methods section needs more detail on how the extent of MNase digestion is chosen to achieve optimal results with MOA-seq. This is described to some extent in the description of control library preparation, but not for the experimental samples.

      We thank the reviewer for noting this unintended omission. We have not updated the Methods section to specify as follows:

      "Digestion patterns were assessed via gel electrophoresis, and the light digest levels ideal for MOA-seq (as per Savadel et al., 2021) were selected as the lightest digest levels that give a pattern of a nucleosomal ladder spanning the entire DNA fragment size range from undigested to mononucleosome bands, as indicated in Figure 1 with the asterisk-marked gel lanes."

      (2) The abstract describes this approach as "native cistrome profiling" but this is misleading since formaldehyde fixation is used.

      We believe the formaldehyde fixation captures native chromatin structure, but indeed we are digesting fixed chromatin and have updated the wording to read as “in situ cistrome profiling.”

      (3) Species- and field-specific jargon and abbreviations need to be clarified on first usage. For example, on page 9: "Downsampling analysis was carried out for two sets of published reference peaks; the CTCF cCRE peak midpoints and for the ERG motif under the ERG ReMap ChIP-seq peaks." The different categories of cCREs were not clearly defined, nor will it be clear what the term ReMap refers to for those outside the field. The sentence after this refers to IDR, which also should be defined.

      We thank the reviewer for highlighting the need for clearer definitions of field-specific terminology and abbreviations. In response, we have revised the manuscript to explicitly define all relevant terms at first mention. Specifically, we now describe the ENCODE candidate cis-regulatory element (cCRE) catalogue and define the individual cCRE categories, including promoter-like (PLS), proximal enhancer-like (pELS), distal enhancer-like (dELS), DNase I–H3K4me3 (K4m3), and CTCF-only regions. We also clarify that ReMap is a curated database of human transcriptional regulator binding peaks derived from ChIP-seq, ChIP-exo, and DAP-seq experiments. Additionally, we now define IDR as the Irreproducible Discovery Rate framework upon first use.

      (4) Figure 4C: Are these motifs examined under MOA sites specifically or anywhere in the genes in question?

      Leading up to and including Figure 4C, we have not yet examined any motifs. Instead, Figure 4C compares gene sets, one defined by our diff-MOA, and those from GO libraries, in this case the "target genes" which are defined by TF-specific studies, primarily ChIP-seq but also related immuno-based mapping techniques. Consequently, the analysis shown in Fig. 4C is not a motif enrichment analysis. Instead, we used the ENRICHR gene set enrichment analysis tool with ENCODE and ChEA consensus transcription factor target gene sets. Thus, the analysis was performed at the gene-set level, and transcription factor motifs were not examined within diff-MOA peaks or elsewhere in the associated genes for Fig. 4C. We note that motif enrichment within diff-MOA peaks was subsequently examined separately in Fig. 6. In Fig. 7, we further examined differentially expressed genes associated with diff-MOA peaks containing enriched transcription factor motifs and used clustering analyses to investigate their regulatory relationships. We have clarified these distinctions in the revised manuscript.

      If the question is about the location of MOA footprints relative to gene structure, we did not examine any MOA sites at any specific location, just overlapping the gene +/- 200 bp, as indicated in Fig. 4B.

      (5) Figure 5B shows that up-DEGs with diff-MOA footprints tend to show more losses of footprints. Do the authors interpret this as a loss of repressor binding?

      Not exclusively, but yes, that is one plausible explanation. That is, the activation (defined by increased RNA levels) via de-repression could be happening. But we also expect these dynamic footprints to be but one component. In other words, we interpret the relationship as consistent with that possibility, but not only that possibility. A logical explanation is that loss of footprint occupancy associated with upregulated genes could be based on displacement of repressive DNA-binding factors, thereby contributing to transcriptional activation. Thus, while loss of repressor binding is a plausible explanation for a subset of these events, additional factor-specific experiments would be required to know for sure in each case. We have added text to the Discussion acknowledging this possibility.

      Reviewer #2 (Public review):

      Summary:

      Singh et al. apply MOA-seq to map transcription factor occupancy genome-wide in HUVECs across a hypoxia time course. The study provides a well-validated, high-resolution view of cistrome dynamics and identifies both HIF1A-associated and independent regulatory programs.

      Major Comments:

      Methodological validation is strong. MOA-seq's ability to map protein-bound DNA at near-nucleotide resolution without factor-specific antibodies is a genuine advance, and the cross-validation against independent ChIP-seq and ENCODE datasets is convincing. As noted, future work with additional biological replicates could further strengthen confidence in the smaller kinetic clusters.

      Regarding additional biological replicates, we have acknowledged this point in the discussion. Importantly, we did subject the replicates to IDR analysis, which we explain in the methods as "In accordance with ENCODE ChIP-seq guidelines (Landt et al., 2012), we further evaluated data quality by assessing pooled pseudo-replicate consistency and self-consistency for each individual replicate (Supplementary Table S2)." This IDR analysis demonstrated consistent peaks between our bioreplicates, meeting ENCODE guidelines. In addition, downsampling analysis demonstrated that our sequencing depth of coverage (Supp Fig 1) was over 10-fold greater than required. We do appreciate that it will be useful to have more biological replicates from other cell types, tissues, or organisms, and hope this study prompts just such future research.

      Imaging-based validation would strengthen the key biological claims. The kinetic clustering and pathway enrichments are computationally inferred. Orthogonal approaches, for example, live-cell fluorescence imaging of HIF1A nuclear translocation to confirm the proposed temporal binding waves, would provide independent experimental support.

      Live-cell imaging could indeed be interesting, but it is beyond our current capacity to add to this study and consider this an exciting future direction, but presence in the nucleus could include both bound and unbound HIF1, so the results may not easily track the DNA-bound HIF1 only.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      In Figure 3B, the x-axis is not labeled.

      Thank you for pointing this out. We have revised Figure 3B by adding the previously missing x-axis label.

      Reviewer #2 (Recommendations for the authors):

      In the abstract, it would be good to define what MOA-seq is and what the cistrome is.

      Thank you for this suggestion. We have revised the abstract to define both MOA-seq (MNase-defined cistrome-Occupancy Analysis sequencing) and the cistrome upon first mention to improve accessibility for readers who may be unfamiliar with these terms.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Summary:

      This paper is an exciting follow-up to two recent publications in eLife: one from the same lab, reporting that slender forms can successfully infect tsetse flies (Schuster, S et al., 2021), and another independent study claiming the opposite (Ngoune, TMJ et al., 2025). Here, the authors address four criticisms raised against their original work: the influence of N-acetyl-glucosamine (NAG), the use of teneral and male flies, and whether slender forms bypass the stumpy stage before becoming procyclic forms.

      Strengths:

      We applaud the authors' efforts in undertaking these experiments and contributing to a better understanding of the T. brucei life cycle. The paper is well-written and the figures are clear.

      Comments on revisions:

      We thank the authors for the revised manuscript and for considering our comments.

      We outline below the 3 points that, in our opinion, remain to be clarified.

      (1) Effect of NAG on slender-form infections in tsetse flies

      The conclusion that "NAG has a negligible effect on slender infections in tsetse flies" based on Figure 1, cannot be fully supported in the absence of a positive control. A relevant positive control is well established in the literature, namely that NAG promotes Tsetse infection by stumpy forms. Without such a control, it is not possible to exclude technical issues (for example, an ineffective NAG treatment), which would yield results similar to those presented in Figure 1.

      We agree that an internal stumpy-form positive control would provide an additional technical reference. However, the enhancing effect of NAG on stumpy-form midgut infections is well established and was also demonstrated under the experimental framework of our original study (Schuster et al. 2021, Figure 2A).

      The purpose of the present Research Advance was therefore not to re-establish the known effect of NAG on stumpy infections, but to test whether slender-form infections require NAG supplementation. Under the conditions tested here, slender bloodstream forms established midgut, proventriculus and salivary-gland infections also in the absence of NAG. We have revised the text accordingly to avoid implying a general absence of NAG effects and to make clear that our conclusion is restricted to slender-form infections under the conditions tested (line 128).

      (2) Infection of non-teneral flies

      Because the experiments shown in Figure 1 (teneral flies) and Figure 2 (non-teneral flies) were not conducted in parallel or under identical conditions, it is important that the figure legends clearly state the parasite numbers used in each case. Specifically, infections of teneral flies were performed with 200 parasites/mL (approximately 4 parasites per bloodmeal), whereas non-teneral infections used 1 × 10<sup>6</sup> parasites/mL (approximately 20,000 parasites per bloodmeal?). At present, this information is scattered across the Methods and Supplementary Tables 1 and 2, making it difficult for readers to immediately appreciate that the parasite load differs by roughly 5,000-fold between these conditions.

      As previously shown by the authors (Schuster et al., 2021) and in the Rotureau laboratory (Tsagmo Ngoune et al.), and as generally expected, the initial parasite dose strongly influences infection outcomes in teneral flies. In this context, it would be informative to know whether the authors have attempted infections of non-teneral flies using lower parasite numbers (noting that Tsagmo Ngoune et al. used a maximum of 10,000 parasites) and what the infection rate was.

      Relatedly, the statement in line 370 appears to be an overgeneralization, as fly age was not directly tested under matched experimental conditions:

      Line 370 - "Here, we unambiguously show that, in the absence of immunosuppressive treatment, slender forms can establish infections in tsetse flies, irrespective of the fly's age or sex."

      We thank the reviewer for highlighting the inconsistent presentation of parasite doses between Figure 1 and 2. We agree this is confusing and have revised the figure legends to clearly state both the parasite concentration (cells/mL) and estimated fly uptake per bloodmeal for each experiment (Lines 143 and 206).

      Regarding experiments with non-teneral flies using lower parasite numbers: We have not tested intermediate doses (e.g., 10,000 parasites/bloodmeal as used by Ngoune et al.) in non-teneral flies. Given that teneral flies already show relatively low infection rates even under optimal conditions, we chose the higher parasite dose (20,000 parasites/bloodmeal) for non-teneral flies to ensure sufficient statistical power for meaningful analysis of infection outcomes across different fly compartments.

      We acknowledge the reviewer's concern regarding the statement in line 370 and have revised this sentence (line 375) to more accurately reflect our experimental conditions, avoiding overgeneralization beyond the specific parameters tested.

      This reads now: “Here, we demonstrate that slender forms can establish infections without immunosuppressive treatment under the conditions tested. This infectivity was observed in both teneral and non-teneral, as well as in both male and female flies, indicating that slender forms retain transmission potential across different fly demographics. However, direct age comparisons under identical parasite doses remain to be tested.”

      (3) Transcriptomic analysis

      Supplementary Figure 8 lacks statistical analysis, which limits its interpretability. Two types of comparisons would be particularly helpful:

      (i) a comparison of PAD1/2 expression levels between slender and stumpy forms at 0 h; and

      (ii) for each gene, a comparison of the overall change in expression (from 0 to 72 h) between infections initiated with slender versus stumpy forms.

      In addition, the figure legend should clarify what "expression levels" refer to. TPM? Normalized counts?

      We appreciate this helpful comment and included statistical analysis for the expression of PAD1 and PAD2 (Supplementary Figure 8) between the two forms for the baseline (0 h) as well as during the differentiation to procyclic forms (0 h to 72 h) by using Welch´s t-test.

      While PAD1 did not show a statistically significant difference in this analysis, PAD2 displayed significant differences in expression dynamics over time. This supports the broader transcriptomic observation that slender- and stumpy-initiated differentiation follow distinct transcriptional trajectories before converging at the procyclic stage.

      We also clarified the figure legends showing the mean log2 counts per million (CPM) values.

      Finally, for the benefit of the field, eLife could encourage publishing a collaborative study in which the Engstler and Rotureau laboratories exchange parasite lines and culture protocols (including media with and without methylcellulose) and perform tsetse fly infections in parallel in their respective laboratories. Such an approach could help resolve the remaining discrepancies and provide a valuable reference for the community.

      We appreciate this constructive suggestion. A collaborative inter-laboratory study in which parasite lines, culture conditions and infection protocols are exchanged between the Engstler and Rotureau laboratories would be a valuable way to address the remaining discrepancies in the field. In particular, parallel infections using matched parasite lines and culture conditions, including media with and without methylcellulose, could provide a useful reference dataset for the community.

      At the same time, such a study would require substantial coordination, reciprocal strain exchange, protocol harmonization and new infection series in two laboratories. It therefore goes beyond the scope of the present Research Advance, which was designed to address the specific methodological concerns raised in response to our original publication. We have restricted our conclusions accordingly and view the proposed collaborative benchmark study as an important direction for future work.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript reports the discovery and characterization of the first bifunctional degrader of tankyrase. Notably, the tankyrase degrader exhibits stronger β-catenin inhibition and tumor growth suppression compared to conventional tankyrase inhibitors. Mechanistically, while tankyrase inhibitors stabilize tankyrase and promote Axin puncta formation - thereby impairing β-catenin degradation - the degrader avoids this effect, resulting in deeper suppression of β-catenin signaling. These findings suggest that targeted degradation of tankyrase offers a novel therapeutic strategy for β-catenin-driven cancers. Overall, this is a compelling study with significant translational potential.

      Strengths:

      (1) The manuscript presents a rigorous and well-executed study on a timely and impactful topic.

      (2) The biochemical and cellular characterization of the tankyrase degrader is thorough, and the comparative analysis with tankyrase inhibitors is insightful.

      (3) The finding that tankyrase stabilization by inhibitors may interfere with Axin function is novel and significant. It aligns with earlier observations (e.g., Huang 2009) that transient tankyrase overexpression can stabilize β-catenin independently of PAR domain activity.

      (4) The use of TNKS1/2 knockout cells expressing catalytically inactive tankyrase to demonstrate β-catenin inhibitory activity of the tankyrase degrader is elegant.

      (5) The finding that the tankyrase degrader has superior anti-proliferative effects in colorectal cancer models has important therapeutic implications.

      Weaknesses:

      (1) A key caveat is that the identified tankyrase degrader also targets GSPT1 for degradation. This raises the possibility that GSPT1 degradation may contribute to the observed β-catenin and tumor growth inhibition.

      (2) The authors address this concern reasonably by showing that DLD1 cells resistant to GSPT1 degradation remain sensitive to the tankyrase degraded.

      (3) To further strengthen this point, the authors might consider generating TNKS1/2 double knockout cells (e.g., in DLD1 or SW480 backgrounds) and demonstrating that the degrader loses its growth-inhibitory effect in these models. However, given the technical challenges of creating double knockouts in cancer cell lines, such experiments could be considered optional.

      We thank the Reviewer for the favorable feedback. The major concern is the collateral degradation of GSPT1. As the Reviewer noted, IWR1-POMA was able to suppress colony formation in DLD-1 cells resistant to a GSPT1/2 degrader (DLD-1R, Figure 6B and S9F), suggesting that TNKS but not GSPT degradation is responsible for growth inhibition.

      We also appreciate that the Reviewer brought it to our attention an important early observation of the TNKS scaffolding effects. Cong reported in 2009 that overexpression of TNKS induced AXIN puncta formation in a SAM but not PARP domain-dependent manner (PMID: 19759537, Ref. 12). We have added this reference to the introduction of TNKS scaffolding in the revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      The ADP-ribosyltransferase tankyrase controls many biological processes, many of which are relevant to human disease. This includes Wnt/beta-catenin signalling, which is dysregulated in many cancers, most notably colorectal cancer. Tankyrase is a positive regulator of Wnt/beta-catenin signalling in that it counters the activity of the beta-catenin destruction complex (DC). Catalytic inhibition of tankyrase not only blocks PAR-dependent ubiquitylation and degradation of AXIN1/2, the central scaffolding protein in the DC, but also tankyrase itself. As a result, blocking tankyrase gives rise to tankyrase accumulation, which may accentuate its non-catalytic functions, which have been proposed to drive Wnt/beta-catenin signalling. Most tankyrase catalytic inhibitors have shown limited efficacy and substantial toxicity in vivo. By developing tankyrase-directed PROTACs, the authors aim to block both catalytic and non-catalytic functions of tankyrase, aspiring to achieve a more complete inhibition of Wnt/beta-catenin signalling. The successfully developed PROTAC, based on the existing catalytic inhibitor IWR1, IWR1-POMA, induces the degradation of both TNKS and TNKS2, blocks beta-catenin-dependent transcription without stabilising the DC in puncta/degradasomes, and inhibits cancer cell growth in vitro. Mechanistically, this points to a scaffolding role of tankyrase in the DC, at least under conditions of tankyrase catalytic inhibition, in line with previous proposals.

      Strengths:

      The study clearly illustrates the incentive for developing a tankyrase degrader, namely, to abolish both catalytic and non-catalytic functions of tankyrase. By and large, the study achieves these ambitions, and the findings support the main conclusions, although the statement that a more complete inhibition of the pathway is achieved requires corroboration. The proteomics studies are powerful. IWR1-POMA constitutes a very useful tool to re-evaluate targeting of tankyrase in oncogenic Wnt/beta-catenin signalling. The paired compounds will benefit investigations of tankyrase scaffolding functions across many different biological systems controlled by tankyrase. The findings are exciting.

      Weaknesses:

      Although the results are promising and mostly compelling, the claim that the PROTACs provide "a deeper suppression of the WNT/β-catenin pathway activity" requires further corroboration, particularly at endogenous tankyrase levels.

      We thank the Reviewer for the encouraging and insightful comments. The major critique concerns whether TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels. IWR1-POMA reduced the level of cytosolic β-catenin more effectively than IWR1 in Wnt3A-stimulated HEK293 cells without protein overexpression (Figure 1D). IWR1POMA also suppressed STF activity more effectively than IWR1 in DLD-1 cells (Figure S8C) and reduced the expression levels of several WNT/β-catenin targets more effectively than IWR1 (Figure 1G and S8D). These results support that TNKS degraders can suppress WNT/β-catenin signaling more effectively than TNKS inhibitors at endogenous TNKS levels.

      There are also some other points that, if considered, would further improve the manuscript, as detailed below.

      (1) Abstract and line 62: Many catalytic tankyrase inhibitors tend to display toxicity, which is likely on-target (e.g., 10.1177/0192623315621192; 10.1158/0008-5472). This constitutes the main limiting factor for these compounds. An incomplete inhibition of Wnt/beta-catenin signalling may contribute to the challenges, but this does not appear to be the dominant problem. A more prominent introduction to this important challenge is probably expected by the field.

      A previous study showed that G007-LK, a selective TNKS inhibitor, exhibited weak efficacy and dose-limiting toxicity at 5‒30 mg/kg BID or 10‒60 mg/kg QD in various mouse xenograft models (PMID: 23539443, Ref. 28). Similarly, G-631, another TNKS inhibitor, also showed dose-limiting toxicity without significant efficacy at 25‒100 mg/kg QD in mice (PMID: 26692561, Ref. 60). However, other studies showed that G007-LK was well-tolerated at 200 mg/kg QD over 3 weeks in mice (PMID: 29316982, Ref. 61), and treating mice with G007-LK at 10 mg/kg QD over 6 months also improved glucose tolerance without notable toxicity (PMID: 26631215, Ref. 62). Importantly, basroparib, a selective TNKS inhibitor, was well tolerated in a recent clinical trial (PMID: 40964966, Ref. 64), and constitutive silencing of both TNKS1 and TNKS2 for 150 days in APC-null mice prevented tumorigenesis without damaging the intestines (PMID: 31337618, Ref. 8). We have included some discussion of the toxicity issue associated with TNKS targeting at the end of the Discussion section.

      (2) The authors do a good job in setting the scene for the need for tankyrase degraders. Their observations relating to the formation of puncta (degradasomes) being tankyrase-dependent are compatible with a previous study by Martino-Echarri et al. 2016 (10.1371/journal.pone.0150484): simultaneous silencing of TNKS and TNKS2 by RNAi abolishes degradasome formation. The paper is cited as reference 17, but only in passing, and deserves more prominence. (It includes an entire paragraph titled "Expression of tankyrases 1 and 2 is required for TNKSi-induced formation of axin puncta").

      Indeed, Henderson’s 2016 paper (PMID: 26930278, previously Ref. 17, now Ref. 18) shed important light on the role of TNKS scaffolding in the DC. However, whereas this study demonstrated that knocking down both TNKS1 and TNKS2 by siRNA prevented G007-LK to induce AXIN puncta, it concluded that “puncta formation requires both the expression and the inactivation of TNKS,” which is inconsistent with our observations that accumulation of either catalytically active or inactive TNKS can promote AXIN puncta formation. The function roles of TNKS scaffolding in the DC also remained unaddressed. We have included additional discussion of Henderson’s findings in the first paragraph the Discussion section.

      (3) Moreover, the scaffolding concept has been discussed comprehensively in other studies: 10.1111/bph.14038 and more recently 10.1042/BCJ20230230. There are also a few studies that focus on targeting the ankyrin repeat clusters of tankyrase to disengage substrates (10.1038/s41598-020-69229-y; 10.1038/s41598-019-55240-5) that illustrate the concept of blocking the scaffolding function. In that sense, the hypotheses are mature, and it is interesting to see some of them supported in this study. The authors could improve how they set their work into the context of these other efforts and proposals.

      Indeed, Guettler demonstrated in 2016 that TNKS scaffolding could promote WNT/β-catenin signaling, which forms the basis of the current work. Meanwhile, whereas there have been efforts to target the SAM or ARC domain to address TNKS scaffolding by Guettler and Lehtiö, our approach of targeting TNKS for degradation is complementary. We have included in the last paragraph of the Discussion section information on efforts to target the ARC or SAM domains as an alternative approach to suppress WNT/β-catenin signaling without promoting TNKS oligomerization (PMID: 31836723 and 32704068, Ref. 66 and 67).

      (4) In several places in the manuscript, the DC is referred to as "biomolecular condensate", at times even as a "classic example", implying that it operates through phase separation. This has not been demonstrated. In fact, super-resolution microscopy indicates that the puncta are not droplet-like (10.7554/eLife.08022), which would argue against the condensate hypothesis.

      Biomolecular condensates are membraneless cellular compartments formed by phase separation of biomolecules, regardless of their physical/material properties (PMID: 28935776 and 28225081, Ref. 22 and 23). Super-resolution microscopy studies by Stenmark (PMID: 26124443, Ref. 17) showed that AXIN, APC, TNKS, and β-catenin interacted with each other to assemble into membraneless complexes, wherein AXIN and APC formed filaments throughout the DC. Peifer has also summarized evidence that supports the condensate nature of the DC (PMID: 30782412, Ref. 9; see also PMID: 26393419). However, we acknowledge that testing the physical properties of reconstituted DC (for example, PMID: 34352208) with TNKS will provide a better understanding of the nature, for example liquid vs. gel, of these condensates.

      (5) It is beautiful to be able to use IWR1 and IWR1-POMA at identical concentrations for direct comparisons. However, this requires the two compounds to bind to tankyrase similarly well and reach the target to a comparable extent. How sure are authors that target engagement is comparable? Has this been evaluated?

      Using a BRET assay, we have confirmed that IWR1-POMA binds to TNKS1 with affinity comparable to that of IWR1. Details of this study is now included in the Results sections, and the data are presented in the Supplementary Information (Fig. S3E–G).

      (6) Figure 1F: It is not immediately apparent how IWR1-POMA shows more complete containment of Wnt/beta-catenin signalling. Most Wnt/beta-catenin targets lie close to the perfect diagonal, so I do not see how the statement "that IWR1-POMA controlled WNT/β-catenin signaling more effectively than IWR1" (in the legend of Figure 1F) is supported. Minimally, an expanded explanation would benefit the reader. Providing the colour-coding legend directly in the figure would help improve clarity. Also, the panel is very small and may benefit from a different presentation in the figure.

      We have updated Fig. 1F to include an inset of Quadrant III for improved clarity and readability. We have also moved Fig. S7C to the main text as Fig. 1G and added an expanded explanation for these figures.

      (7) Figure 2: The conclusion of a "deeper suppression" of signalling relies on overexpression of tankyrase in an otherwise tankyrase-null background. Have the authors attempted to measure reporter activity or endogenous gene expression without tankyrase overexpression, in Wnt3a-stimulated cells (in the context of a normal Wnt/beta-catenin pathway) or CRC cells at the basal level? Non-catalytic activity in a similar assay has previously been observed upon tankyrase overexpression (10.1016/j.molcel.2016.06.019). Whether or not there is a substantial scaffolding effect at endogenous tankyrase levels after tankyrase inhibition remains unconfirmed, and the PROTAC is a valuable tool to address this important question. The findings presented in Figure S7C and D go some way towards answering this question - these data could be presented more prominently, and similar assays could be performed in other cell systems.

      IWR1-POMA suppressed STF activity more effectively than IWR1 in APC-mut DLD-1 and SW480 CRC cells without TNKS overexpression (Fig. S8C). Similarly, IWR1-POMA provided a deeper suppression of STF signals in HeLa cells transfected with AXIN1 and β-catenin while expressing endogenous TNKS (Fig. 4G). These results suggest that inhibitor-induced TNKS scaffolding plays a significant role at endogenous TNKS expression levels. Following the reviewer’s suggestion, Fig. S7C is now Fig. 1G.

      (8) Line 237/238: "TNKS accumulation negatively impacts the catalytic activity of the DC (Figure 5D)" - the data do not show this. Beta-catenin levels are a surrogate readout for DC function (phosphorylation and ubiquitylation). Minimally, this requires rewording, with reference to beta-catenin levels.

      We have rephrased "TNKS accumulation negatively impacts the catalytic activity of the DC" as "TNKS accumulation negatively impacts the exchange of β-catenin in the DC."

      (9) Line 303-304: Beta-catenin is thought to exchange at beta-catenin degradasomes; this is clear from previous FRAP assays and the observation that phospho-beta-catenin accumulates in degradasomes upon proteasome inhibition (10.1158/1541-7786.MCR-15-0125). However, degradasome size hasn't, to my knowledge, been related to activity. Can this be clarified, please?

      We apologize for confusing β-catenin phosphorylation with β-catenin abundance. Here, we refer the catalytic activity of the DC to as the ability of the DC to promote β-catenin degradation rather than the kinetics of β-catenin phosphorylation. It is commonly observed that AXIN stabilization by TNKS inhibitors increases the DC size and reduces the β-catenin levels. As such, the induction of AXIN puncta by TNKS inhibitors is frequently used as an indicator of WNT/β-catenin pathway inhibition. However, we have found that, TNKS inhibition drives TNKS accumulation, which reduces the ability of the DC to promote β-catenin degradation. We agree that the DC only primes β-catenin but does not catalyze its degradation. We have revised our manuscript as follows: "increasing the local concentration of the DC components improves its 'effective activity'[50,51]."

      (10) There are previous hypotheses/proposals that the sensitivity of CRC cells to tankyrase inhibition correlates with APC truncation or PIK3CA status (10.1158/1535-7163.MCT-16-0578; 10.1038/s41416-023-02484-8). Have the authors considered expanding their cell line panel (Figure S7) to sample a wider range of cell lines, including some that are wild-type with regard to APC or Wnt/beta-catenin signalling in general? This would be a valuable addition to the work. Quantitated colony formation data could be moved to the main body of the manuscript.

      We have so far tested the effects of IWR1-POMA on the proliferation of DLD-1, SW480, HT-29, HCT116, and RKO cells (Fig. 6A and 6B). While a heterozygous Ser45 deletion in CTNNB1 confers resistance to IWR1-POMA, we did not observe sensitivity associated with APC or PIK3CA status. The ability of IWR1-POMA to suppress the growth of RKO cells expressing wild-type APC is consistent with a previous report that knockdown of both TNKS1 and TNKS2 stabilized PTEN to suppress cell proliferation and glycolysis in vitro and tumor growth in vivo (PMID: 25547115, Ref. 48) independently of the β-catenin pathway. We have added this new information as well as quantification of the colony growth results (Fig. S8A, S8B, S9A, S9F, and S9G) to the revised manuscript.

      (11) The manuscript only mentions toxicity (i.e., therapeutic window) in the last sentence of the Discussion section. As this is THE main challenge with tankyrase inhibitors (as mentioned above), can the authors expand their discussion of this aspect? Is there an expectation that PROTACs may be less toxic?

      As discussed above, evidence for on-target toxicity of WNT/β-catenin inhibition is mixed. Yet, the absence of dose-limiting toxicity for basroparib at doses up to 360 mg QD in human (PMID: 40964966, Ref. 64) is encouraging. PROTAC works by catalyzing target degradation, which is different from traditional catalytic inhibitors that require continuous target occupancy at a high level. It remains unclear whether the observed on-target toxicity of TNKSi is associated with TNKS accumulation at high doses, akin to the cytotoxicity induced by PARP1-trapping upon catalytic inhibition. We have included a brief discussion of the toxicity issue in the final paragraph of the Discussion section.

      (12) Figures 3, 4, 5A: For fluorescence microscopy experiments, can these be quantified, and can repeat data be included?

      We have included quantification data and replicate information for Fig. 3–5.

      (13) Figure 4, S6: An additional channel illustrating the distribution of cells (e.g., nuclei, cytoskeleton, or membrane) would be helpful for orientation and context for the AXIN1 signal.

      We have included cell outlines or nuclear staining for Fig. 3, 4, S6, and S7.

      (14) How were cytosolic fractions of cells prepared to assess cytosolic beta-catenin levels? This detail is missing from the methods.

      We have updated the Methods section to include additional details on the preparation of the cytosolic fractions of cells.

      Reviewer #3 (Public review):

      In this manuscript, Wang et al employ a chemical biology approach to investigate the differences between the enzymatic and scaffolding roles of tankyrase during Wnt β-catenin signalling. It was previously established that, in addition to its enzymatic activity, tankyrase 1/2 also plays a scaffolding function within the destruction complex, a property conferred by SAM-domain-dependent polymerization (PMID: 27494558). It is also known that TNKS1/2 is an autoregulated protein and that its enzymatic inhibition leads to accumulation of total TNKS proteins and stabilization of Axin punctae (through the scaffolding function of TNKS1/2), leading to rigidification of the DC and decreased β-catenin turnover. The authors surmised that this could, in part, explain the limited efficacy of TNKS1/2 catalytic inhibition for the treatment of colorectal cancers. To test this hypothesis, they evaluated a series of PROTAC molecules promoting the degradation of TNKS1/2 to block both the catalytic and scaffolding activities. They show that IWR1-POMA (their most active molecule) promotes more efficient suppression of beta-catenin-mediated transcription and is more active in inhibiting colorectal cancer cell and CRC patient-derived organoids growth. Mechanistically, the authors used FRAP to demonstrate that catalytic inhibitors of TNKS led to a reduced dynamic assembly of the DC (rigidification), whereas IWR1-POMA did not affect the dynamics.

      Overall, this is an interesting study describing the design and development of a PROTAC for TNKS1/2 that could have increased efficacy where catalytic inhibitors have displayed limited activity. Knowing the importance of the scaffolding role of TNKS1/2 within the destruction complex, targeting both the catalytic and scaffolding roles certainly makes sense. The manuscript contains convincing evidence of the different mechanisms of the PROTAC vs catalytic inhibitors. Some additional efforts to quantify several of the experiments and to indicate the reproducibility and statistical analysis would strengthen the manuscript. Ultimately, it would have been great to evaluate the in vivo efficacy of IWR1-POMA in an in vivo CRC assay (APCmin mice or using PDX models); however, I realize that this is likely beyond the scope of this manuscript.

      We thank the Reviewer for the helpful suggestions.

      I have some recommendations listed below for consideration by the authors to strengthen their study:

      (1) The title is slightly misleading, as it is already known that the scaffolding function of TNKS is important within the DC. The authors should consider incorporating the PROTAC targeting aspect in the title (e.g., PROTAC-mediated targeting of tankyrase leads to increased inhibition of betacat signaling and CRC growth inhibition).

      We have modified the title accordingly to "Targeting tankyrase scaffolding in the β-catenin destruction complex by PROTAC overcomes the limitation of catalytic inhibitors in cancer."

      (2) The authors should comment in the manuscript on the bell-shaped curve obtained with treatment of cells with the PROTACs (Figure S2C). This likely indicates tittering of the targets within a bifunctional molecule with increasing concentration (and likely reveals the auto-inhibition conferred by the catalytic inhibition alone).

      As suggested by the Reviewer, the bell-shaped dose-response likely originated from the formation of non-productive binary protein-ligand complexes at high PROTAC concentrations. We have added a sentence to clarify this unique behavior of PROTAC molecules.

      (3) The authors comment that using G007-LK as warehead was unsuccessful, but they do not show data. Do the authors know why this was the case?

      The structure-activity relationship of PROTACs is often unpredictable, as both the kinetics and thermodynamics of target and E3 ligase binding play important roles in promoting efficient target degradation. We have include data on G007-LK based PROTACs (Fig. S2D) in the revised manuscript.

      (4) Throughout the manuscript, the authors need to do a better job at quantifying their results (i.e., the western blots and the IF). For example, the degradation of TNKS1/2 in Figure 1D is not overly convincing. Similarly, the IF data in Figure 3 needs to be quantified in some ways. Along the same lines, the effect of IWR1-POMA treatments on the proliferation of cells and organoids should be quantified using viability assays... There is also no indication of how many times these experiments were performed and whether the blots shown are representative experiments. The quantification should include all experiments.

      We have included quantification of the immunofluorescence images, colony formation data, and Western blots in the revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) For clarity, can the authors use the official gene names, TNKS and TNKS2?

      We favor using TNKS1 and TNKS2 when referring to the protein for clarity and use TNKS for simplicity when referring to both proteins.

      (2) Line 92: The authors refer to TNKS2 "induction" - it remains unclear what is meant by "induction".

      We have changed "without induction" to "under basal conditions".

      (3) Can the authors please display molecular weight markers for Western blots throughout?

      (4) Line 144: The description "significantly more effectively" refers to Figure S5A, which shows a single, non-quantified Western blot. I don't think significance has been tested, and this statement should be reworded, or quantified aggregate data provided.

      We have added a Supplementary Information file showing molecular weight markers and quantification of Western blots.

      (5) Line 226: "plateaued at a much lower level" - can this be expressed more quantitatively in the text?

      We have included more quantitative information on the FRAP results.

      (6) Line 249: Can the authors repeat the cross-reference to Figure S7A here?

      We have repeated the cross-reference to the figures.

      (7) Line 266: The description of the experiment using the GSPT1/2 degrader CC-90009 would benefit from a brief recap of the purpose as not every reader will be familiar with this common PROTAC off-target. This is a very thorough analysis, though, and commendable.

      We have added background information on GSPT1 degradation to the revised manuscript.

      (8) Figure 1A: Can the number of repeats and the type of repeats be indicated, please?

      (9) Figure 2: Does n refer to biological or technical repeats?

      (10) Figure 5B, D: How many separate experiments are the data based on?

      (12) Figure S3D, S9A, D: number and types of repeats and the nature of the displayed data and error bars need to be included, please.

      (13) Figure S6B, S7B: I can see three data points, but it would still be helpful to state the number and type of repeats in the legend.

      (14) Figures S9A, S9D: There is value in showing the cumulative data from several repeats in the main figure (Figure 6, which currently is only qualitative) rather than the supplementary material.

      (15) Where single Western blots are shown, can the authors indicate how many experiments they are representative of?

      We have included the number of biological repeats for all data.

      (11) Figure S2C: For most graphs, the main response of interest occurs at low compound concentrations. The y-axis scale does not always help the reader to appreciate the effects, as the response seems small against the magnitude of the hook effect. Interrupting the y-axis as in the final panel may help, with y-axis scales consistent over all panels in the figure.

      We have updated Fig. S2C to emphasize on the degradation efficacy.

      (16) The authors may want to give further method details for some of their assays to facilitate replication of their experiments in the future. For example, the STF assay description is currently quite minimalistic. I assume the assay is fairly robust, though. Other details include cell media (general media details and specific additives and their concentrations in the 3D spheroid formation assay), etc. A general look at the methods section will likely be beneficial.

      We have updated the Methods section to provide more detailed experimental information.

      Reviewer #3 (Recommendations for the authors):

      (1) In Figure 2A, one of the most important findings of the manuscript is that IWR1-POMA induced promoted deeper suppression of beta-catenin-mediated transcription. This seems to be the case only at 3.2uM. Is it statistically significant? What are the data points on this graph? What are the error bars?

      We have included statistical analysis as Fig. S5G.

      (2) On Figure 2C and 2D, do the authors know why the TNKS20M1054V mutant is much better at promoting signaling than the TNKS1-PD ? Is it expression levels?

      It is indeed interesting that TNKS2-M1054V promoted significantly stronger WNT signaling than TNKS1-PD. The basis for its strong scaffolding effect is unclear.

      (3) In Figure 4C, the authors claim that when cells are treated with IWR1-POMA, AXIN1 is distributed diffusely throughout the cytoplasm. It appears that small punctae are visible.

      Quantitative analysis (Fig. 4F) suggest that the size of AXIN1 puncta upon IWR1-POMA is rather insignificant.

      (4) Label on Figure 1D has a spelling error TNKS1/2.

      Corrected.

    1. Author response:

      We thank the editors and reviewers for their thoughtful assessment of our manuscript, and for recognizing openretina as a valuable and timely resource for the retinal modelling community.

      We are especially glad that the reviewers appreciated the motivation of the project, the focus on standardization and reproducibility, and the potential of the platform to support systematic benchmarking and community-driven model development.

      We also understand the concerns raised. In the revision of the manuscript, we will strengthen the conceptual discussion of how predictive models, including the current “Core + Readout” models, can contribute to retinal neuroscience alongside more mechanistic and circuit-based approaches. This is a central matter for us, and one that some of us have recently addressed in a broader review on current trends in retina modelling (see https://doi.org/10.1016/j.visres.2026.108854). We will draw on this perspective to better articulate when predictive models are useful, where their limitations lie, and how openretina can provide infrastructure for comparing functional, normative and mechanistic models within a shared framework.

      We will also clarify the scope and limitations of the in-silico analysis methods provided within openretina. This will include a more explicit discussion of how MEIs, gradient-field analyses, and model-weight visualisations should be interpreted.

      Furthermore, we will add more information that will help the reader better judge different aspects of dataset quality, including, for example, spike-sorting or calcium-processing information and explainable-variance distributions. We note, however, that there are many subtle details about experimental workflows that are difficult to capture in compact indicators. In addition, we will make it clearer that the manuscript represents a snapshot of a living resource: The website, dataset cards, documentation, and repository will be the primary source of this information, especially as new datasets are contributed.

      Finally, we will of course address the technical clarifications raised by the reviewers, with the aim of making the manuscript more accessible overall.

      We are grateful for the reviewers’ constructive comments and believe that addressing these points will make our presentation of openretina clearer and more useful to the community.

    1. Author response:

      We thank the editors and reviewers for their thoughtful comments. Below, we list our provisional responses to the reviewers’ major points:

      On the rationale for CaMKIIα versus Thy1-driven stimulation and physiological relevance: We agree that we did not make clear the motivation for using CaMKIIα-driven stimulation, distinct from the Thy1-driven paradigm in our previous work (Williams et al., 2026). Using the Thy1 driver, both excitatory and inhibitory cells received direct theta drive. In contrast, CaMKIIα expression is largely restricted to principal neurons. Comparing these models lets us isolate a "driven I-cell" PING mechanism from the "E cell recovers first" mechanism relevant when interneurons are also directly driven.

      Regarding physiological relevance, Gonzalez-Sulser et al. (2014) found that septal GABAergic projections selectively and directly inhibit mEC interneurons, rather than exciting either principal cells or interneurons, implying that theta drive in vivo likely acts through rhythmic disinhibition of interneurons rather than direct excitation of any cell type. Neither the Thy1 nor the CaMKIIα paradigm reproduces this disinhibitory mechanism: both rely on excitatory optogenetic drive rather than rhythmic inhibition of interneurons, and replicating the natural drive (tonic excitatory tone plus rhythmic, interneuron-selective inhibition) is technically difficult in acute slices, which are largely quiescent without exogenous stimulation. We therefore view CaMKIIα and Thy1 as complementary approximations, each isolating a different circuit interaction. If forced to choose, we’d argue that the CaMKIIα is a better model of disinhibition of excitatory neurons. We will revise the Discussion regarding this point.

      On reproducibility of the voltage imaging findings: We thank the reviewer for this comment and agree that clarification is warranted.

      The voltage imaging dataset combines two levels of analysis with different sample sizes. The population-level firing and spike-correlation analyses (Fig. 5F–H) are pooled across multiple imaging sessions (n = 240 neurons). The spatial clustering analysis of subthreshold voltage correlations (Fig. 6, and the corresponding example traces in Fig. 5A–E) are drawn from a single representative recording session, as the reviewer correctly notes. We have voltage imaging data from 14 fields of view (1 FOV per slice) across 6 mice (240 neurons total; 3–41 neurons per FOV). In revision, we will extend the clustering and spatial-correlation analysis from Fig. 6 across sessions to assess whether the reported organization is reproducible, rather than relying on a single example. We will also revise the text to distinguish clearly which analyses are single-session versus pooled.

      On restricting the computational model of excitatory neurons to stellate cells: We modeled stellate cells as the excitatory population because they are the principal cells reciprocally connected to fast-spiking PV+ interneurons (Fuchs et al., 2016), the interneuron class most directly implicated in theta-nested gamma. Pyramidal cells, by contrast, are primarily connected via 5-HT3a-positive interneurons (Fuchs et al., 2016), with the exception of a subset of "intermediate" pyramidal cells that do show reciprocal PV+ connectivity. Our model, which captures the full measured heterogeneity of stellate cell and PV+ interneuron intrinsic properties and their reciprocal connectivity, is, to our knowledge, the most biophysically constrained implementation of this specific microcircuit to date. Incorporating the PV+-connected intermediate pyramidal population is a natural next step. Because this refinement, which requires more experimental data, is nontrivial and beyond the scope of this study, we will note this explicitly as a limitation of the current model in the revised Discussion.

      In vivo comparison (temporal/phase-locking): We agree that grounding our findings in existing in vivo data strengthens the study and will add these comparisons to the revision.

      Our whole-cell recordings reproduce the temporal organization in vivo and provide further insights into cell-type differences between the principal cells. All cell types were strongly phase-locked to theta, while gamma phase-locking declined across successive spikes, with stellate cells decoupling after the first spike and pyramidal cells after the second. This earlier decoupling in stellate cells may contribute to their weaker theta rhythmicity reported in freely moving rats (Ray et al., 2014; Tang et al., 2014). In extracellular recordings from behaving mice, spike-train cross-correlation identifies putative monosynaptic excitatory connections (1–4 ms) from principal cells onto fast-spiking interneurons (Latuske et al., 2015); the excitation-to-inhibition offset we measured is of comparable magnitude, here resolved as a synaptic-current delay in electrophysiologically classified cell types.

      We note that bursting and theta engagement have been assigned inconsistently across in vivo datasets. Bursty cells are preferentially classified as putative stellate by spikepattern classifiers (Latuske et al., 2015), while anatomically identified pyramidal cells are reported as the bursty, theta-rhythmic population in other work (Ebbesen et al., 2016). Because our cell-type assignments are based on subthreshold intrinsic properties (membrane sag, time constant) rather than spike patterning, our phase-locking results are independent of this classification ambiguity.

      In vivo comparison (spatial organization): We agree high-density silicon-probe datasets are the appropriate reference here. To our knowledge, the anatomical distribution of gamma-locked spiking in superficial mEC has not been characterized in vivo. The highest-density available recordings (Gardner et al., 2022) analyze population activity in the decoded state rather than tissue coordinates, do not examine gamma, and are restricted to grid cells. We regard the dissociation we observe between spatially clustered subthreshold input and spatially distributed spiking as a principal advance of the present study, and as a testable prediction for future high-density recordings.

      Ebbesen CL, Reifenstein ET, Tang Q, Burgalossi A, Ray S, Schreiber S, Kempter R, Brecht M. 2016. Cell Type-Specific Differences in Spike Timing and Spike Shape in the Rat Parasubiculum and Superficial Medial Entorhinal Cortex. Cell Reports 16:1005–1015. DOI: https://doi.org/10.1016/j.celrep.2016.06.057

      Fuchs EC, Neitz A, Pinna R, Melzer S, Caputi A, Monyer H. 2016. Local and Distant Input Controlling Excitation in Layer II of the Medial Entorhinal Cortex. Neuron 89:194–208. DOI: https://doi.org/10.1016/j.neuron.2015.11.029

      Gardner RJ, Hermansen E, Pachitariu M, Burak Y, Baas NA, Dunn BA, Moser M-B, Moser EI. 2022. Toroidal topology of population activity in grid cells. Nature 602:123–128. DOI: https://doi.org/10.1038/s41586-021-04268-7

      Gonzalez-Sulser A, Parthier D, Candela A, McClure C, Pastoll H, Garden D, Sürmeli G, Nolan MF. 2014. Gabaergic projections from the medial septum selectively inhibit interneurons in the medial entorhinal cortex. Journal of Neuroscience 34:16739–16743. DOI: https://doi.org/10.1523/JNEUROSCI.1612-14.2014, PMID: 25505326

      Latuske P, Toader O, Allen K. 2015. Interspike Intervals Reveal Functionally Distinct Cell Populations in the Medial Entorhinal Cortex. Journal of Neuroscience 35:10963–10976. DOI: https://doi.org/10.1523/JNEUROSCI.0276-15.2015

      Ray S, Naumann R, Burgalossi A, Tang Q, Schmidt H, Brecht M. 2014. Grid-Layout and Theta-Modulation of Layer 2 Pyramidal Neurons in Medial Entorhinal Cortex. Science 343:891–896. DOI: https://doi.org/10.1126/science.1243028

      Tang Q, Burgalossi A, Ebbesen CL, Ray S, Naumann R, Schmidt H, Spicher D, Brecht M. 2014. Pyramidal and Stellate Cell Specificity of Grid and Border Representations in Layer 2 of Medial Entorhinal Cortex. Neuron 84:1191–1197. DOI: https://doi.org/10.1016/j.neuron.2014.11.009

      Williams B, Vedururu Srinivas A, Baravalle R, Fernandez FR, Canavier CC, White JohnA. 2026. Fast spiking interneurons autonomously generate fast gamma oscillations in the medial entorhinal cortex with excitation strength tuning ING–PING transitions. eneuro ENEURO.0452-25.2026. DOI: https://doi.org/10.1523/ENEURO.0452-25.2026

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how Ca2+ levels inside the RGCs' mitochondria relate to whether these cells survive or die after injury to the optic nerve. The authors used advanced in vivo fundus live imaging techniques in mice to watch these changes unfold in real time, combined with genetic and drug-based tools to alter calcium flow into these compartments. Their central finding is a striking paradox: cells that naturally survive injury tend to have higher baseline calcium levels in these compartments, yet experimentally reducing calcium entry protects the broader population of cells from death.

      Strengths:

      The authors are applying sophisticated biosensors to track cellular chemistry in living animals over days and weeks. The tools and methods are creative and direct to detect the longitudinal RGC degeneration with mito-Ca2+ imaging. The topic and research aspect are novel and attractive. The results are significant, showing a clear relationship between the mito-Ca2+ regulatory machinery and cell survival.

      Weaknesses:

      The details of the mitochondrial-located signal of the Ca2+ sensor need to be further proved in the mito-matrix or between the mito-membranes. The study primarily describes a correlation and a surprising experimental outcome without fully explaining the underlying biological reasons for the paradox. While the evidence supporting the phenomenon is good, the mechanistic insight into why high calcium is linked to survival, or why lowering it helps after injury, remains limited.

      We appreciate Reviewer #1’s assessment of our manuscript. We also agree that we should have more clearly indicated that our mitochondrial Ca2+ sensor (Cox8-Twitch2b) is localized to the mitochondrial matrix. The Cox8-mitochondrial localization peptide is a well-established tool first identified in 1992 by Rizzuto and colleagues (Rizzuto, Simpson and Pozzan, 1992). We should have cited this work in our manuscript and will add it to our references. Further, as discussed in our submission, Cox8-Twitch2b has previously been validated for mitochondrial Ca2+ measurements in CNS axons (Witte et al., 2019). Thus, given the decades of use and characterization for this toolset, and the fact that we have pharmacological data supporting mitochondrial matrix localization of Cox8-Twitch2b, we do not feel it is strongly necessary to demonstrate mitochondrial matrix versus inner membrane space localization. However, we could attempt immuno-electron microscopy if this is deemed critical.

      We also agree that the mechanism by which reducing mitochondrial Ca2+ is protective would be satisfying and strengthen this study. But we feel it is beyond the scope of this project. It is likely manifold since mitochondrial Ca2+ impacts many vital cellular functions relevant to pathology including metabolism and apoptosis. We ultimately believe that an adequate investigation of these mechanisms would significantly slow down the dissemination of the core novel findings presented herein.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by McCraken and colleagues provides a continuation of their 2023 study (Cell Reports 42:113165) characterizing calcium regulation in retinal ganglion cells (RGCs) after acute optic nerve damage (a 10s crush using an intraorbital approach). This work is principally focused on how mitochondrial calcium stores change in both RGCs that are resilient and susceptible to injury. They report that resilient RGCs typically exhibited high calcium levels, but paradoxically, manipulating mitoCa2+ levels was more protective when the stores were reduced. Overall, regardless of susceptibility, mitoCa2+ levels decreased after injury, which is opposite to other reports that mitoCa2+ increases in degenerating neurons. The manipulation of mitoCa2+ was conducted both pharmacologically (Ru265) and by overexpression or knockdown of a primary calcium uniporter MCU. The evaluation of mitoCa2+ was conducted by using a reporter (Twitch2b) that was targeted to the mitochondria.

      Strengths:

      Many of the experiments are elegant and well-performed.

      Weaknesses:

      (1) Some experiments require further controls to validate that reagents are doing what they are intended to do.

      We agree with Reviewer #2 that our AAV manipulations of shMCU and MCU overexpression should be analyzed to verify how they alter mitochondrial Ca2+. To do this, we will co-express gene therapy vectors to lower and raise MCU expression with mito-Twitch2b biosensor and perform direct measurements of mitochondrial Ca2+. We will then determine if there is a relationship between gene expression level (inferred by mCherry intensity) and mitochondrial Ca2+ within samples, and if mean mitochondrial Ca2+ levels in treatments are higher or lower than mCherry reporter only controls.

      (2) Some findings can have alternate interpretations that are not considered.

      We will expand our Results and Discussion sections to broaden the interpretations of our data.

      (3) There is a broad generalization to the biology of all RGCs that may not be biologically relevant to different RGC subtypes.

      We agree that a more fine-grained understanding of RGC mitochondrial Ca2+ diversity would make interpretations of our data stronger. In our revisions, we will thus expand the number of RGC families in which we directly measure homeostatic mitochondrial Ca2+ levels. To do this, we will perform in vivo mito-Twitch2b measurements, collect and fix retinal wholemounts and immunostain for ON-OFF-direction selective RGCs using the marker CART and F-RGCs using the marker Foxp2. This will provide a complement of well-surviving RGC types (alpha and intrinsically photosensitive RGCs already examined) and poorly-surviving types.

      Reviewer #3 (Public review):

      Summary:

      Following previous work that demonstrated a relationship between higher homeostatic cytosolic calcium and lower retinal ganglion cell (RGC) apoptosis following injury to their axons, McCracken et al. investigated whether homeostatic calcium levels of the endoplasmic reticulum (ER) or mitochondria provide additional insights into the mechanisms by which calcium influences RGC survival. Their study reveals that homeostatic mitochondrial calcium shows a similar positive correlation with RGC survival. Despite that correlation, pharmacologic or genetic methods to lower mitochondrial calcium improved, rather than reduced, the survival of injured RGCs, while a genetic approach intended to increase mitochondrial calcium resulted in more RGC loss. These findings highlight the complexities of calcium regulation in modulating neuronal survival and raise important questions of how homeostatic levels of mitochondrial calcium affect stress responses that themselves can be either neuroprotective or neurodegenerative.

      Strengths:

      This study tackles an intriguing hypothesis that differences in calcium ion homeostasis in specific organelles may contribute to differences in survival of various RGC subtypes after optic nerve injury. This is a technically demanding question, and a primary strength of this work is its attention to, and meticulous reporting of, appropriate controls and, where applicable, seemingly contradictory results. Among these are careful evaluation of the effects of drug (or vehicle) delivery and genetic manipulations with and without injury and over extended time courses. The combination of thoughtful pharmacologic and genetic approaches makes for a thorough analysis of a challenging set of questions. The result is a study that provides a helpful perspective on the complicated roles that calcium, and especially mitochondrial calcium, can play across neuronal insults, neuronal types, and neuronal subtypes.

      Weaknesses:

      Given the paradoxical results, it would be helpful to have a clearer picture of how strongly the overexpression and knockdown of MCU altered the mitochondrial calcium levels. There may be potential for extraordinarily strong effects that would need to be tuned by using different shRNAs or promoters to more closely align with the observed differences between surviving RGCs and those that die. The investigation includes a relatively small number of resilient RGC subtypes, using the markers SPP1 and TBR2, raising questions of how generalizable the trend is between mitochondrial calcium levels and RGC resilience. The analysis and implications of Figure 3D might benefit from including not only the provided 50:50 split between "high" and "low" but also views of the data after splitting into thirds, fourths, and perhaps even fifths. The authors' inference that higher homeostatic calcium in more resilient RGCs may result in chronic mitochondrial stress is intriguing and worthy of more experimental investigation than is currently provided.

      We agree with the feedback from Reviewer #3, especially as it aligns with input from other reviewers. As these points agree with aspects above we will briefly reiterate our proposed revisions. We will validate the true effects on mitochondrial Ca2+ levels after gene therapy treatments by co-injecting AAV-mito-Twitch2b and AAV-shMCU or AAV-MCU. We will measure mitochondrial Ca2+ levels and correlate these levels with mCherry reporter expression intensity to determine the effect size of these treatments, and compare sample mean mitochondrial Ca2+ levels with those of mCherry control AAV.

      To further map the variance in homeostatic mitochondrial Ca2+ levels to RGC types we will perform in vivo mito-Twitch2b imaging, and then immunostain for ON-OFF-direction selective RGCs (CART) and F-RGCs (Foxp2), two poorly surviving RGC types.

      Lastly, we agree with Reviewer #3 that finer delineation between mitochondrial Ca2+ levels and their relationship to survival may be informative. We will split RGCs into smaller subgroups based on homeostatic mitochondrial Ca2+ levels and examine their survival outcome.

      Overall, we thank the Reviewers for their feedback, and believe the suggested changes will greatly strengthen our study.

      REFERENCES

      Rizzuto R., Simpson A.W. and Pozzan T. (1992). Rapid changes of mitochondrial Ca2+ revealed by specifically targeted recombinant aequorin. Nature, 358 (6384): 325-327.

      Witte M.E., Schumacher A-M., Mahler C.F., Bewersdorf J.P., Lehmitz J., Scheiter A., Sanchez P., Williams P.R., Griesbeck O., Naumann R., Misgeld T. and Kerschensteiner M. (2019). Calcium influx through plasma-membrane nanoruptures drives axon degeneration in a model of multiple sclerosis. Neuron, 101(4): 615-624.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Chen et al. describe metabolic phenotypes in Dp16 Down Syndrome mice, specifically the Dp(16)1Yey/+ mice - segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs. The group has performed metabolic phenotyping data in chow and high-fat diets, as well as undertaking a transcriptomic and metabolomic approach in tissues such as white and brown adipose tissues, liver, skeletal muscle, and hypothalamus to reveal both shared and sex-specific differences. The group describes sexual dimorphism in body weight, body temperature, food intake, and physical activity. Core shared features are insulin resistance, glucose intolerance, impaired lipid clearance, and dyslipidaemia in the Dp16 mice. They report tissue signatures of immune activation and a pro-inflammatory state, ER and oxidative stress, fibrosis, impaired glucose and fatty acid catabolism, altered lipid and bile acid profiles, and reduced mitochondrial respiration in Dp16 mice.

      Strengths:

      Overall, this is a good study with detailed, comprehensive data from an excellent group who have previously published on metabolic phenotyping of 2 other Down Syndrome mouse models. Although somewhat descriptive, it does certainly add to the current field and understanding of strengths and weaknesses of Down Syndrome mouse models, as well as identifying new features whilst strengthening previously suggested mechanisms.

      Weaknesses:

      Many aspects of this study have been described in other Down syndrome mouse models, though there are certainly aspects that are new. It would be useful if the authors could do a direct critique and comparison with previous publications in the area, utilizing the same Down Syndrome mouse model. There are also a few limitations in the number of animals used and the interpretation of the data that should be acknowledged.

      We have cited all relevant publications using Down syndrome mouse models. Regarding the Dp16 model, we have cited and discussed the only other study addressing metabolic aspects beyond body weight (Reference #138; PMID: 39803786). While that study reported glucose intolerance, insulin resistance, and defective insulin secretion, we did not measure pancreatic insulin content in our mice. Crucially, while the previous study found no sexual dimorphism, our study observed extensive sexual dimorphism in body weight gain, tissue-specific gene expression, and serum and liver metabolite changes.

      Regarding sample size, we used 6 mice per genotype per sex for transcriptomic and metabolomic analyses; this is constrained by the cost of performing these omics-type analyses. For mitochondrial respiration assays, we used 9–10 mice, and for most other in vivo and ex vivo assays, we utilized 12–15 mice, with some assays exceeding 20. We believe these sample sizes are robust and appropriate for this study.

      Reviewer #2 (Public review):

      Summary:

      Human DS is associated with metabolic dysfunction in humans, but the precise details of this have not been studied in detail. Here, the authors use a mouse model of DS to study systemic metabolic and transcriptional responses in key metabolic tissues to provide a deep understanding of the metabolic changes associated with DS. As part of his work, the authors also aimed to help inform the selection of a mouse model that best reflects the metabolic profile of DS, through comparison with other DS model metabolic data.

      The data presented in this model will be of interest to those in the field of metabolism. The immediate impact is unclear, but the breadth of data presented makes this a very useful resource.

      Strengths:

      (1) This work builds on other comprehensive analyses that the authors have performed in other DS mouse models.

      (2) The authors note common metabolic disturbances between male and female mice (e.g., insulin resistance) alongside clearly sexually dimorphic phenotypes (e.g., body weight). Studying both sexes in this context is important.

      (3) The authors have written the paper in a way that integrates a large number of observations well. There is complex data, and a high degree of sexual dimorphism. The study has generated a valuable and wide-ranging dataset comprising molecular, biochemical, and physiological data that will be useful for further, more mechanistic studies of metabolism in DS.

      (4) For specific observations, like the findings of altered body temperature in male and female mice, the authors undertake follow-up hypothesis-driven analyses of BAT mitochondria and specific hormones. Although these analyses do not explain the change in temperature, they ensure the study is not purely descriptive in nature.

      Weaknesses:

      (1) Assessing metabolism using dynamic testing is a strength. ITT, GTT and LTTs are included.

      (2) The dosing for GTTs, ITTs and LTTs was performed per body weight. But the mice under chow and HFD had different body weights. This may compromise the interpretation of the data. Further, ITTs are presented as percentage change, and this can be heavily influenced by baseline glucose measures. The changes appear quite dramatic, so can the authors plot the raw data instead?

      We have updated the ITT data plots to show raw glucose values instead of percentage change. Regarding the dosing, we believe basing it on body weight is an appropriate approach. This method is consistent with nearly all published rodent studies, as blood volume and metabolic tissues such as skeletal muscle and adipose tissue scale with body weight. Adjusting for weight prevents potentially erroneous conclusions. As for the diet groups, we compared WT and Dp16 mice only within the same diet group (Chow or HFD) rather than across different diets. We believe this ensures a valid and appropriate comparison for our study.

      (3) In addition, throughout the manuscript, it is not clear which tissues are the most dominant in disrupting metabolism. The ITT and GTT are composite measures across tissues. Tissue-specific analyses using a clamp technique or isolated tissues may provide more clarity here.

      Our data suggest a systemic metabolic deficit across multiple tissues, supported by tolerance tests, pan-tissue transcriptomic analyses, and liver and serum metabolite profiling. This is consistent with the triplication of genes in Down syndrome, several of which have known metabolic roles as highlighted in our discussion. We do not have evidence to support the role of a dominant tissue that contributes to the systemic metabolic dysfunction.

      Regarding the suggestion to use a clamp technique, we agree this would effectively determine whether insulin resistance is localized in the liver or skeletal muscle. However, we do not currently have the necessary equipment at Johns Hopkins University to perform these experiments. Conducting this work would require sending separate cohorts of WT and Dp16 male and female mice (on both chow and HFD) to an NIH-funded Mouse Metabolic Phenotyping Centre (MMPC). While we appreciate the value of this approach, we believe such labor-intensive experimentation falls beyond the scope of the present study.

      (4) One of the aims of the study was "to help inform the selection of mouse model that best reflects the metabolic profile of DS". The discussion does not contain a comparison between the previous work on different strains and relative to known human data.

      We chose not to include a comparison of different mouse models in the "Discussion" section because we previously highlighted the widely used Down syndrome models (Ts65Dn, Tc1, and TcMAC21) and their associated caveats in the "Introduction." Given the significant limitations of those models such as hypermetabolism in TcMAC21 and the presence of 41 triplicated protein-coding genes unrelated to human chromosome 21 we focused our in-depth metabolic analyses on the Dp16 model, which does not share these issues. We felt that restating this information in the "Discussion" would be unnecessarily repetitive.

      (5) Data availability. Raw metabolomic data should be made available.

      We have uploaded all metabolomics data, along with details regarding sample processing and data analysis, to the Metabolomics Workbench, an NIH-funded public repository. We have updated the "Methods" and "Data Availability" sections of the manuscript to include this information and the corresponding access link.

      Reviewer #3 (Public review):

      Summary:

      The article by Chen et al. describes the comprehensive metabolic profiling of DP16 mice, a Down syndrome model that carries a duplicated segment of the mouse chromosome syntenic to human chromosome 21. The authors note that this model is superior to previously used models, based on genetics, as ~65% of the chromosome 21 orthologues. The metabolic phenotypes also appear to be more consistent with those observed in humans with Down Syndrome. The study lays the groundwork for a more detailed genetic dissection of dosage-sensitive genes that contribute to the metabolic deficits observed in Down Syndrome.

      Strengths:

      There is an enormous amount of data in this manuscript, and the methods are described with adequate attention to detail. A strength of the manuscript is that both male and female mice were analyzed, so that concordant and discordant phenotypes were identified. Both males and females had evidence of insulin resistance. Transcriptomic and metabolomic data revealed impaired pathways for lipid metabolism, a pro-inflammatory state, reduced mitochondrial health and oxidative stress. Although the effects of a high-fat diet on weight gain were divergent, this diet caused worsened insulin resistance in both males and females.

      The discussion is excellent. Limitations of the study are well described. This reviewer does not identify any critical missing data.

      Weaknesses:

      It might have been helpful to have included blood pressure measurements, given the differences in 19-Nor-deoxycorticosterone. The discussion references several articles that describe sex-dependent differences in metabolic phenotypes in humans with Down syndrome, and it might have been helpful to state more explicitly whether these differences correlate with those observed here in mice.

      We appreciate the suggestion of blood pressure measurements. While we agree this is an important metric, given the metabolic focus of the present study and the significant volume of data already presented, we feel that blood pressure analysis is beyond the current scope and better suited for a follow-up study.

      Our study highlights sex differences in metabolic phenotypes in individuals with Down syndrome. While most published human studies focus on a limited set of parameters such as body weight, adiposity, serum lipoprotein profile, and fasting lipid/glucose levels our mouse data remain generally concordant with these findings. Beyond these standard measurements, we also observed substantial sex differences in pan-tissue transcriptomes as well as serum and liver metabolites.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) A major question is how these findings compare to data that have previously been published. For example, Lamantia et al. Bone 2024 and Dard et al. European Journal of Pharmacology 2025 both report no changes in body weight using the same Dp(16)1Yey Down syndrome mouse model? There is also a recent publication on liver dysfunction in Down Syndrome using the same mouse model. It would be useful to understand some of the similarities and differences of what is being reported by Dunn et al. Cell Rep 2026. In this assessment, there is an in-serum alanine transaminase (ALT) level, which was not the case in Dunn et al?

      For the Lamantia et al. Bone 2024 study, the authors only measured the body weights of Dp16 mice at 6 weeks of age. Our findings at 6 weeks align with Lamantia et al., showing no weight differences between Dp16 and WT mice of either sex (Fig. 2A and C). For the Dard et al. 2025 study, the authors only measured the body weights of Dp16 mice at 12 weeks old (P90) and observed no differences in body weights between genotype of either sex. At 12 weeks of age, we also did not observe body weight differences between Dp16 male mice and WT littermates (Fig. 2A). However, at 12 weeks of age, the Dp16 female mice clearly gained more weight compared to WT littermates (Fig. 2C). Our study tracked weights weekly from 6 to 16 weeks, revealing that while Dp16 females start at weights similar to WT littermates, the groups diverge over time. The reason for the difference between our findings and the single-point measurement by Dard et al. is unclear. Notable variables include:

      Mouse Sourcing: We obtained all cohorts and littermate controls from Jackson Laboratory, while Dard et al. bred their mice in-house.

      Diet: We used Envigo standard chow (catalogue # 2018SX). Dard et al. did not specify the chow used in their study.

      It remains uncertain whether these or other environmental factors contribute to the observed weight differences in female mice.

      In the Dunn et al study (Cell Rep 2026), they also performed metabolic analyses on serum and liver tissue in Dp16 mice. Consistent with their metabolic analyses of serum and liver tissue in Dp16 mice, we also observed the upregulation of multiple bile acids, including taurochenodeoxycholic, tauromuricholic, taurolithocholic, and lithocholic acids. Furthermore, our findings align with theirs regarding the transcriptomic and biochemical signatures of hepatic inflammation and fibrosis. However, there are two notable differences between our studies:

      (1) Liver Injury Markers: We observed an elevation in serum ALT, whereas the Dunn et al. study did not.

      (2) Sex Differences: We identified significant sex differences in the Dp16 transcriptome and metabolome. In contrast, Dunn et al. reported minimal to no sex differences and consequently combined male and female data for all analyses.

      Because Dunn et al. combined male and female data, a sex-stratified comparison between our results (separated by sex) and theirs was not feasible.

      (2) It would be important to understand trends in wild-type animals compared to Dp16 mice. For example, the sex specific and non-specific features - are any of these described in obesogenic wild-type animals fed on a high-fat diet? I.e., are the same features at play and just exacerbated in Dp16, or is this a Dp16-specific feature of systemic metabolism?

      Published literature indicates that WT females typically gain significantly less weight on a high-fat diet (HFD) than WT males. However, our data suggest that the weight gain patterns observed in Figure 6A and C are specific to the Dp16 genotype. Dp16 females gained substantially more weight during the first six weeks of HFD before WT females caught up. In contrast, Dp16 males showed robust initial weight gain comparable to WT controls, but their weight plateaued after seven weeks while WT controls continued to gain, leading to a clear divergence (Fig. 6A).

      Other metabolic parameters also appear specific to the Dp16 model. On a standard chow diet, WT mice of both sexes generally do not exhibit glucose intolerance, insulin resistance, dysregulated lipoprotein profiles (VLDL-TG), or an impaired capacity to handle lipid loads. We observed all of these features in our Dp16 male and female mice (Fig. 3). Furthermore, transcriptomic analyses of Dp16 mice on standard chow revealed gene signatures of inflammation, fibrosis, and oxidative stress that are absent in WT mice.

      When challenged with HFD, while WT mice typically develop glucose intolerance and insulin resistance, the triplicated genes in Dp16 mice significantly exacerbated this metabolic deterioration. This is reflected in the worsening of glucose control and insulin sensitivity observed in our tolerance tests.

      In summary, most of these metabolic features are specific to Dp16 mice on a standard chow diet and are further exacerbated when combined with a high-fat diet.

      (3) Food intake data is difficult to interpret when weight has already diverged, as bigger animals will eat more food. Hence, the higher food may be a consequence rather than a cause of the weight gain (data in Figure 1).

      The reviewer makes a valid point. Since physical activity and energy expenditure do not differ significantly between Dp16 females and WT controls (Fig. 2F), the observed increase in food intake may indeed contribute to the higher body weights in Dp16 female mice.

      To rigorously confirm this, food intake would need to be measured between 6 and 8 weeks of age, prior to the divergence in body weight. Unfortunately, we did not measure food intake at that earlier time point.

      (4) The n numbers seem to vary significantly. For example, the use of n=6 for metabolic studies is generally rather small and underpowered. For the seahorse data, another concern is the snap freezing of samples before Seahorse assessment. For example, snap freezing of samples has been shown to increase certain metabolites. Freeze-thaw tissues often show a significant reduction in optical redox ratio.

      Regarding the transcriptomics and metabolomics studies, we utilized N=6 mice per tissue per sex. While we agree that a larger sample size is always preferable, the high cost of OMICS analyses covering 144 RNA-seq and 48 metabolomics samples limited our capacity to increase this number. However, N=6 remains a robust and standard approach for these specific assays. For the majority of our other in vivo and ex vivo data, we employed a higher sample size of 12-15 mice per genotype per sex to ensure statistical rigour. For a few assays, we have sample size of over 20.

      Regarding the respirometry analysis, we acknowledge the limitations of using frozen tissue. We chose this method because it allowed us to perform Seahorse assays on multiple tissues from 9-10 mice, which is a significant sample size for this type of analysis. The alternative isolating mitochondria from fresh tissue would have restricted our ability to process multiple tissues from a large number of animals on the same day due to the length of the protocol. We believe this trade-off was necessary to maintain a high sample size across various tissues.

      (5) For oestradiol measurements, were the samples taken at the same times within the estrous cycle? This may affect the comparability of female Dp16 and WT mice?

      Regarding our protocol, blood samples were collected between 11:00 AM and noon, with food removed two hours prior. While we did not specifically monitor the oestrous cycle of the female mice, serum samples for both the Dp16 females and WT littermates were collected on the same day and at the same time to ensure comparability across the groups.

      (6) Body weight reduction and organ size reduction on an HFD are especially interesting. Could enhanced inflammation and fibrosis be the root cause of this? Are there other mouse models where this is the reason?

      On a high-fat diet, we observed a reduction in iWAT and gWAT fat depot weights in both male and female Dp16 mice, which is consistent with their lower overall body weights (Fig. 6 - figure supplement 3). Conversely, Dp16 females fed a high-fat diet showed increased heart and kidney weights. Despite their lower adiposity, the Dp16 mice on this diet exhibited greater insulin resistance and glucose intolerance (Fig. 7). This suggests that the worsening of glucose control is independent of obesity. While we observed signatures of inflammation and fibrosis, we do not yet have direct mechanistic evidence demonstrating that these factors causally impaired glucose and lipid metabolism.

      (7) The authors are circumspect throughout to avoid over-claiming, as the majority of data is observational. One exception: "Many bile acids serve as ligands for nuclear hormone receptors (e.g., FRX and TGR5) that control various aspects of glucose and lipid metabolism (74, 75), and extensive changes in circulating bile acids are contributing, at least in part, to the systemic metabolic phenotypes in Dp16 mice." The authors have not shown a direct link between bile acids and metabolism in this model. Please edit.

      We have edited the text accordingly.

      Minor:

      (1)"Most human studies at the whole-body level are limited to assessing the impact of trisomy 21 on food intake, adiposity, physical activity level, and energy expenditure in adolescents or adults with DS"

      While we were uncertain of the reviewer's specific intent regarding the suggested changes, we have rephrased the sentence for clarity.

      (2) It is somewhat surprising that T3 is elevated, although there are reports of T3 elevation in visceral obesity in humans (e.g., Sun Nam et al., Obes Res Clin Pract, 2010).

      We observed that T3 levels did not differ by genotype in mice of either sex when fed a standard chow (Fig. 2 - figure supplement 5). However, we noted elevated T3 levels in both male and female Dp16 mice on a high-fat diet (Fig. 6 - figure supplement 2). While increased T3 levels correlated with higher physical activity and a modest increase in metabolic rate in Dp16 females, this was not observed in males (Fig. 6). We do not currently have a clear explanation for these findings. Given that individuals with Down syndrome often present with hypothyroidism and lower T3 levels, this discrepancy may reflect a species-specific difference between humans and mice.

      (3) Please can the authors clarify the percentage gene coverage, as this is quoted as ~58% of Hsa21 gene orthologs or ~65% of the Hsa21 gene orthologs, where the same reference is used.

      We apologize for the confusion. The number of triplicated genes in Dp16 mice corresponds to ~58% of Hsa21 genes (PMID: 26765563). We have corrected the typographical error in the text.

      (4) "segmental duplication model carrying a majority of the triplicated Hsa21 gene orthologs" for this given percentage majority sounds too strong, and the use of percentage is recommended.

      We have modified the text accordingly.

      (5) It is puzzling that in female gWAT with 7 triplicated Hsa21 gene orthologs (Rbm11, Chodl, Cldn8, Sh3bgr, Igsf5, Itgb2l, and Tmprss2). Could this be a technical issue? Was the reduced expression quantified by RT-Q-PCR?

      We have examined the normalized counts in the RNA-seq data for the seven genes in question, and the results do not appear to be an artifact. The sample size for this data is six mice per tissue per sex. In general, we prefer utilizing raw and normalized counts from RNA sequencing because there is a linear relationship between transcript amount and raw counts that is independent of housekeeping genes. In contrast, RT-qPCR involves mRNA amplification and requires expression to be normalized by one or more housekeeping genes (such as GAPDH, β-actin, 36B4, or ubiquitin) under the assumption that their levels remain constant.

      (6) The difference in body temperature is of interest. In male Dp16 mice, there is an increase in core temperature and a lowering of body temperature in females. In female Dp16 mice, higher estradiol levels have been stated by the authors to contribute to lower body temperature and higher physical activity (69-72). I am uncertain if the references are all relevant, as some relate to ovariectomized animals. No explanation is given for males.

      We currently do not have an explanation for why Dp16 males on a chow diet exhibit higher core body temperature, while Dp16 females show lower body temperatures. Although elevated T3 levels can increase body temperature, we have ruled this out; our data indicates there are no significant differences in T3 levels between genotypes for either sex on a chow diet.

      (7) The authors find a higher percentage heart weight in Dp16 mice on HFD and comment in the discussion that this is in keeping with "high-fat diet-induced cardiac hypertrophy". From what I can see, no histology has been performed to justify this statement. Furthermore, it would be useful to understand which animals had congenital heart disease in the first instance.

      We have modified the text accordingly. Unfortunately, we do not have histology data on the heart to inform us on whether some of our mice had congenital heart disease.

      Reviewer #2 (Recommendations for the authors):

      (1) The authors should comment on the dosing method of glucose/insulin/lipid in the tolerance tests to acknowledge that differences in body weight may affect these tests. In addition, I encourage the authors to present ITT data as raw data, and not % change.

      In response to the reviewer’s comments, we have updated the ITT data plots to show raw data rather than percentage change. Regarding the dosing methodology, we maintain that basing dosage on body weight is appropriate. This approach is consistent with the vast majority of published rodent studies, as blood volume and metabolic tissues—such as skeletal muscle and adipose tissue—scale with body weight. Standardizing dose independently of body weight could lead to erroneous conclusions.

      (2) It would be useful for the authors to include a discussion on the likely specific tissue involvement in the whole-body metabolic disturbance. From my reading of the manuscript, there seems to be data suggesting functional and transcriptional dysfunction across most tissues, but do the authors suggest there is a dominant tissue in this regard?

      Due to the triplication of large number of genes on human chromosome 21, people with Down syndrome exhibit deficits across most organ systems (PMID: 32029743). Metabolic homeostasis also involves multiple tissues and cell types (adipose tissues, liver, skeletal muscle, pancreas, gut, hypothalamus, and immune cells). Most of the triplicated genes do express across these tissues. Our data indicate metabolic dysregulation across adipose tissues (white and brown), liver, skeletal muscle, and hypothalamus. Given the complex genetic perturbations of the Down syndrome mouse model, we do not think that there is a dominant tissue that contributes disproportionately to the systemic metabolic dysfunction phenotypes we observed in the Dp16 mice. Rather, we think that the metabolic phenotype is due to the combined deficits across multiple organs and tissues. As we do not have data to support the disproportionate contribution of any one tissue, we therefore did not speculate on the dominant contribution of any single tissue in the Discussion.

      (3) Related to this, muscle lipid is thought to be a major driver of muscle insulin resistance. Do the authors have measures of muscle lipid accumulation? This might be particularly interesting in the HFD models.

      Unfortunately, we did not measure lipid content in the skeletal muscle during this study. For the chow-fed mice, the entire gastrocnemius muscle was used for RNA isolation to perform RNA sequencing, and no tissue remains for additional analysis. Regarding the HFD-fed group, skeletal muscle was not collected at the termination of the study. As a result, we are unable to provide the requested lipid analysis data.

      (4) For mitochondrial analyses - do the authors have measures of total tissue mitochondria, and might changes in mitochondria abundance be driving some of these differences?

      For all our mitochondrial respiration analyses, we normalized the data to mitochondrial content as quantified by the MTDR assay (PMID: 32432379; PMID: 39704485). These results indicate that for a given amount of mitochondrial content, respiration as measured by the Seahorse assay is reduced in Dp16 mouse tissues, specifically in the BAT and liver.

      (5) To broaden the scope and interest, can the authors compare the transcriptional or metabolomic data to what has been found in non-DS insulin resistance (humans or mice), for example? This may help to highlight the key changes in metabolism that are causal for specific phenotypes.

      Overall, this is a comprehensive assessment of metabolism in a DS model.

      We appreciate the reviewer’s suggestion. However, given the vast number of published datasets on non-DS insulin resistance in both humans and mice, comparisons would yield varying results depending on the specific datasets selected. Consequently, we feel that such an analysis is beyond the scope of this study. We would like to highlight that many of the processes dysregulated in Dp16 mice as identified through our pan-tissue transcriptomes and metabolomes align with those frequently observed in non-DS insulin resistance. These include signatures of chronic low-grade inflammation, fibrosis, ER and oxidative stress, and impaired glucose and lipid metabolism.

      Reviewer #3 (Recommendations for the authors):

      It is slightly disconcerting that Figure 5 - Figure Supplements 2-5 are referred to in the text before the data in Figure 5 are discussed. It might make sense to indicate that the data are discussed further below (assuming that the authors do not wish to renumber these figures).

      We have fixed this issue raised by the reviewer.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript "A stable cryogenic fluorescence microscope for correlative super-resolution light and electron microscopy," the authors demonstrate a new cryogenic light microscopy design and characterize its temperature and spatial stability. The manuscript does a good job of reviewing the state of the field and highlights the need for improved cryogenic microscope stages. The system avoids challenges associated with vacuum-based designs, particularly vacuum transfer systems that can be difficult to engineer, while also showing minimal ice contamination and drift, which are the primary challenges associated with open cryostat systems.

      Strengths:

      The key strengths of the manuscript are the simple design and the significant level of detail provided in the description of the cryogenic stage. This represents a valuable step forward for the field by providing a home-built, non-vacuum stage design that others can emulate.

      We thank the reviewer for their positive assessment and strive to address the weaknesses they have constructively raised below.

      Weaknesses:

      There are only minor weaknesses or issues to address, which, if resolved, would strengthen the manuscript overall.

      (1) A key element of the design gets little attention, which is the plastic cap for the objective. It is not entirely clear to the reader how this is being used except as something of a thermal break between the cryogen environment and the objective, but there are some questions. Is the objective housing touching the plastic cap? Where is the front of the cap relative to the front objective lens? Is the front objective lens exposed to the cryogenic environment? Could the authors provide some 3D views of that in an SI figure? This would help clarify.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we will include a new supplementary figure (Fig. S2) providing detailed 3D views of the copper adapter, microscope objective, and plastic cap. The figure will show that the rim surrounding the front lens of the objective is covered by the plastic cap to provide thermal insulation between the objective housing and the cryogenic environment (Fig. S2b). We will also clarify that the front surface of the cap is levelled with the front objective lens to maintain the full working distance of the objective while allowing for axial movement of the z-stage. Finally, we will explicitly state that the front objective lens is exposed to the cryogenic environment (cold nitrogen gas).

      (2) The refilling system is not shown in the diagrams provided in Figure 1 and S1 in sufficient detail. How is the system mechanically coupled to the dewar on the microscope stage? Are there any concerns about coupling vibrations onto the table?

      To minimise vibrations arising from the nitrogen refilling pumps, the cryostat and liquid nitrogen tubing are mechanically decoupled from the microscope cage system, objective, translation stages, and sample. Specifically, the cryostat and nitrogen tubing are supported independently on a laboratory jack and surround the cage system without rigid mechanical contact. In the revised manuscript, we will update Fig. S1a, b to illustrate the liquid nitrogen tubing and refilling system more clearly. In addition, we will include a new supplementary figure (Fig. S3) to show the detailed cryostat design, refilling tubing, and temperature sensor position.

      (3) There is a description on page 6 that a rectangular aperture is used to align the excitation with the position and orientation of the sample. I know the authors are using this for excitation of the lamella, but without saying so in this text, it is confusing. I would consider stating that this is for future work involving excitation of lamella and then citing their preprint.

      We agree that the purpose of the rectangular aperture should be made clearer. In line with the suggestion from the reviewer, in the revised manuscript, we will briefly explain that the aperture is intended for selective illumination, such as in applications to cryo-FIB lamellae, and will cite our recent preprint describing this approach.

      (4) In Figure 2d, the z-drift is shown with the focus lock correction applied. This is highly relevant, but I also think it would be good to plot the z position plus the stage position in an SI figure. This will give a better idea of the mechanical stability of the system. Also, in this figure, I wonder if the authors could comment on the source of the jumps in lateral position. For example, just before 30 minutes. Lastly, I would make the lower plot have a tighter y-axis range. It is hard to see anything, hence the inset.

      We thank the reviewer for this suggestion. In the revised manuscript, we will include an additional supplementary figure (Fig. S4) showing the axial drift measured without focus-lock correction to illustrate the intrinsic mechanical stability of the microscope. We will also clarify that the periodic lateral displacement observed along the x-direction (approximately 300 nm amplitude with a period of ~22 minutes) arises from slight lateral repositioning accompanying z-stage stepping during focus-lock operation, likely due to mechanical coupling between the axes of the translation stage. We will revise the lower panel of Fig. 2d by reducing the y-axis range to improve data visibility.

      (5) The ice contamination looks minimal in Figure 3. I think it would benefit the manuscript to have lower magnification images as well, to show the level of ice contamination across a representative square. This would be good, but only if the authors have it in hand.

      We agree that this would be useful. In the revised manuscript, we will update Fig. 3 to include two additional low and intermediate-magnification cryo-EM images showing a representative grid square and a zoomed-in region of it, including a few grid holes. These images provide an overview of the ice contamination across a substantially larger field of view.

      (6) In Figure 4b, the y-axis is unclear. It looks like it has been normalized. Consider revising.

      The y-axis in Fig. 4b represents the localization rate (number of detected localizations per frame) within the selected ROI in Fig.4c and was not normalized. The values were calculated in SMAP by binning the localization frames into 100 temporal bins and dividing the number of localizations in each bin by the corresponding bin width, resulting in units of localizations per frame. Therefore, values close to 1 indicate approximately one localization detected per frame at that time point. To avoid potential confusion regarding the interpretation of this representation, we will replace this plot in the revised manuscript with a more explicit visualization showing the number of detected localizations per defined number of frames as a function of time (frame number) for the specific ROI shown in Fig. 4c.

      (7) A fluorescence intensity trace for the data shown in Figures 4c and f would be helpful to show the single-molecule behavior.

      In the revised manuscript, we will add fluorescence intensity traces corresponding to the single-molecule events shown in Fig. 4c and Fig. 4f to further demonstrate their single-molecule emission characteristics.

      Reviewer #2 (Public review):

      Summary:

      This manuscript reports the development of a cryo super-resolution fluorescence microscopy system. The authors demonstrate that they can achieve a mechanical and thermal stability that is sufficient to perform cryo-SMLM over the course of several hours. Focus instability is compensated for by tracking a fluorescent bead for its movement in the axial direction and adjusting the sample stage accordingly during data acquisition. Lateral instabilities are corrected after data acquisition. An enclosure around the microscope allows to significantly reduce ice contamination during cryo-SMLM imaging and sample transfer. The authors show an example of correlative cryo-SMLM and cryo-ET imaging achieved with their microscope system, which depicts the distribution of FtsZ-rsEGFP2 in E. coli.

      Strengths:

      The authors have designed a microscopy system for SR-cryo-CLEM, which achieves high stability while reducing complexity and costs substantially when compared to vacuum-insulated systems (e.g., Hoffman et al., 2020). They also provide software for controlling the microscope and data acquisition. This lowers the barrier for other labs to implement SR-cryo-CLEM into existing cryo-ET workflows. Reduction of ice contamination helps to increase throughput, which is currently one of the biggest bottlenecks for SR-cryo-CLEM.

      We thank the reviewer for their critical assessment, and for their suggestions below which we have used to improve the manuscript.

      Weaknesses:

      To correct for focus drift, the authors track a fluorescent bead in the far-red channel. This is possible for bacterial samples as used in this work, as beads can easily be introduced to surround the cells.

      Recommendations:

      (1) It is not discussed how this can be achieved in other samples than bacterial samples, such as lamellae in mammalian cells. Here, it would be much more difficult to introduce bright point-like markers with far-red fluorescence that would be distributed in the entire cell to capture at least one in the final lamella. Furthermore, it might be important to know for readers whether the far-red channel has to be sacrificed entirely for the focus correction.

      We thank the reviewer for highlighting this point. We agree that focus stabilization strategies for cryo-FIB lamellae are likely to differ from those used for the individual bacterial cell samples. For lateral drift correction, the presence of a single continuously detectable bright feature within the field of view is sufficient. Importantly, this feature does not need to be a fluorescent bead; any stable signal that can be continuously detected by the camera can serve as a suitable reference for drift correction. We will expand the Discussion to describe potential strategies for stable cryo-SMLM imaging, including the use of intrinsic sample or lamella features for autofocus, minimal fiducial-based approaches, and the practical implications of dedicating the far-red channel to focus stabilization.

      Furthermore, in the revised manuscript, we will include a new supplementary figure (Fig. S4) demonstrating the intrinsic axial stability of the microscope in the absence of active focus-lock correction. These measurements show that the system remains within the objective's depth of focus for a relatively long time, providing adequate stability for experiments in which far-red fluorescent fiducial beads are unavailable, such as cryo-FIB lamella imaging.

      (2) The authors show an application of SR-cryo-CLEM imaging of FtsZ-rsEGFP2 in E. coli. In the chosen correlative example (Figure 4d.f), no clear structure can be seen in the fluorescent images. The overview image (Figure 4d) shows no distinct signal in the cell, as it is shown for the non-correlative example in Figure 4a. The cryo-SMLM image (Figure 4f) does not show any ring-like features or accumulations of signals at the constriction site, as would be expected for a projecting along the optical axis. A clearer application example, which would show how increased resolution in cryo fluorescence microscopy enables resolving certain structural details or adds information not accessible in cryo electron tomography, would have strengthened the work. Particularly if taking into consideration that bacteria have a strong auto-fluorescence in the green range (Dahlberg et al., 2020), which could lead to high background or false positive localizations when using green fluorophores as labels.

      We thank the reviewer for this thoughtful comment. We agree that a correlative example displaying more pronounced structural features would further illustrate the capabilities of cryo-SMLM. However, the primary aim of the present work is the development and characterization of a robust cryogenic super-resolution microscope for reliable cryo-SMLM and correlative cryo-CLEM, rather than the demonstration of new biological applications. The utility of correlative cryo-SMLM/cryo-ET for resolving cellular structures has already been established in previous studies, including those employing rsEGFP2-labelled targets.

      The correlative dataset presented here is intended to demonstrate the compatibility of the microscope with cryo-CLEM workflows rather than to provide detailed biological insight. Moreover, the use of intact E. coli cells imposes inherent limitations on the ultrastructural information accessible by cryo-electron tomography; overcoming these limitations would typically require specimen thinning, for example, by cryo-focused ion beam (cryo-FIB) milling, which is beyond the scope of the present work.

      Regarding the concern about auto-fluorescence, elevated background fluorescence is not unique to bacterial samples or green fluorescent proteins but is a general consideration in cryo-SMLM that depends on the specimen and imaging conditions. While auto-fluorescence may reduce image contrast, it does not affect the conclusions of this work, which focuses on the design and performance of the microscope.

      (3) Access to CAD drawings (particularly for custom-made parts, such as cryostat or humidity enclosure) and a parts list is highly important for other researchers who would like to set up this SR-cryo-CLEM system in their own lab or institution. This is currently missing and, therefore, creating a hurdle for a wider adaptation of the technique.

      Thank you for this useful suggestion. In the revised manuscript, we will make available the complete SolidWorks CAD files for all custom-designed components, together with a comprehensive parts list and the full assembly corresponding to Fig. S1 as supplementary materials.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study examines how luminescence can be used to measure bacterial population dynamics during antimicrobial treatment by comparing it directly with optical density and colony counts. The authors aim to determine when luminescence reflects changes in population size and when it instead captures metabolic or physiological states induced by drug exposure. By generating parallel datasets under controlled conditions, the work provides a detailed view of how these three common measurements relate to one another across a range of drug treatments.

      Strengths

      The study is technically strong and thoughtfully designed. Measuring luminescence, optical density, and colony counts from the same cultures allows the authors to make clear and informative comparisons between methods. The data are compelling, and the analyses highlight both agreements and divergences in a way that is easy to interpret. The manuscript also succeeds in showing why these divergences arise. For example, the observation that filamentation and metabolic shifts can sustain luminescence even when colony counts drop provides valuable information on how different readouts capture distinct aspects of bacterial physiology. The writing is clear, the figures are effective, and the work will be useful for researchers who need high-throughput approaches to quantify microbial population dynamics experimentally.

      Weaknesses:

      The study also exposes some inherent limitations of luminescence-based measurements. Because luminescence depends on metabolic activity, it can remain high when cells are damaged or unable to resume growth, and it can fall quickly when drugs disrupt energy production, even if cells remain physically intact. These properties complicate interpretation in conditions that induce strong stress re-sponses or heterogeneous survival states.

      In addition, the use of drug-free plates for colony counts may overestimate survival when filamented or stressed cells recover once the antibiotic is removed, making differences between luminescence and colony counts harder to attribute to killing alone. Finally, while the authors discuss luminescence in the context of clinically relevant concentration ranges, the current implementation relies on engineered laboratory strains and does not directly demonstrate applicability to clinical isolates. These limitations do not detract from the technical value of the work but should be kept in mind by readers who wish to apply the method more broadly.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Luminescence limitations. We agree that the lack of a direct link between light intensity and a population property such as biomass or cell number is the main limitation of the luminescence method. To further emphasise this, we have expanded the Discussion in the revised manuscript.

      Drug-free plates. The use of drug-free plates is intentional. As we measure a time series, the question at each point is how many cells are alive at each time point. Cells that are stressed but viable at time t contribute correctly to the count at t. How long they survive under the respective treatment is captured by the subsequent timepoints.

      Filaments. Recovery of plated filamented cells should not inflate this estimate. A single plated filamentous cell is expected to yield either zero (death before division) or one single colony, regardless of in how many parts it separates, as all descendants are part of the same cluster. However, if the cells divide before plating, CFU can overestimate survival. Having that said, we have no indication that this occurred in our experiments, since in all observed discrepancies, CFU-based estimates were equal to or lower than those obtained from luminescence and the time cells spent in dilution was kept short.

      Clinical applicability. We agree with the reviewer that the method is not practical for ad-hoc pharmacodynamic studies of clinical isolates. What we instead provide is an E. coli-based model system to explore clinically relevant treatment conditions, which we address in the revised manuscript. We believe that constructing analogous bioluminescent model strains in other clinically relevant species would be a valuable direction for future work.

      Reviewer #2 (Public review):

      Summary:

      This preprint proposes luxCDABE-based luminescence as a high-throughput alternative (or complement) to CFU time-kill assays for estimating antimicrobial rates of population change at super-MIC concentrations, by comparing luminescence- and CFU-derived rates across 20 antimicrobials (22 assays) and attributing divergences primarily to filamentation (luminescence closer to biomass/volume than cell number) and changes in culturability/carryover (CFU undercounting viable cells).

      Strengths:

      The authors do not merely report discrepancies; they experimentally validate the biological causes. Specifically, they successfully attribute the slower decline of luminescence in certain drugs to bacterial filamentation (maintaining biomass despite halted division) and the rapid decline of CFU in others to loss of culturability or carryover effects.

      The inclusion of 20 antimicrobials spanning 11 classes provides a robust dataset that allows for broad categorisation of drug-specific assay behaviours.

      The study critically exposes flaws in the “gold standard” CFU method, specifically regarding antimicrobial carryover (demonstrated with pexiganan) and the potential for CFU to overestimate cell death in the presence of VBNC (viable but non-culturable) states induced by drugs like ciprofloxacin.

      The use of chromosomal integration for the lux operon to minimise plasmid copy-number effects and the validation of linearity between light intensity and cell density establish a solid technical foundation.

      Weaknesses:

      The study is conducted exclusively using Escherichia coli. While E. coli is a standard model organism, the paper claims to evaluate luminescence as a generalisable high-throughput tool. Many of the discrepancies observed are driven by filamentation. However, distinct morphological responses occur in other critical pathogens (e.g., Staphylococcus aureus does not filament in the same way).

      The authors propose that luminescence data can be corrected using microscopyderived volume data to better align with CFU counts. The primary appeal of luminescence is high-throughput efficiency. If a researcher must perform timelapse microscopy to calculate cell volume changes to “correct” their luminescence data, the high-throughput advantage is lost.

      The paper argues that for ciprofloxacin, CFU underestimates viability because cells remain intact and impermeable to propidium iodide. While the cells are metabolically active and membrane-intact, if they cannot divide to form a colony (even after drug removal/dilution), their clinical relevance as “living” pathogens is debatable.

      Some other comments:

      The use of a population dynamical model to simulate filamentation effects is excellent. The finding that light intensity tracks volume ($\psi_V$) better than cell number ($\psi_B$) is a key theoretical contribution.

      The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      The use of bootstrapping to estimate rate distributions is appropriate and robust.

      Conclusion:

      Muetter et al. provide a compelling argument that luminescence is a reliable, highthroughput alternative to CFU for super-MIC investigations, particularly when the quantity of interest is biomass. The paper effectively warns researchers that discrepancies between CFU and luminescence are often biological (filamentation, VBNC) rather than methodological failures.

      We thank the reviewer for reading our paper thoroughly and for the helpful feedback.

      Generalisability. We agree that the alignments and divergences reported for specific drugs may not transfer directly to other species, which may elongate differently (e.g. cocci) or show different physiological responses to treatment. Constructing analogous model strains — for example based on S. aureus to cover a broader range of morphologies and clinically relevant species would therefore be an interesting follow-up project, and we have adjusted the Discussion to make this clearer. We nevertheless believe that the broader conclusions (larger cells emit more light) of the paper likely hold across species.

      Volume correction. We agree that requiring microscopy would undermine the high-throughput advantage of the luminescence assay. It was not our intention to propose this as a practical approach, nor to imply that the luminescence signal needs a correction. Taken on its own, the signal can be interpreted as the cumulative metabolic output of the population, which is closely linked to biomass, and that measure is valuable in itself for many applications. We used the volume correction only to demonstrate that luminescence tracks biomass more closely than cell number: by adjusting the luminescence distribution with the measured volume change, it moves towards the CFU distribution. We have revised the Discussion to prevent this from being misunderstood as a required step.

      Culturability vs. clinical relevance. We agree that the dynamics of culturable cells are highly relevant, especially in a clinical context. Our aim was to explain the observed differences between CFU and luminescence by highlighting that culturability and viability are not always identical, without implying that one measure is inherently superior to the other — we leave it to the reader to decide which metric best suits their needs.

      Linear elongation. The model assumes linear elongation for mathematical convenience, which, depending on the specific strain and drug mechanism, could be incorrect. Its purpose is to demonstrate that a shift of the mean cell volume to a new, higher equilibrium under treatment can cause an initial peak in the luminescence signal despite a declining population. This remains true for non-linear elongation models, though the shape, height and position of the peak may change. We have adjusted the Results to make this clearer.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors present luminescence as a practical measurement of population decline under antibiotic exposure. One aspect that could be clarified is how the method behaves when tolerance arises from phenotypic heterogeneity, such as the presence of small, metabolically quiet survivors. Because luminescence reflects metabolic activity and biomass, the signal will be dominated by metabolically active cells, making rare tolerant subpopulations difficult to detect. A short discussion of how luminescence performs in these heterogeneous scenarios, and whether complementary assays are needed to capture long-lived tolerant cells, would strengthen the manuscript.

      Yes, that is a valid concern and we thank the reviewer for raising this point.

      Heterogeneity in cell-specific luminosity alone does not bias population-level rate estimates. A bias can arise, however, when specific luminosity correlates with a second factor — most importantly, the decline rate under treatment.

      We agree with the reviewer’s suggestion that brighter cells plausibly die faster than tolerant, metabolically quiet ones. When one subpopulation dominates the light signal, we expect minimal bias, as the rate estimate primarily reflects that subpopulation. However, in a transition phase when both subpopulations contribute roughly equally to the light signal, luminescence likely overestimates the decline.

      We added a corresponding caveat to the Discussion (lines 581–583).

      (2) The manuscript shows that filamentation can influence ψ_I by altering biomass and metabolic activity independently of cell number. However, antibiotic exposure can also trigger other stress responses and metabolic shifts that change energy fluxes, redox balance, and biosynthetic activity. Since luminescence depends on metabolic state and substrate availability, these additional physiological transitions may also affect ψ_I in ways not directly tied to birth or death processes. It would be useful to comment on whether such responses, beyond filamentation, are likely to influence luminescence dynamics across different drug classes or treatment conditions.

      We thank the reviewer for raising this point and agree that there is no biological law strictly linking luminosity to a single population property such as biomass or cell number, and changes in the metabolism most likely affect Ψ<sub>I</sub> as well.

      Transitioning to a new metabolic steady state biases Ψ<sub>I</sub>; once the new steady state is reached, however, the rate estimate should no longer be affected.

      Looking across drug classes, drugs that primarily lyse cells (polymyxins and, to a lesser degree, beta-lactams targeting PBP1) did not show noticeable deviations between Ψ<sub>I</sub> and Ψ<sub>CFU</sub>, and — perhaps counterintuitively — neither did ribosome-inhibiting drugs.

      For the remaining cases, we were able to attribute part of the discrepancy between Ψ<sub>I</sub> and Ψ<sub>CFU</sub> to changes in biomass or loss of culturability, though drug-induced metabolic changes may also contribute to the residual differences.

      We clarify this in the Discussion (lines 569–579).

      (3) The authors quantify survival using colony counts on drug-free medium. Because filamentation can be a reversible state that persists during antibiotic exposure, plating on drug-free medium may capture recovery potential rather than in-treatment viability. Filamented or stressed cells that cannot divide in the presence of a drug may nevertheless form colonies once the drug is removed. Clarifying how this recovery step affects ψ_CFU would help readers interpret differences between luminescence-based and colony-based measurements, especially in cases where transient tolerant states are present.

      We thank the reviewer for raising this point.

      Our CFU assay estimates the number of culturable cells at each time point; the rate Ψ<sub>CFU</sub> is then inferred from how this number changes across time points. Plating on drug-free medium is intentional, as it maximises the probability that a culturable cell is detected at each snapshot. Whether those cells would have continued dividing or died under continued treatment is captured by the subsequent time points.

      Filamentation interacts with the probability of colony formation in several, partly opposing ways:

      (1) It can increase the death rate, as for ceftazidime and cefepime, which is part of the kill effect captured by Ψ<sub>CFU</sub>;

      (2) Entanglement between filaments may reduce the number of colonies per plated bacterium;

      (3) Conversely, if a filament divides upon drug removal, its fragments form a cluster that — stochastically — is very likely to produce one (but not multiple) colony.

      The only scenario in which CFU could overestimate bacterial density is if a filament separates into individual cells in the liquid phase before plating; we have no indication that this occurred in our experiments.

      We addressed this concern in our response to the public comment.

      (4) A brief discussion comparing luminescence to fluorescent reporter systems could be helpful. Fluorescent proteins typically require a chromophore maturation step before becoming detectable, which introduces a delay between the underlying cellular event and the appearance of the signal. In contrast, as far as I understand, lux reporters emit light immediately once the enzymatic components and substrates are present, without a maturation stage. Highlighting this distinction may help readers understand why luminescence is well-suited for tracking rapid changes in population physiology under antibiotic exposure. However, the manuscript also notes that luminescence can lag slightly behind very rapid killing (particularly for AMPs), but the temporal dynamics of signal shutdown are not explored in detail. Because lux reflects metabolic activity rather than viability, a short delay between irreversible damage and the loss of light is biologically expected. It may help readers if the authors could expand on the mechanism underlying this delay in order to clarify when ψ_I is likely to track true biomass decline and when residual metabolic activity might mask early killing events.

      On fluorescent reporters:

      We thank the reviewer for this suggestion.

      Under some conditions, change rates can also be measured using fluorescence, provided the number of fluorescent molecules per bacterium remains constant. This requires a balance between production, maturation, degradation and dilution, which is only established if the growth rate and conditions remain constant over a sufficiently long period (typically hours).

      For measuring population decline, however, the key issue is that cell death does not inactivate fluorescent proteins: once matured, they emit independently of the cell’s metabolic state and decay only with the protein’s half-life, which is typically slower than the kill rates of interest.

      We added a clarification to the Introduction (lines 58–60).

      On the lux signal lag:

      We thank the reviewer for raising this point. The short lag between luminescence and CFU decline could in principle arise from two mechanisms: (i) luminescence declining more slowly than the actual cell number (residual light from dead cells), or (ii) CFU declining more steeply than the actual cell number (damaged but still viable cells failing to form colonies).

      Mechanism (i) splits into two sub-cases:

      (i.a) Dead but impermeable — the lux reaction could in principle continue for a short while if enough components are retained in the cell. However, a metabolically active, impermeable cell is difficult to classify as dead in the first place, making this scenario conceptually awkward.

      (i.b) Dead and permeable (lysed) — the lux components dilute into the medium, and by mass-action the reaction rate should drop rapidly (though not instantly). Any residual signal after lysis should therefore be short-lived.

      Mechanism (ii) — damaged (e.g. permeable) cells may be particularly sensitive to plating on agar (e.g. due to oxidative stress), resulting in a declining probability of colony formation.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse, making (i.b) and/or (ii) the likely explanations. Based on our experimental data, we cannot distinguish between these possibilities and therefore limit ourselves to reporting the observed discrepancy.

      We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded the discussion there.

      (5) In lines 85–89, the authors state that “high-throughput OD and luminescence measurements at sub-MIC concentrations provide valuable insights into drug effects on growth rates, [but] the super-MIC range is clinically more relevant,” and they present luminescence as a way to investigate super-MIC population dynamics. While super-MIC behaviour is indeed important for pharmacodynamics and resistance evolution, it is not clear that the specific luminescence implementation used here has direct clinical relevance. The study relies on a chromosomally integrated reporter in a laboratory strain, and the manuscript does not demonstrate that this approach can be applied to clinical isolates or diagnostic workflows. It may be helpful to moderate the claim of “clinical relevance” and frame the method more clearly as a high-throughput experimental tool that can inform clinically relevant questions, rather than as an assay ready for clinical application.

      We agree and have moderated the framing accordingly (lines 100–104).

      Reviewer #2 (Recommendations for the authors):

      (1) The conclusions regarding “biomass vs. cell number” may not apply equally to non-rod-shaped bacteria or species with different stress responses. The authors must explicitly discuss this limitation in the Discussion.

      The broad conclusion that bigger cells emit more light likely holds across morphologies, since it rests on the principle that more cellular material means more metabolic activity and therefore more light. The quantitative relationship between cell size and luminosity, however, may differ across species, shapes and conditions, for two reasons. First, chromosome copy number: whether drug-induced morphological changes are accompanied by chromosome replication and therefore an increase in lux operon copy number — varies across species and drug mechanisms. Second, the surface-to-volume ratio likely modulates mass-specific metabolism; some morphological changes preserve it (e.g. purely lateral elongation) while others do not.

      The more specific conclusions about which drug classes produce alignment or divergence between CFU and luminescence may also not transfer directly, as drug mechanisms can act differently across species.

      We already note this limitation in the Discussion (lines 594– 597) and have expanded the wording there.

      (2) The manuscript should clarify that luminescence is a superior metric for biomass without correction, rather than framing the volume correction as a necessary step to mimic CFU. The divergence should be embraced as a feature (biomass tracking), not a bug that needs fixing via labor-intensive microscopy.

      We agree with the framing and will make it clearer; it was actually our intention to clarify which method does what, rather than judge one as better or worse.

      We removed the “correction” sentence from the Discussion to make this clearer.

      (3) The authors should refrain from definitively stating CFU “underestimates” viability and instead use more precise terminology, such as “reproductive capability” vs. “metabolic integrity.”

      We agree with the reviewer that measuring culturability is a property, not a flaw, of CFU. Our intention was to emphasise that when CFU is used as a proxy for viability (which it often is), it can yield lower values than the actual number of survivors. We tried to make that distinction explicit in the manuscript (e.g. in lines 317–322).

      We would also like to note that in the case of antimicrobial carryover, CFU can genuinely underestimate culturability itself, not only viability.

      Regarding the suggested reproductive capability vs. metabolic integrity framing: we agree that metabolism and luminescence are closely linked. What held us back from drawing that link directly is that metabolism is hard to quantify, being the cumulative output of a diverse set of processes.

      (4) The model assumes linear elongation. The authors should briefly comment on whether this holds true for the specific drug mechanisms tested (e.g., PBP inhibition vs. DNA gyrase inhibition).

      Linear elongation is a mathematically convenient simplification whose only purpose in the model is to allow the population to converge to a new equilibrium volume under treatment. Assuming constant volume-specific luminosity, we showed that this produces an initial peak in light intensity before the signal declines in parallel with Ψ<sub>B</sub>. The exact shape, height and position of this peak depend on the volume growth model used, but the qualitative pattern — peak followed by parallel decline — holds for other growth models as well. We now clarify this in lines 230–235.

      (5) The authors suggest the carryover effect is due to a delay between cell death and cessation of luminescence. This “lag time” is a critical physical constraint of the lux system (likely related to ATP depletion or enzyme decay) and should be quantified or discussed in more detail as a fundamental “speed limit” for the assay.

      The origin of the lag between luminescence and CFU is an interesting question, but one we cannot definitively answer. We can, however, discuss the potential mechanisms:

      A dead but impermeable cell could in principle continue to emit residual light for some time. We note, though, that calling a metabolically active, impermeable cell “dead” is a question of definition we would rather not discuss here.

      In our case, the discrepancy was observed specifically for pexiganan, where cells can be assumed to lyse. Under lysis, the lux components dilute quickly into the medium, and by mass-action the reaction rate should drop rapidly — though not necessarily instantaneously.

      A plausible alternative to a delayed cessation of the light signal is that the probability of colony formation drops rapidly after permeabilisation, for example because permeable cells are sensitive to oxidative stress when plated on agar.

      Based on our experimental data we cannot distinguish between these mechanisms, so we limit ourselves to reporting the observed discrepancy. We have moved the interpretation from the Results to the Discussion (lines 523–542) and expanded on the candidate mechanisms there.

      Additional revisions

      Beyond the changes prompted by the reviewers’ comments, we made the following revisions to the supplementary information:

      We corrected the Λ matrix (converted row 2, col 4 from 0 → 2)

      We removed the line numbering

    1. Author response:

      Reviewer 1 (Public review):

      Summary:

      The authors describe a clever genetic system based on rapamycin-inducible expression of a beta-galactose reporter. The authors compare this spectrophotometer-based readout to the parasite reduction rate version 2 (PRR v2) recently described by some of the same authors and based on incorporation of [3H]-hypoxanthine. The results are generally comparable, with some differences for slower-acting compounds. The authors report that this format is better suited for higher-throughput studies and requires less time to quantify the time-dependent onset of parasiticidal action compared with the PRR v2.

      Strengths:

      This is a very well-executed and well-described body of work with a comprehensive set of analyses.

      Weaknesses:

      The authors should revise their text to also describe other methods used to quantify parasite growth. This method saves time compared to the PRR v2 but is too complex for simple screening of antiplasmodial activity of agents tested alone. Its value lies in assessing the speed of action of compounds tested in combination.

      We thank reviewer 1 for the supportive feedback and for raising some important points.

      Many antimalarials have quite specific times of action. Are these MULT-i<sup>2</sup> assays, and the comparator PRR v2 assays, conducted with asynchronous cultures? This should be described in the methods and referred to in the text (apologies if I missed some references).

      We thank the reviewer for this important comment. Both, the MULT-i<sup>2</sup> and PRR v2 assays were performed using asynchronous parasite cultures. This information is included in the Methods section together with the relevant references. To improve clarity, we will also explicitly state this in the main text.

      The authors correctly state that flow cytometry-based readouts, such as with MitoTracker alone, can limit throughput and that MitoTracker alone can produce spurious results. The authors should cite work from other labs that combine MitoTracker with a nuclear dye, such as SYBR Green I. I think others have also been used, such as YoYo-1, which overcomes the limitations of using MitoTracker alone. Also, many labs use a nuclear dye such as SYBR Green I in a spectrophotometer-based format that enables rapid processing of plates at scale (96, 384, or even 1536 wells per plate). Luciferase-based screens have also been used in large-scale screening campaigns. The introduction should cite these various approaches, especially as the MULT-i<sup>2</sup> method is quite a complex screen with an initial period of drug exposure (up to 3 days) followed by a five-day phase initiated by rapamycin addition to induce expression of the beta-gal sensor.

      We thank the reviewer for this helpful suggestion. In the Introduction we will mention and describe alternative approaches for assessing parasite viability. This will also include the work by Maiga et al., which combines MitoTracker with a nuclear dye to improve the reliability of flow cytometry-based readouts. We will revise the text to explicitly mention the use of dual straining to make this discussion more explicit.

      We agree that several additional methods, such as luciferase-based reporter systems, have been successfully applied in antimalarial screening. However, these approaches are primarily designed to assess parasite growth inhibition rather than directly measuring parasite viability after drug exposure, which is the focus of the present study. Readout methods used to assess parasite viability in a PRR assay setup are so far based on HRP2-ELISA (de Carvalho et al.), MitoTracker and SYBR green staining (Maiga et al.) and [<sup>3</sup>H]-hypoxanthine incorporation (Sanz et al.; Walz et al.) as cited in the manuscript. Many other readout methods to assess parasite growth have other limitations as briefly discussed in Hellingman et al, 2024. A comprehensive comparison and review of all available readout methods would therefore be beyond the scope of this manuscript.

      It would be helpful for authors to provide some indication of the cost comparison between the PPR v2 and MULT-i<sup>2</sup>.

      We thank the reviewer for this valuable suggestion. We agree that a comparison of the costs associated with the PRR v2 and MULT-i<sup>2</sup> assays would be informative, but while the consumable costs provide one measure of assay expense, we consider the reduction in hands-on time and the simplified workflow to be the main contributors to the overall cost advantage of the MULT-i<sup>2</sup> assay. These reductions in labor requirements are subject to large regional differences and impossible for us to access. Nevertheless, together with the increased throughput and the reduced labor, make the MULT-i<sup>2</sup> assay more cost-effective for larger-scale applications compared with the PRR v2 assay.

      Also, the authors should indicate whether these reagents will be deposited in a repository such as BEI Resources. They should also indicate conditions for other groups to request these materials, such as whether an MTA is required.

      We thank the reviewer for this important suggestion. The engineered parasite line will be made available for non commercial use to other researchers upon request. An MTA will be required excluding commercial use of the provided strains. The detailed code used for data analysis is available upon request, and an example code file has already been included as a Supplementary File.

      The pharmacological models are interesting, but likely well out of the range of expertise of many labs. Has code been deposited into public repositories that make it possible for other labs to implement these analyses?

      We thank the reviewer for this valuable comment. We agree that implementation of pharmacological modeling approaches can represent a barrier for laboratories without prior experience in pharmacometric analysis, particularly due to the requirement for specialized software such as NONMEM. To facilitate implementation, an example code is provided in the Supplementary File. The final model was developed using a forward–backward selection approach for parameter estimation and model refinement as described in the Methods section. These additions should help other researchers adapt the approach to their own datasets.

      Reviewer 2 (Public review):

      Summary

      Antimalarial combination therapy is the standard of care for malaria, a disease that impacts hundreds of millions of people annually. Combination therapy is crucial for effectively treating the disease and delaying the emergence of drug resistance. Despite the importance of choosing appropriate partner antimalarials for combination therapy, drug interactions are typically evaluated late in the course of drug development. Standard in vitro assays that determine synergistic, antagonistic, or additive interactions between drug combinations rely on measuring inhibition of parasite proliferation, which is inadequate for translation to pharmacodynamic models for parasite clearance in the patient. Direct measurement of parasite viability under drug treatment has previously relied on methods that are labor and resource-intensive, limiting applications to single compounds and single concentrations. Here, Hellingman et al make use of an inducible chemiluminescence reporter to measure cell viability and apply this novel approach to quantify drug interactions. The methodology is a significant improvement upon prior methods, requiring significantly fewer resources, half the time, and substantially less handling than the standard PRR v2 assay, whilst maintaining high resolution and sensitivity.

      They assess the limit of detection for the improved method and cross-reference their results for single drugs at a single concentration with the currently standard PRRv2 assay. The authors next established analytical methods to characterize the impact of drug combinations on parasite viability using the GDPI pharmacodynamic model and compared their MULT-i<sup>2</sup> assay to the prior cPRR approach. Their refined workflow allowed them to comprehensively evaluate the known synergistic combination between atovaquone and proguanil with greater resolution than the comparable cPRR assay and identified additional interaction parameters between the fast-acting antimalarials piperaquine and pyrimethamine. Overall, the authors demonstrate that their inducible lacZ system provides significant advantages compared with prior approaches to determine parasite viability. They convincingly demonstrate the strengths of their approach by characterizing two antimalarial combinations at much greater resolution than previously possible with prior methods. The system and methods established here will be particularly useful for evaluating novel antimalarial combinations with chemical series in preclinical evaluation and to optimize future antimalarial therapies.

      Strengths:

      The streamlined approach relies on induction of the lacZ enzyme only after drug washout. As opposed to when stably expressed, this allows the authors to estimate parasite viability without undergoing serial dilutions to estimate viable parasite titers. This innovation vastly reduced resource and time intensity, enabling greater throughput for parasite viability estimation. The established methodology and analysis pipeline enabled the testing of 49 drug combinations for parasite viability in the MULT-i<sup>2</sup> assay compared to only 9 in the conventional cPRR assay. This provided improved resolution in the ability to estimate drug combination parameters in a pharmacodynamic model. The ability to comprehensively characterize combination pharmacodynamic properties in vitro will have important implications for downstream modelling of in vivo combinations, and for optimizing future antimalarial combination therapies.

      The authors made good use of modelling and AICc for parametric estimation and model evaluation to demonstrate the advantages of the richer dataset afforded by the MULT-i<sup>2</sup> assay.

      We thank reviewer 2 for her/his appreciation of our work.

      Weaknesses:

      The authors correctly identified a range of confounding effects that lead to artefacts in their assay results when compared to the cPRR assay. For instance, the authors observed reduced signal at high parasite density during recovery due to overgrowth and likely enzyme degradation, and suggested residual signal may remain from non-proliferating sexual stage parasites surviving drug treatment that would not be detected in the cPRR assay.

      Measurement of parasite viability in the MULT-i<sup>2</sup> assay was achieved by extrapolating the chemoluminescence signal to that of a serial dilution of parasites made at the initiation of drug treatment. How did the authors account for differing levels of enzyme expression at early (e.g., ring) vs late stage parasites (trophozoite or schizonts)? Were cultures synchronized prior to initiation of assays? Could differences in life-cycle progression following drug treatment be an additional confounding factor that may account for differences with the PRR v2 assay?

      We thank the reviewer for raising this important point. All, the MULT-i<sup>2</sup> and PRR v2 assay were performed using asynchronous parasite cultures. We will clarify this in the revised manuscript.

      We agree that parasite developmental stages may influence the MULT-i<sup>2</sup> readout, as LacZ expression levels differ between parasite stages, with differences observed between ring stages and more mature trophozoite/schizont stages as published by Hellingman et al., 2024. This represents a potential source of variability, as the MULT-i<sup>2</sup> assay quantifies the amount of expressed reporter enzyme rather than directly measuring parasite numbers at the time of readout. The use of asynchronous cultures minimizes the impact of stage-specific effects by providing a mixed parasite population representative of the natural distribution of developmental stages. Nevertheless, we acknowledge that differences in parasite stage progression following drug exposure may contribute to variation in the extrapolated parasite numbers and may partially explain differences observed between the MULT-i<sup>2</sup> and PRR v2 assay measurements. We will add this consideration to the Discussion.

      The addition of an inducible element is an improvement of their earlier lacZ/β-gal<sup>SENSOR</sup> (PMID: 41575867); however, the authors fail to explain why this is an improvement and how this adds additional merit over the initial system. While the authors compare their new assay to the PRR v2, they fail to compare it to their own non-inducible lacZ/β-gal<sup>SENSOR</sup> system. Their non-inducible system already showed superiority to the cPRR assays, and it would be good to show how they compare and what the advantages of the new system are over the old. e.g., how is the signal-to-noise improved?

      We thank the reviewer for this important comment. The main improvement provided by the inducible system is the temporal separation of parasite growth/drug exposure from reporter expression. In the original non-inducible lacZ/β-gal<sup>SENSOR</sup> system, reporter expression occurs continuously throughout the assay, resulting in accumulation of β-galactosidase during parasite growth/drug exposure and therefore an increasing background signal. Consequently, quantification relies on endpoint reporter levels and does not allow the reporter expression window to be standardized independently of parasite exposure history.

      In contrast, in the MULT-i<sup>2</sup> system, reporter expression is initiated only after addition of rapamycin post antimalarial drug washout. This prevents reporter accumulation during the drug exposure window and ensures a defined reporter enzyme accumulation window after drug exposure. Importantly, this allows parasite numbers to be extrapolated from a calibration curve generated at the time of induction, which would not be possible with the non-inducible system because reporter expression would continue after drug removal and would depend on the previous culture history.

      We will revise the manuscript to more clearly describe these advantages and to emphasize that the key benefit of the inducible system is not simply an increase in signal intensity, but improved control of reporter expression, reduced background accumulation, and the ability to perform quantitative parasite reduction rate measurements.

      How does the sensitivity compare? How quickly does the can the signal be detected after induction? They show signal after 48h, but it would be very useful to the community to look at earlier timepoints as well and compare them to the uninduced line and a line that has been induced 48h earlier to match the expression patterns throughout the lifecycle (something like 2h,4h,6h, 12h, and 24h).

      We thank the reviewer for this important suggestion. We acknowledge that the sensitivity of the MULT-i<sup>2</sup> readout depends on both the initial parasite density and the duration of the induction period and that a detailed characterization of the induction kinetics, including earlier time points after rapamycin addition, would provide additional information on the sensitivity and temporal resolution of the MULT-i<sup>2</sup> system.

      In the present study, we focused on the time window relevant for application of the assay in a PRR assay workflow and routine drug screening setting. Earlier time points (<24 h after induction) were therefore not systematically evaluated. The selected time points were chosen based on the expected kinetics of the loxP-DiCre recombination system, which has previously been reported to achieve high recombination efficiency within one asexual parasite cycle, (Collins et al., 2013) and shown with own data in this study, as well as on practical considerations for implementation in routine workflows.

      Is the chemiluminescence signal for the i-lacZ induced parasites comparable to the stably expressed lacZ parasites previously characterized by the group? If so, do the authors consider this inducible iteration a complete replacement for PRR assays?

      We thank the reviewer for this question. The chemiluminescence signal obtained with the inducible lacZ (i-lacZ) parasites is comparable to that observed with the previously characterized constitutively expressing lacZ parasites. However, the inducible system provides an important additional advantage by avoiding continuous β-galactosidase production and accumulation during parasite growth, thereby reducing background signal and enabling a controlled reporter expression window.

      We do not consider the MULT-i<sup>2</sup> assay to be a replacement for classical PRR assays. Rather, we consider it a complementary approach that enables more efficient screening and characterization of drug combinations, particularly by providing information on the time-dependent onset of parasiticidal activity in a higher-throughput format. Promising combinations identified using MULT-i<sup>2</sup> assay can subsequently be investigated in more extensive PRR assays.

      Reviewer 3 (Public review):

      In this manuscript, the authors strived to develop a highly efficient drug survival assay for in vitro cultured human malaria parasites P. falciparum. This was done by generating a transgenic P. falciparum line using a creLox strategy that allows detection of (presumably) viable parasites by a β-lactamase assay. To estimate the Limit of quantification of the recombined P. falciparum NF54i-lacZ, the authors ultimately designed a protocol in which viable parasites are detected by the luminescence of β-D-galactoside generated by β-lactamase within the transgenic parasites. For this, the parasite must be incubated with rapamycin for 120 hours to induce CreLox recombinase, which places β-lactamase under an active promoter. Using this assay, termed MULT-i<sup>2</sup>, the author shows interactions between two antimalarial drug pairs that were previously demonstrated by another assay. In the case of pyronaridine and piperaquine pair, the NULT-i<sup>2</sup> assay generated some additional insights compared to the previous assay, presumably by virtue of including more concentration datapoints. In conclusion, the authors argue that the MULT-i<sup>2</sup> assay is much less resource-intensive and time-consuming and can be applied on a large scale at a much lower cost and with the highest efficiency.

      Overall, the data generated in this manuscript are clear and well represented, and I am convinced that MULT-i<sup>2</sup> provides yet another of many drug assays for malaria parasites and could be put to good use. However, I struggle to fully appreciate the merit of his study, as the manuscript reads more like a technical document than a scientific study.

      We thank reviewer 3 for her/his appreciation of our work.

      I particularly lack an understanding of the strengths and weaknesses/limitations of the MULT-i<sup>2</sup> methodology and, thus, its applicability. I also do not fully appreciate the need for such an elaborate luminescence-based experimental setup. It would be good if some of these issues were addressed.

      Specifically:

      (1): The whole assay is based on detecting parasites by luminescence after 120 hr (5 days) after drug exposure. During that time, presumably the parasites that survived the drug pressure regrow to a detectable level and, at the same time, perform efficacious CreLox-based recombination to produce β-D-galactoside for detection. Is this necessary? How superior is this detection method to other methods, such as Fluorescence-assisted Cell Sorting (FACS), etc? Moreover, the 5-day growth-CreLox-β-D-galactoside production could introduce a series of confounding effects. In my view, more studies (beyond comparisons with a single existing method) would be useful for understanding this entire process.

      We thank the reviewer for raising this important point regarding the rationale, applicability, and limitations of the MULT-i<sup>2</sup> methodology.

      Quantification of viable parasites after drug exposure remains challenging, particularly when surviving parasites are present at low frequencies or require extended recovery periods. Current approaches, such as the parasite reduction ratio (PRR) assay based on [<sup>3</sup>H]-hypoxanthine incorporation, provide sensitive measurements of replicating parasites but are labor-intensive, require specialized infrastructure, and are not easily scalable for large numbers of drug combinations. Alternative approaches based on HRP2 detection no longer rely on radioactive readouts but generally provide lower sensitivity, particularly when quantifying low levels of surviving parasites within a shorter time frame.

      The MULT-i<sup>2</sup> assay was developed to address these limitations by combining a highly sensitive chemiluminescent β-galactosidase readout with an inducible reporter system. The 5-day induction period after drug exposure serves as a controlled gene expression step, allowing surviving parasites to recover and produce sufficient reporter signal for sensitive quantification using a standard plate reader. This approach enables higher-throughput assessment of parasiticidal activity while avoiding radioactive readouts and reducing the need for labor-intensive dilution-based approaches.

      We acknowledge that the recovery and reporter expression period introduces additional biological steps compared with direct parasite detection methods and may therefore represent a potential source of variability. The MULT-i<sup>2</sup> assay is not intended to replace all existing viability measurements but rather to provide a complementary screening tool for investigating larger numbers of drug combinations. More detailed comparisons with additional detection platforms, including fluorescence-based approaches such as flow cytometry, would be valuable; however, a comprehensive comparison of all available parasite viability readouts was beyond the scope of this study. We will add more explanations to the Discussion including the strengths and limitations.

      (2) Related to that above, how would MULT-i<sup>2</sup> perform in case of drugs that do not necessarily kill all parasites, such as artemisinin? In the case of artemisinin, it is becoming evident that at least a small fraction of the parasite revives after treatment via a temporary dormancy state. This has, in fact, also been shown for other drugs such as mefloquine, pyrimethamine, etc. Would such a situation produce a range of false readings? In general, in its current state, it is hard to see what the limitations of this method are, which makes it hard to decide whether to use it for a particular application.

      We thank the reviewer for raising this important point regarding the interpretation and applicability of the MULT-i<sup>2</sup> assay. We agree that distinguishing between growth inhibition assays and viability-based assays is essential when interpreting the response to drugs that induce temporary parasite dormancy or delayed recovery.

      The MULT-i<sup>2</sup> assay was specifically developed as a viability-based approach and therefore differs fundamentally from conventional IC50 assays, which primarily measure inhibition of parasite growth during drug exposure and may not capture parasites that survive treatment through temporary growth arrest or dormancy. Similar to the PRR assay, the MULT-i<sup>2</sup> assay measures the ability of surviving parasites to recover and proliferate after drug exposure. Therefore, parasites that temporarily enter a dormant state but subsequently resume replication are expected to contribute to the measured signal rather than representing false-positive or false-negative results.

      This is illustrated by the artemisinin experiments presented in this study, where the MULT-i<sup>2</sup> assay captures the recovery of surviving parasites following treatment as it does the PRR v2 assay.

      (3) Given the stated cost and labor efficiency of MULT-i<sup>2</sup>, it is disappointing to see only two applications for two drug pairs: atovaquone/proguanil and piperquine/pyronaridine, for both of which their interactions were already known. The manuscript would benefit greatly if the authors demonstrated more drug interactions and identified (and ultimately validated) new ones. This would certainly make MULT-i<sup>2</sup> method more attractive. In particular, it would be nice to see if one could use MULT-i<sup>2</sup> for studies of triple combinations as enthusiastically suggested.

      We thank the reviewer for this valuable suggestion. We agree that demonstrating additional applications, including triple-drug combinations, would further highlight the potential of the MULT-i<sup>2</sup> assay.

      The primary aim of this study was to validate the MULT-i<sup>2</sup> methodology against the established PRR v2 assay and to demonstrate that the new platform can reproduce known parasiticidal interaction profiles while providing a more scalable workflow. For this reason, we selected well-characterized drug combinations, including atovaquone/proguanil and piperaquine/pyronaridine, which provide suitable benchmark systems for comparison with previous PRR data.

      Although evaluation of a larger number of novel combinations and triple-drug regimens would be highly valuable, generating corresponding PRR datasets for direct comparison was beyond the scope of the current study.

      (4) Throughout the manuscript, the authors claim that MULT-i<sup>2</sup> is considerably less expensive and can be done much faster than previous methods. In my view, this is not exactly a scientific argument. The cost of an assay depends heavily on the cost of reagents and labor, which are subject to market price fluctuations. The efficiency and time consumption can very much depend on laboratory organization, etc. Unless the author could specifically demonstrate where and how these assays are cheaper and faster, I suggest not discussing this.

      We thank the reviewer for this important comment. We agree that absolute assay costs can vary depending on local reagent prices, labor costs and laboratory infrastructure.

      When comparing both methods under the same laboratory conditions, the total assay duration of the MULT-i<sup>2</sup> assay is shorter than that of the PRR assay (11 days (MULT-i<sup>2</sup>) compared with approximately 21–28 days (PRR) according to published protocols). In addition, the MULT-i<sup>2</sup> assay reduces labor-intensive processing steps and enables higher-throughput measurements using a plate reader for readout. These factors contribute to reduced workload and improved scalability, independent of fluctuations in individual reagent or personnel costs.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Wang et al. reports the potential involvement of an asymmetric neurocircuit in the sympathetic control of liver glucose metabolism.

      Strengths:

      The concept that the contralateral brain-liver neurocircuit preferentially regulates each liver lobe may be interesting.

      Weaknesses:

      However, the experimental evidence presented did not support the study's central conclusion.

      We thank the reviewer for recognizing the conceptual novelty of our work and for constructive comments aimed at enhancing its rigor and clarity. In response, we carried out targeted experiments to address the points raised, including: (i) further characterization of LPGi projections to vagal and sympathetic circuits; (ii) evaluation of potential pancreatic involvement; and (iii) validation of the specificity of chemogenetic activation within the proposed circuit. All new experiments, figures, and text have been incorporated, and corresponding revisions are highlighted for ease of review.

      (1) Pseudorabies virus (PRV) tracing experiment:

      The liver not only possesses sympathetic innervations but also vagal sensory innervations. The experimental setup failed to distinguish whether the PRV-labeling of LPGi (Lateral Paragigantocellular Nucleus) is derived from sympathetic or vagal sensory inputs to the liver.

      Thank you for raising this important point. We fully agree that the liver receives both sympathetic and vagal sensory innervation, and we acknowledge that PRV-based tracing alone does not definitively distinguish between these two pathways. This represented a limitation of the original experimental design.

      Based on established anatomical literature as well as our experimental observations, vagal sensory neuron cell bodies reside in the nodose ganglion (NG), and their central projections terminate predominantly in the nucleus of the solitary tract (NTS) (Nature. 2023;623(7986):387-396; Curr Biol. 2020;30(20):3986-3998.e5.), which is located in the dorsomedial medulla. In contrast, the LPGi, together with other sympathetic-related nuclei, is predominantly distributed in the ventral medulla (Cell Metab. 2025;37(11):2264-2279.e10; Nat Commun. 2022;13(1):5079).

      To determine whether the LPGi contains neurons that modulate the liver via vagal sensory pathways, we performed two complementary experiments.

      First, we conducted CGRP immunohistochemistry on brainstem sections, using the NTS, a well-established visceral sensory centre, as a positive control. While abundant CGRP-positive cell bodies were detected in the NTS as expected, few to no CGRP-positive cell bodies were observed in the LPGi (Figure S1G). These results strongly support that the LPGi neurons labeled in our PRV tracing predominantly belong to sympathetic efferent circuits rather than vagal sensory pathways.

      Second, to examine whether LPGi neurons send axonal projections to sensory ganglia, we injected hSyn-Cre combined with DIO-Axon-EGFP into the LPGi and examined both the dorsal root ganglia (DRG) and nodose ganglia (NG). No Axon-EGFP-positive signals were detected in either ganglion (Figures S1H-S1J), indicating that LPGi neurons do not directly innervate sensory ganglia. In other words, PRV cannot retrogradely trace to the LPGi via the NG or DRG.

      These additions have been incorporated into the revised manuscript, with the Result 1 clearly documenting that these findings confirm that the LPGi specifically regulates sympathetic, rather than vagal sensory, inputs to the liver.

      (2) Impact on pancreas:

      The celiac ganglia not only provide sympathetic innervations to the liver but also to the pancreas, the central endocrine organ for glucose metabolism. The chemogenetic manipulation of LPGi failed to consider a direct impact on the secretion of insulin and glucagon from the pancreas.

      Thank you for this important comment. We agree that the celiac ganglia (CG) provide sympathetic innervation not only to the liver but also to the pancreas, which plays a central role in glucose homeostasis through the secretion of both insulin and glucagon. Therefore, the potential pancreatic implications associated with LPGi chemogenetic manipulation are worth careful consideration.

      To address this concern, we measured circulating glucagon and insulin levels following chemogenetic manipulation of the LPGi<sup>GAD1</sup> neurons. We found that neither glucagon nor insulin levels changed significantly under our experimental conditions, which indicated that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated by changes in pancreatic hormone secretion (Figure S2G).

      These additions have been incorporated into the revised manuscript, with the Result 2 clearly documenting that these findings confirm that the hyperglycemic effect induced by LPGi activation is unlikely to be mediated indirectly via altered pancreatic endocrine output.

      (3) Neuroanatomy of the brain-liver neurocircuit:

      The current study and its conclusion are based on a speculative brain-liver sympathetic circuit without the necessary anatomical information downstream of LPGi.

      Thank you for raising this important point. A clear anatomical definition of the downstream pathways linking the brain to the liver was essential for interpreting the proposed brain-liver sympathetic circuit.

      The present study (Figure 4A) provides direct anatomical evidence supporting the organization of the brain–liver sympathetic neurocircuit. These observations are consistent with our recent detailed characterization of the brain-liver sympathetic circuit published in Cell Metabolism (Cell Metab. 2025;37(11):2264–2279). In that study, we showed that LPGi GABAergic neurons inhibit GABAergic neurons in the caudal ventrolateral medulla (CVLM). Disinhibition of CVLM reduced GABAergic suppression of rostral ventrolateral medulla (RVLM) neurons, which are key excitatory drivers of sympathetic tone. RVLM neurons project to sympathetic preganglionic neurons in the sympathetic chain (Syc). These neurons synapse with postganglionic sympathetic neurons in ganglia such as the celiac-superior mesenteric ganglion (CG-SMG). Postganglionic sympathetic fibers then innervate the liver, releasing norepinephrine (NE) to activate hepatic β<sub>2</sub>-adrenergic receptors and stimulate hepatic glucose production (HGP).

      Together, these data establish a coherent anatomical basis for the proposed brain-liver sympathetic pathway and clarify the downstream organization relevant to the functional experiments presented in figure 4A and Author response image 1..

      Author response image 1.

      Tracing scheme (Left) and whole-mount imaging (Right) of PRV-labeled brain-liver neurocircuit. Scale bars, 3,000 (whole mount) or 1,000 (optical sections) μm.

      (4) Local manipulation of the celiac ganglia:

      The left and right ganglia of mice are not separate from each other but rather anatomically connected. The claim that the local injection of AAV in the left or right ganglion without affecting the other side is against this basic anatomical feature.

      Thank you for raising this important anatomical point. We fully acknowledge that the left and right CG in mice are interconnected, and that unilateral viral injection could theoretically affect the contralateral side. The CG-SMG complex serves as a major sympathetic hub that regulates visceral organ functions. Recent transcriptomic, anatomical, and functional studies have revealed that the CG-SMG is not a homogeneous structure but is composed of molecularly and functionally distinct neuronal populations. These populations exhibit specialized projection patterns and regulate different aspects of gastrointestinal physiology, supporting a model of modular sympathetic control. (Nature. 2025 Jan;637(8047):895-902). Therefore, we were aware of this phenomenon during the initial stages of these experiments.

      To minimize unintended spread to the contralateral CG, we took two complementary approaches.

      First, we optimized the injection strategy by using an extremely small injection volume (100 nL per site), with a very slow infusion rate (50 nL/min), and fine glass micropipettes. With these refinements, contralateral viral spread was rarely observed.

      Second, and importantly, all animals included in the final analyses were subjected to post hoc anatomical verification. After completion of the experiments, CGs were collected, sectioned, and examined for viral expression. As shown in Supplementary Figure 5F, only mice in which viral expression was strictly confined to the targeted CG, with no detectable infection in the contralateral ganglion, were included in the presented data.

      Together, these measures ensure that our local manipulation of the intended CG produced the reported effects. We have revised the Methods section to more explicitly detail these technical precautions, and the legend for Figure S5F clearly states its role in validating injection specificity.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Wang and colleagues aims to determine whether the left and right LPGi differentially regulate hepatic glucose metabolism and to reveal decussation of hepatic sympathetic nerves.

      The authors used tissue clearing to identify sympathetic fibers in the liver lobes, then injected PRV into the hepatic lobes. Five days post-injection, PRV-labeled neurons in the LPGi were identified. The results indicated contralateral dominance of premotor neurons and partial innervation of more than one lobe. Then the authors activated each side of the LPGi, resulting in a greater increase in blood glucose levels after right-sided activation than after left-sided activation, as well as changes in protein expression in the liver lobes. These data suggested modulation of HGP (hepatic glucose production) in a lobe-specific manner. Chemical denervation of a particular lobe did not affect glucose levels due to compensation by the other lobes. In addition, nerve bundles decussate in the hepatic portal region.

      We thank the reviewer for the thorough and constructive evaluation of our manuscript. In direct response, we undertook comprehensive revisions to enhance the rigor and clarity of the study, including: (i) correcting ambiguous or misleading terminology about anatomical resolution and sympathetic circuit organization; (ii) expanding the Methods section with complete experimental details, improved image presentation, and explicit justification of our viral and genetic approaches; and (iii) strengthening data interpretation by addressing issues related to sparse PRV labeling, projection heterogeneity, and the functional implications of double-labeled neurons.

      Strengths:

      The manuscript is timely and relevant. It is important to understand the sympathetic regulation of the liver and the contribution of each lobe to hepatic glucose production. The authors use state-of-the-art methodology.

      Weaknesses:

      (1) The wording/terminology used in the manuscript is misleading, and it is not used in the proper context. For instance, the goal of the study is "to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism..." (see abstract); however, the authors focus on the brainstem (a single structure without hemispheres). Similarly, symmetric is not the best word for the projections.

      We thank the reviewer for raising these critical points regarding terminology and conceptual framing. We acknowledge that certain phrases in our original manuscript may have been overly broad or ambiguous, particularly in describing the scope of sympathetic heterogeneity and the specificity of neural projections. Due to practical constraints and the scope of our study, our investigation focused on the brainstem, which represents the final common pathway for these lateralized commands. We acknowledge that terms referring to the cerebral hemispheres do not accurately describe our study. We have revised the manuscript to ensure accurate and consistent terminology.

      Below are specific examples:

      Original 1: This study aims to investigate whether cerebral hemispheres differentially regulate hepatic glucose metabolism and localize the site of sympathetic crossover to the liver.

      Revised 1: “This study aimed to determine whether the central nervous system exerts lateralized control over hepatic glucose metabolism and to localize the site of peripheral sympathetic crossover to the liver.”

      Original 2: These findings demonstrate that the brain exerts lobe-specific, lateralized control of hepatic glucose metabolism via symmetric brain-liver sympathetic pathways.

      Revised 2: “These findings demonstrate that the brainstem can exert lobe-specific, lateralized control of hepatic glucose metabolism via bilaterally projecting brain–liver sympathetic pathways.”

      Original 3: The cerebral hemispheres exhibit pronounced functional asymmetry, [1,2] a phenomenon traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation.[3,4]

      Revised 3: “Pronounced functional lateralization within the central nervous system (CNS) is a well-documented phenomenon, [1,2] traditionally associated with cognitive and motor processes such as language, voluntary movement, and spatial navigation [3,4].

      (2) Sparse labeling of liver-related neurons was shown in the LPGi (Figure 1). It would be ideal to have lower magnification images to show the area. Higher quality images would be necessary, as it is difficult to identify brainstem areas. The low number of labeled neurons in the LPGi after five days of inoculation is surprising. Previous findings showed extensive labeling in the ventral brainstem at four days post-inoculation (Desmoulins et al., 2025). Unfortunately, it is not possible to compare the injection paradigm/methods because the PRV inoculation is missing from the methods section. If the PRV is different from the previously published viral tracers, time-dependent studies to determine the order of neurons and the time course of infection would be necessary.

      We sincerely thank the reviewer for these detailed and constructive comments regarding the PRV tracing experiments. We fully agree that careful presentation and interpretation of the anatomical data are essential for ensuring rigor and transparency. We address each point in detail below.

      (1) Image magnification and anatomical context of LPGi labeling

      We agree that the original images did not sufficiently convey the broader anatomical context of the LPGi. Due to fluorescence quenching in previous sections, we repeated the PRV retrograde tracing experiment and performed statistical analysis. In the revised manuscript, we replaced the original panels in Figure 1 and Figure S1 with new images that include lower-magnification overviews of the brainstem, alongside higher-magnification views of the LPGi (Figure 1). These images clearly delineate the LPGi with respect to established anatomical landmarks and atlas boundaries. Image contrast and resolution were optimized to allow unambiguous identification of PRV-labeled neurons and surrounding structures.

      (2) Sparse LPGi labeling at 5 days post-injection and methodological details

      We apologize for the omission of the detailed PRV injection protocol in the original Methods section. We deliberately used small-volume, local injections (1 µL per liver lobe) to minimize viral spread and to restrict labeling to circuits specifically connected to the targeted hepatic region. This sparse labeling was consistent with the use of small, spatially restricted injections designed to minimize off-target spread and preferentially label higher-order upstream neurons. This information has now been added, including the PRV strain, viral titer, injection volume, precise injection coordinates, and surgical procedures. All new figures, legends, and Method details have been incorporated, with changes clearly highlighted for ease of review.

      These additions have been incorporated into the revised manuscript, in Figure 1B-D, Figure S1C-F. Furthermore, we also added details of the Methods.

      (3) Not all LPGi cells are liver-related. Was the entire LPGi population stimulated, or was it done in a cell-type-specific manner? What was the strain, sex, and age of the mice? What was the rationale for using the particular viral constructs?

      We thank the reviewer for this insightful and important question. We agree that not all neurons within the LPGi are liver-related, and we apologize that our rationale was not clearly articulated in the original manuscript.

      (1) Our decision to target GABAergic neurons in the LPGi using GAD1-Cre mice was based on prior experimental evidence rather than an assumption about the entire LPGi population. In our previous study (Cell Metab. 2025;37(11):2264-2279.e10), we performed single-cell RNA sequencing on retrogradely labeled LPGi neurons following liver tracing. These analyses revealed that the majority of liver-projecting LPGi neurons are GABAergic in nature. Based on these findings, we chose to selectively manipulate GABAergic neurons in the LPGi rather than the entire LPGi neuronal population, to achieve greater cellular specificity and to minimize potential confounding effects arising from heterogeneous neuron types within this region. We regret that this rationale was not clearly described in the original submission and have now revised the manuscript to explicitly state this reasoning (Results section 2, paragraph 2: “Prior single-nucleus RNA sequencing and immunofluorescence analyses demonstrated that liver-projecting LPGi neurons are predominantly GABAergic.”).

      (2) In addition, we apologize for the omission of mouse strain, sex, and age information in the Methods section. These details have been fully added.

      (3) We selected AAV-based viral vectors, specifically the AAV9 serotype, due to their well-established efficiency in transducing neurons in the brainstem, relatively low toxicity, and widespread use in circuit-level chemogenetic and optogenetic studies. When combined with Cre-dependent viral constructs in GAD1-Cre mice, this approach enabled selective and reliable manipulation of LPGi GABAergic neurons.

      (4) The authors should consider the effect of stimulation of double-labeled neurons (innervating more than one lobe) and potential confounding effects regarding other physiological functions.

      We thank the reviewer for raising this important point. We agree that neurons innervating more than one liver lobe could, in principle, introduce potential confounding effects and may reflect higher-order integrative autonomic neurons.

      This consideration is consistent with a key finding of the cited study: the CG-SMG contains molecularly distinct sympathetic neuron populations (e.g., RXFP1<sup>+</sup> vs. SHOX2<sup>+</sup>) that exhibit complementary organ projections and separate, non‑overlapping functions. Specifically, RXFP1<sup>+</sup> neurons innervate secretory organs (pancreas, bile duct) to regulate secretion, while SHOX2<sup>+</sup> neurons innervate the gastrointestinal tract to control motility. This functional segregation supports the concept of specialized autonomic modules rather than a uniform, “fight-or-flight” response, reinforcing the need for careful interpretation of circuit-specific manipulations. (Nature. 2025;637(8047):895-902; Neuron. 2026;114(3):463-478.e7).

      In our PRV tracing experiments, the proportion of double-labeled neurons was relatively small, suggesting that the majority of labeled LPGi neurons preferentially associate with individual hepatic lobes. Nevertheless, we recognize that activation of this minority population could contribute to broader physiological effects beyond strictly lobe-specific regulation. We have therefore added a paragraph in the second paragraph of the Discussion (Paragraph 2: “A small subset of LPGi neurons was double-labeled after bilateral PRV injections, suggesting a fraction of these neurons projects bilaterally to both sides of the liver. Such neurons may support interlobar coordination.”).

      (5) The authors state that "central projections directly descend along the sympathetic chain to the celiac-superior mesenteric ganglia". What they mean is unclear. Do the authors refer to pre-ganglionic neurons or premotor neurons? How does it fit with the previous literature?

      We thank the reviewer for pointing out this imprecise wording. We agree that the original phrasing was anatomically inaccurate and potentially confusing. The pathways we intended to describe involve brainstem premotor neurons that project to sympathetic preganglionic neurons in the spinal cord. These preganglionic neurons then innervate neurons in the CG-SMG, which in turn provide postganglionic input to the liver.

      We have revised the manuscript to clearly distinguish premotor from preganglionic neurons (Results section 4, paragraph 1: “Using whole-mount clearing, we visualized the brain–liver sympathetic circuit and found that preganglionic neurons in the spinal cord send descending fibers through the sympathetic chain (SyC) to innervate postganglionic neurons in the CG-SMG (Figure 4A). Further whole-mount TH immunostaining showed that TH-positive sympathetic cell bodies within the CG-SMG project to the liver along the hepatic vasculature (Figure 4B and Figure S5F). These observations suggest that the nerve bundles likely decussate at the porta hepatis before entering the individual hepatic lobes.”).

      (6) How was the chemical denervation completed for the individual lobes?

      We thank the reviewer for raising this important methodological concern. We agree that potential diffusion of 6-OHDA is a critical issue when performing lobe-specific chemical denervation, and we apologize that our original description did not sufficiently clarify how this was controlled.

      In the revised Methods section, we provided a detailed description of the denervation procedure, including the injection volume and concentration of 6-OHDA, as well as the physical separation and isolation of individual hepatic lobes during application to minimize diffusion to adjacent tissue.

      To directly assess the specificity of the chemical denervation, we included immunofluorescence and Western blot analyses demonstrating a selective reduction of sympathetic markers in the targeted lobe (Figure 3C), with minimal effects on non-targeted lobes. These results support the effectiveness and relative spatial confinement of the 6-OHDA treatment under our experimental conditions.

      We thank the reviewer for highlighting this point, which has helped us improve both the clarity and rigor of the manuscript.

      (7) The Western Blot images look like they are from different blots, but there are no details provided regarding protein amount (loading) or housekeeping. What was the reason to switch beta-actin and alpha-tubulin? In Figures 3F -G, the GS expression is not a good representative image. Were chemiluminescence or fluorescence antibodies used? Were the membranes reused?

      We thank the reviewer for this careful and detailed evaluation of the Western blot data. We apologize that insufficient methodological detail was provided in the original submission.

      (1) We would like to clarify that the protein bands shown within each panel were derived from the same membrane. To improve transparency, we provided full, uncropped images of the corresponding membranes in the supplementary materials. In addition, detailed information regarding protein loading amounts, gel conditions, and housekeeping controls has also been added to the Methods section.

      (2) The use of different loading controls (β-actin or α-tubulin) reflects a technical consideration rather than an experimental inconsistency. In our experiments, the molecular weight of the TH (62kDa) was too close to that of α-tubulin (55kDa), and β-actin (42kDa) was therefore used to avoid band overlap and to ensure accurate quantification.

      (3) Regarding the GS signal shown in Figures 3F–G, we agree that the original representative image was suboptimal. This appears to be related to antibody performance rather than sample quality. To address this, we repeated the Western blot from Figures 3F–G using a newly validated antibody. The original tissue samples had been aliquoted and stored at −80 °C, allowing reliable re-analysis.

      (4) All Western blot experiments were detected using chemiluminescence, and membrane stripping and reprobing procedures are now explicitly described in the Methods section.

      We thank the reviewer for highlighting these issues, which significantly improve the rigor and clarity of our data presentation. All new figures and legends have been incorporated, with changes clearly highlighted for ease of review.

      (8) Key references using PRV for liver innervation studies are missing (Stanley et al, 2010 [PMID: 20351287]; Torres et al., 2021 [PMID: 34231420]; Desmoulins et al., 2025 [PMID: 39647176]).

      We thank the reviewer for pointing out these important and highly relevant references that were inadvertently omitted in our initial submission. The studies by Stanley et al. (Proc Natl Acad Sci U S A, 2010), Torres et al. (Am J Physiol Regul Integr Comp Physiol, 2021), and Desmoulins et al. (Auton Neurosci, 2025) represent key PRV-based retrograde tracing work that has mapped central neural circuits innervating the liver and thus provide essential context for our anatomical analyses.

      We agree that the inclusion of these studies is necessary to properly situate our findings within the existing literature. Accordingly, we incorporated citations to these references in the revised manuscript and discussed their relationship to our results.

      Reviewer #3 (Public review):

      Summary:

      This study found a lobe-specific, lateralized control of hepatic glucose metabolism by the brain and provides anatomical evidence for sympathetic crossover at the porta hepatis. The findings are particularly insightful to the researchers in the field of liver metabolism, regeneration, and tumors.

      Strengths:

      Increasing evidence suggests spatial heterogeneity of the liver across many aspects of metabolism and regenerative capacity. The current study has provided interesting findings: neuronal innervation of the liver also shows anatomical differences across lobes. The findings could be particularly useful for understanding liver pathophysiology and treatment, such as metabolic interventions or transplantation.

      Weaknesses:

      Inclusion of detailed method and Discussion:

      We sincerely thank the reviewer for the positive and constructive feedback, which significantly enhances both the methodological rigor and the broader biological interpretation of our study. In direct response, we revised the Discussion to elaborate on the potential physiological advantages of a lateralized and lobe-specific pattern of liver innervation. Furthermore, we expanded the Methods section to include a comprehensive description of the quantitative analysis applied to PRV-labeled neurons. Together, these revisions strengthened the manuscript’s clarity, depth, and relevance to researchers in hepatic metabolism, regeneration, and disease.

      (1) The quantitative results of PRV-labeled neurons are presented, and please include the specific quantitative methods.

      We thank the reviewer for this helpful suggestion. We have added a detailed description of the quantitative methods used to analyze PRV-labeled neurons in the revised Methods section. We have now provided detailed information in the Methods section, including the criteria used for cell counting, the anatomical boundaries of the brain regions analyzed, the delineation of regions of interest, and the normalization procedures applied to derive the reported neuron counts. These additions have been incorporated into the revised Methods, with all changes clearly indicated for ease of review.

      (2) The Discussion can be expanded to include potential biological advantages of this complex lateralized innervation pattern.

      We appreciated the reviewer’s suggestion. We have expanded the Discussion to include a paragraph addressing the potential biological significance of lateralized liver innervation. We highlight that this asymmetric organization could allow for more precise, lobe-specific regulation of hepatic metabolism, enable integration of distinct physiological signals, and potentially provide robustness against perturbations. The additional discussion content has been highlighted in the revised version as indicated (Discussion section, paragraph 3: “Bilateral LPGi activation produced additive effects, indicating that both sides of the brainstem can cooperatively regulate hepatic metabolism in a spatially segregated manner. This pattern suggests that hepatic glucose output can be modulated in a lobe-specific, rather than uniform whole-organ, manner.”).

      Reviewer #4 (Public review):

      Summary:

      The studies here are highly informative in terms of anatomical tracing and sympathetic nerve function in the liver related to glucose levels, but given that they are performed in a single species, it is challenging to translated them to humans, or to determine whether these neural circuits are evolutionarily conserved. Dual-labeling anatomical studies are elegant, and the addition of chemogenetic and optogenetic studies is mechanistically informative. Denervation studies lack appropriate controls, and the role of sensory innervation in the liver is overlooked.

      We sincerely appreciate the reviewer's thoughtful evaluation and fully agree that findings derived from a single-species model must be interpreted with caution in relation to human physiology. In direct response, we revised the manuscript to explicitly clarify that all experimental data were obtained in mice and to provide a discussion of the limitations regarding direct extrapolation to humans. Concurrently, we expanded the Discussion section by integrating our findings with recent human and translational studies, including a multicenter clinical trial demonstrating that catheter-based endovascular denervation of the celiac and hepatic arteries significantly improved glycemic control in patients with poorly controlled type 2 diabetes, without major adverse events (Signal Transduct Target Ther. 2025;10(1):371). While our current work focuses on defining the anatomical organization and functional asymmetry of this circuit in mice, the clinical findings suggest that the core principles, sympathetic control of hepatic glucose metabolism via CG-liver pathways, may be conserved and of translational relevance. Additionally, we clarified the interpretation of TH labeling and expanded the discussion of hepatic sensory and parasympathetic innervation, acknowledging their important roles in liver-brain communication and identifying them as key directions for future research. Collectively, these revisions provide a more balanced, clinically informed, and rigorous framework for interpreting our findings.

      Specific Weaknesses - Major:

      (1) The species name should be included in the title.

      We thank the reviewer for this suggestion. We agree that the species should be clearly indicated. The findings presented in this study were obtained in mice using tissue clearing and whole-organ imaging approaches. Due to technical limitations, these observations are currently restricted to the mouse strain. We have updated the title (Symmetric brain-liver circuits mediate lateralized regulation of hepatic glucose output in mice) and clarified the species used throughout the manuscript.

      (2) Tyrosine hydroxylase was used to mark sympathetic fibers in the liver, but this marker also hits a portion of sensory fibers that need to be ruled out in whole-mount imaging data

      We thank the reviewer for pointing this out. We acknowledge that TH labels not only sympathetic fibers but also a subset of sensory fibers. We have added a limitation of this point in the revised manuscript. In addition, using SyGlass (2.4.0) three-dimensional reconstruction, we observed TH-positive nerve fibers originating from the CG-SMG extending along the porta hepatis and penetrating into the liver parenchyma. Given that the CG-SMG is a well-established sympathetic ganglion innervating visceral organs (Nature. 2025 Jan;637(8047):895-902.), these nerve fibers can be definitively identified as sympathetic. In parallel, we collected DRG from spinal segments T1-6 and T7-12 five days after intrahepatic PRV injection. While T7-12 DRG are known to contain sensory neurons innervating the liver, only a sparse number of PRV-positive neurons were detected in these segments (Anat Rec A Discov Mol Cell Evol Biol. 2004 Sep;280(1):827-35. Auton Neurosci. 2024 Jun;253:103174). The additional figure and discussion content have been highlighted in the revised version as indicated (Discussion section, paragraph 6: “Third, although whole-mount TH immunostaining with three-dimensional reconstruction revealed sympathetic nerve bundles projecting from the CG to the liver, TH is not entirely specific and can also label a subset of sensory neurons. More selective approaches, such as genetic targeting of sympathetic lineages, will be important for further validation.”).

      Author response image 2.

      Representative immunofluorescence images of PRV-labeled neurons (EGFP) in DRG from the spinal segments T1-6 (bottom) and T7-12 (top) following PRV injections into the liver lobes. Scale bars, 200μm

      (3) Chemogenetic and optogenetic data demonstrating hyperglycemia should be described in the context of prior work demonstrating liver nerve involvement in these processes. There is only a brief mention in the Discussion currently, but comparing methods and observations would be helpful.

      We thank the reviewer for this suggestion. Previous studies largely relied on electrical stimulation to modulate liver innervation, which provides relatively coarse control of neural activity (Eur J Biochem. 1992;207(2):399-411). By contrast, our use of chemogenetic and optogenetic approaches allows selective, cell-type-specific manipulation of LPGi neurons. We revised the Discussion to place our functional data in the context of prior work, highlighting how these more precise approaches improve understanding of the contribution of liver-innervating neurons to hyperglycemia. The newly added discussion has been clearly labeled in the response to facilitate your review (Discussion section, paragraph 3: “This spatial organization is likely obscured by conventional electrical stimulation, which indiscriminately activates heterogeneous sympathetic fibers. By contrast, chemogenetic and optogenetic approaches permit selective, cell type-specific manipulation of LPGi neurons, thereby revealing the contralateral and lobe-specific architecture of brain-liver sympathetic control”).

      (4) Sympathetic denervation with 6-OHDA can drive compensatory increases to tissue sensory innervation, and this should be measured in the liver denervation studies to implicate potential crosstalk, especially given the increase in LPGi cFOS that may be due to afferent nerve activity. Compensatory sympathetic drive may not be the only culprit, though it is clearly assumed to be. The sensory or parasympathetic/vagal innervation of the liver is altogether ignored in this paper and could be better described in general.

      We thank the reviewer for this insightful and important comment, which highlights a potential alternative interpretation of our findings. We agree that chemical sympathetic denervation with 6-OHDA may induce compensatory changes in non-sympathetic inputs, including sensory and parasympathetic (vagal) innervation of the liver.

      Conceptually, we agree with the reviewer’s perspective that the central nervous system operates as a highly integrated homeostatic regulatory system, continuously receiving and integrating a broad range of afferent signals. These inputs include, as noted by the reviewer, hepatic sensory and vagal afferents (Science. 2024;386(6722):673-677), as well as centrally derived interoceptive signals such as brain glucose, temperature sensing, even the pulsation of cerebral vascular system (Cell Metab. 2025;37(11):2264-2279.e10.; Cell Metab. 2022;34(6):888-901.e5; Science. 2024;383(6682):eadk8511). The CNS integrates these diverse signals and generates coordinated efferent outputs to maintain systemic homeostasis.

      From this viewpoint, the changes in c-FOS activity that we observe in the LPGi likely represent only a limited snapshot of this broader integrative process, rather than evidence of a single dominant pathway. We acknowledge that compensatory sensory or parasympathetic mechanisms, in addition to altered sympathetic drive, contributed to the observed LPGi activation following hepatic sympathetic denervation.

      We further acknowledge that, due to limitations in scope and experimental focus, we did not directly assess sensory or parasympathetic innervation of the liver in the present study. As appropriately pointed out by the reviewer, a more comprehensive characterization of hepatic neural inputs would provide a more complete picture of the underlying neurocircuitry. To address this, we expanded the Discussion and explicitly noted this limitation, including a more balanced discussion of potential crosstalk among sympathetic, sensory, and parasympathetic pathways and how these may collectively influence LPGi activity. For your convenience, the newly added discussion text has been distinctly marked in the manuscript (Discussion section, paragraph 4: “Although enhanced sympathetic output appears to mediate much of this compensation, our findings suggest that the underlying regulation extends beyond a purely descending pathway. In particular, c-FOS activation in the contralateral LPGi after unilateral 6-OHDA-mediated denervation suggests that the loss of peripheral input may be sensed through an ascending neural pathway, centrally integrated, and translated into compensatory sympathetic output to the intact hepatic lobes. These results therefore support a model in which hepatic glucose production is regulated by an integrated afferent-central-efferent loop, with our current analyses primarily resolving its efferent component.”).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      Although the findings are interesting, this reviewer has major concerns about the experimental design, methodology, results, and interpretation of the data. Experimental details are lacking, including basic information (age, sex, strain of mice, procedures, magnification, etc.).

      We thank the reviewer for this important recommendation. We agree that comprehensive reporting of experimental details is essential for rigor and reproducibility.

      In the revised manuscript, we added complete information regarding mouse strain, sex, age, and sample size for each experiment. In addition, detailed descriptions of surgical procedures, viral constructs, injection parameters, imaging magnification, and analysis methods have been incorporated into the Methods section.

      These revisions ensured that all experiments are described with sufficient technical detail and clarity to allow accurate interpretation and replication of our findings. Experimental details have been incorporated, and corresponding revisions are highlighted for ease of review.

      Reviewer #3 (Recommendations for the authors):

      Addressing a few questions might help:

      (1) The study found that liver-associated LPGi neurons are predominantly GABAergic. It would be informative to molecularly characterize the PRV-traced, liver-projecting LPGi neurons to determine their neurochemical phenotypes.

      We thank the reviewer for this insightful suggestion. We agree that molecular characterization of liver-projecting LPGi neurons is important for understanding their functional identity.

      This issue has been addressed in detail in our recent study (Cell Metab. 2025;37(11):2264-2279.e10), in which we performed single-cell RNA sequencing on retrogradely traced LPGi neurons connected to the liver. These analyses demonstrated that the majority of liver-projecting LPGi neurons are GABAergic, with a defined transcriptional profile distinct from neighboring non–liver-related populations.

      Based on these findings, the current study selectively targeted GABAergic LPGi neurons using GAD1-Cre mice. We have explicitly cited these molecular results in the revised manuscript to clarify the neurochemical identity of the PRV-traced LPGi neurons. New text has been incorporated, and corresponding revisions are highlighted for ease of review.

      (2) Is it possible to do a local microinjection of a sodium channel blocker (e.g., lidocaine) or an adrenergic receptor antagonist into the porta hepatis? That would potentially provide additional evidence for the porta hepatis as the functional crossover point.

      We appreciated the reviewer’s thoughtful suggestion. Although pharmacological blockade at the porta hepatis can modulate local neural activity, this approach is inherently limited in its ability to distinguish between ipsilateral and contralateral inputs. Consequently, it may not provide definitive evidence for neural crossover at this specific site.

      In our view, the anatomical evidence provided by whole-mount tissue clearing, dual-labeled tracing, and direct visualization of decussating nerve bundles at the porta hepatis offers a more definitive demonstration of sympathetic crossover. Pharmacological blockade would affect both crossed and uncrossed fibers simultaneously and therefore would not specifically resolve the anatomical organization of this decussation.

      Nevertheless, we agree that functional interrogation of the porta hepatis represents an interesting direction for future work, and we acknowledge this possibility in the Discussion (Paragraph 6: “Fourth, although our data support a peripheral decussation at the porta hepatis, direct validation of this crossover site was not feasible with local pharmacological blockade, as currently available approaches lack sufficient spatial specificity and would likely perturb multiple neural components. Future studies employing more selective inhibitory strategies will be required to directly test this possibility.”).

      (3) It is possible to investigate the effects of unilateral LPGi manipulation or ablation of one side of CG/SMG on liver metabolism, such as hyperglycemia?

      We thank the reviewer for this important suggestion. Because unilateral LPGi manipulation was already examined in our study (Figure 2D), we focused here on unilateral ablation of the CG to further assess lateralized sympathetic control of hepatic metabolism. We successfully performed unilateral CG ablation without LPGi manipulation, but observed no significant change in blood glucose compared with the sham group (Author response image 3A and 3B). To determine whether glucose homeostasis was nonetheless affected, we further performed glucose tolerance tests (GTT) and insulin tolerance tests (ITT) (Author response image 3C and 3D). Neither test showed significant impairment after unilateral ablation, suggesting that compensatory neural mechanisms and/or hormonal homeostatic regulation may be recruited to preserve systemic glucose homeostasis.

      Author response image 3.

      (A) Blood glucose levels in mice subjected to left- or right-sided CG ablation via 6-OHDA treatment (n = 6). (B) Representative images of ablation of CG. Scale bars, 100 μm. (C and D) Blood glucose levels during GTT (C, n = 6) and ITT (D, n = 6) in mice with left- or right-sided CG ablation.

      Reviewer #4 (Recommendations for the authors):

      In the abstract and elsewhere, the use of the term 'sympathetic release' is unclear - do you mean release of nerve products, such as the neurotransmitter norepinephrine? This should be more clearly defined.

      We thank the reviewer for pointing out this ambiguity. We agree that the term “sympathetic release” was imprecise. In the revised manuscript, we explicitly referred to the release of sympathetic neurotransmitters, primarily norepinephrine, from postganglionic sympathetic fibers.

      We revised the wording throughout the manuscript to ensure accurate and consistent terminology and to avoid potential confusion regarding the underlying neurobiological mechanisms.

      Original: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased sympathetic release, glucose production, and glycogen depletion.”

      Revised: “Following unilateral hepatic denervation, contralateral LPGi activation induced metabolic compensation in the remaining innervated lobes, characterized by increased norepinephrine release, glucose production, and glycogen depletion.”

    1. Author response:

      The following is the authors’ response to the previous reviews

      We thank you for the time you took to review our work and for your feedback! The main changes to the manuscript are:

      We added a paragraph to the Discussion addressing differences in visuomotor mismatch responses recorded over frontal and occipital electrodes, and their possible interpretation.

      We added time-frequency power and phase-locking analysis as supplementary figures to the manuscript.

      We added a statement in the Discussion emphasizing the importance of performing these experiments with denser EEG channel coverage.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, Solyga, Zelechowski & Keller study human visuomotor mismatch responses as an alternative instantiation of prediction errors to classic oddball paradigms. Using VR, they created a condition in which participants were moving around thereby creating a visuomotor coupling between physical movement and visual flow. To attempt to isolate the contribution of specifically movement-related predictions in this condition, they contrasted it to a condition in which participants were seated and rewatching their movement trajectory during the 'active' condition. Visuomotor mismatches were created by temporarily decoupling movement and visual experience by halting the VR display as participants continued to move.

      The core finding of the paper is that participants exhibit a positively-valenced response to the visuomotor decoupling in the active but not in the passive condition. Since walking speed only insignificantly slows down following decoupling events in the active conditions, the authors argue that this difference can not be accounted for by "changes in participants' behavior or to simple visual offset responses" with the latter being equal across both conditions. The following reinstatement of the coupling in turn does not differ between the two conditions. The authors additionally show that this mismatch response differs from visual onset responses elicited by checkerboard inversions and that it's "qualitatively" stronger than more commonly studied auditory oddball mismatch responses.

      The design with its focus on ecological validity is impressive, well-rationalized and the results are well illustrated. I additionally appreciate the control analyses with regards to changes in walking speed and playback DOF and, now added, additional participants who experience the passive condition before the active. I have a couple of questions/comments.

      My main question in round 1 regarded the isolation of visuomotor mismatch. Although the comparison with a seated control seems like a very sensible way to control for simple visual responses, there seem to be more differences than just a break in visuomotor coupling between the conditions. I therefore wonder whether the reduced offset response in the seated condition may be, in part, explained differently. For example, given that participants always conduct the active condition before rewatching their movement in the seated condition, it seemed likely that there is a component of learning across the session that flow will sometimes be halted. This is confirmed with the analyses. The explanation that there is a visuomotor component here is given further weight by their conduction of an additional group of participants who perform the conditions in the reverse order, so this has strengthened the manuscript considerably. However, it does of course remain an imperfect control because the visual stimulus is now different between the conditions for these participants. It's the best that can be achieved with this type of paradigm though and of course it yields a great deal of ecological validity.

      The reviewer is correct. But one should keep in mind that our result here stands in the context of a considerable amount of work on mouse cortex investigating responses to very similar visuomotor mismatches. There we can we have much additional evidence to argue that the cortical response to a visuomotor mismatch is a prediction error. We would argue, it is the best one can do in human experiments.

      I was also wondering whether the authors may consider the findings in frontal electrodes more closely given that the title of the paper focuses on a specifically occipital effect. Their further analyses have confirmed that there are likely interesting frontal effects. From a theoretical point of view, the spatial dissociation in adaptation effects, which were stronger in frontal and weaker in occipital areas, seems interesting and perhaps worth discussing, especially given the interpretation that "mismatch processing may initially arise in sensory visual areas before engaging higher-order frontal regions." How come the frontal decrease in responses is not accompanied by an analogous decrease in its supposed occipital source? Could these two responses reflect different kinds of prediction error signals (i.e. objective vs subjective)?

      We have added a paragraph to the Discussion addressing the differences between signals recorded over frontal and occipital electrodes, as suggested.

      I remain concerned that the authors fight too defensively that they have absolutely isolated visuomotor prediction mechanisms with this paradigm. It's a nice, informative study, but it seems odd to argue there are no other possible explanations. One picks a design to optimize some features but they will always come at some cost to others. Prioritising ecological validity, which is a justifiable aim, necessarily usually weakens some control over confounds.

      We are not sure what the reviewer is referring to here. We certainly do not think (or are aware of having argued) that a visuomotor prediction error is the only possibly interpretation of the response. In the last paragraph of our response to the reviewers point 3 in the last revision, we explicitly discuss that the interpretation of the responses as a prediction error is only one possible interpretation. Our argument is that it is the most likely given the evidence.

      To outline my reasoning fully: My concerns wrt generic influences of action on perception are reflected in Fig 1. The P1 is smaller when walking than sitting. It seems likely that the mismatch response reflects something about extrapolation or prediction, because it is larger when walking. However, it's not necessarily sensorimotor prediction. Even if you remove action from the equation, the flow can be extrapolated or predicted most of the time in a way it cannot so well when the video is halted. Of course the sitting condition somewhat controls for it, but when it came second the visual flow disruptions were more predictable here. A reduction in effects over time is indeed confirmed with their analyses. They now have conducted a study with the conditions in the reverse order and they find the same thing. But of course this necessitates non-identical visual flow because the sitting condition is playing the previous participant's flow. So it is likely that across all of these comparisons, it is the visuomotor mismatch that is especially salient. It's just that each comparison is a bit messy/confounded. It would strengthen the manuscript if there were some consideration given to the other processes likely at play here.

      We would be happy to add additional considerations to other processes. If the reviewer has anything specific in mind, we can add that, but it would need to be somewhat concrete with some theoretical basis. We share the reviewer’s intuition, but unless this can be formalized to the point of being experimentally testable, we do not see any value in discussing it in the manuscript.

      Regarding the reason for a difference in visual responses in walking vs sitting state is, this is not entirely clear to us. Predictive processing would provide one possible explanation. Assuming the precision weighting of predictions is higher during walking, the sudden appearance of a visual stimulus might lead to stronger stimulus history prediction errors than when just sitting. But this is rather speculative.

      As a more minor point in response to our previous review, whether particular accounts represent an 'orthodox' view at present does not determine whether they raise logical issues in need of consideration. The authors may have missed that the papers in question consider mechanisms underlying the attenuation of particular pieces of information ‘from perception’. Not perceptual processing. We have one percept at any one moment in time and must understand how different population types synergistically generate that percept.

      Please excuse, the reviewer is correct, the orthodoxy of an idea is not relevant. For dubious reasons, we chose to euphemize what we actually meant to say here. With regards to circuit implementations of predictive processing (we cannot and do not intend to speak to interpretations of predictive processing that relate to conscious perception much of V1 activity is likely not consciously perceived – we assume this is what the reviewer is referring to by “we have one percept”) – the reviewers interpretation was not unorthodox, but rather incorrect (which is what we should have said). The statement that “the brain predictively ‘cancels’ expected action outcomes from perception” is incorrect in the context of sensory processing – based on both theoretical models of predictive processing, and more importantly physiological evidence. If the point was only in regards to conscious perception, we also suspect the statement is wrong, but even if it were correct, don’t see how it pertains to our work.

      Similarly a little strange is the way in which the authors aggressively defend the position that self-generated motion is 'the strongest' type of prediction. Sure, we probably experience the effects of our actions more often than ambulances. But what about objects obeying laws of gravity or others' faces being structured and moving in systematic ways? It is hard to quantify, such that presumably many scientists would be skeptical of such a claim, and it is not needed logically to justify the importance of examining mechanisms enabling action to shape perceptual processing. I'd assume it better to fight the battles you need to (and can) fight, such that the robust claims carry more weight.

      We believe it is absolutely essential for the progress of the field that we start to emphasize the differences between something that is “predictable in principle” and “predicted by the brain”. There is likely indeed a hierarchy of predictability that looks something like this:

      (1) Sensorimotor coupling

      (2) Laws of physics

      (3) Behavior of other living things

      (4) Artificial, human-made statistical relationships

      Almost all published experiments are based on the fourth type of prediction. Indeed, why not use physics simulations instead of oddballs and MMN? We absolutely should! But the field tends to revert to artificial couplings. As a direct consequence of this, the number of papers appearing recently (from both human and mouse fields), that are built on the following premise:

      (1) Expose an animal or human to an artificial coupling between A and B (e.g. an oddball, or a global oddball, or any of a myriad other constructions).

      (2) Probe for prediction error responses to the violation of the artificial coupling.

      (3) Find no prediction error responses and conclude predictive processing is wrong.

      Is utterly baffling. The fallacy here is of course the assumption that if something is predictable in principle, the brain must predict it. Thus, we are, and will continue to be strong on this point, and we think it is essential that we – as a field – are.

      Hope these comments are helpful.

      Reviewer #2 (Public review):

      Summary:

      This study investigates whether visuomotor mismatch responses can be detected in humans. By adapting paradigms from rodent studies, the authors report EEG evidence of mismatch responses during visuomotor conditions and compare them to visual-only stimulation and mismatch responses in other modalities.

      Strengths:

      Authors use a creative experimental design to elicit visuomotor mismatch responses in humans.

      The study provides an initial dataset and analytical framework that could support future research on human visuomotor prediction errors.

      Weaknesses:

      Methodological issues (e.g., volume conduction) make it difficult to confidently attribute the observed mismatch responses to activity in visual cortical regions. This could be alleviated by increasing the number of channels.

      We have added a discussion of this.

      The authors successfully demonstrate that visuomotor mismatch paradigms can, in principle, be applied in human EEG. This approach provides a translational bridge between rodent and human work on predictive processing.

      Reviewer #3 (Public review):

      Solyga, Zelechowski, and Keller present a concise report of an innovative study demonstrating clear visuomotor mismatch responses in ambulating humans, using a mobile EEG setup and virtual reality. Human subjects walked around a virtual corridor while EEGs were recorded. Occasionally, motion and visual flow were uncoupled, and this evoked a mismatch response that was strongest in occipitally placed electrodes and had a considerable signal to noise ratio. It was robust across participants and could not be explained by the visual stimulus alone.

      This is an important extension of their prior work in mice, and represents an elegant translation of those previous findings to humans, where future work can inform theories of e.g. psychiatric diseases that are believed to involve disordered predictive processing. For the most part, the authors are appropriately circumspect in their interpretations and discussions of the implications. The paper in its current form represents an important addition to the literature.

      The authors have included analyses of the auditory mismatch using temporal electrodes, referenced to Cz (and therefore should exhibit a mismatch positivity). This added data clearly and convincingly shows that the sensorimotor mismatch is, indeed, stronger than the passive auditory MMN.

      The reference electrode placed at Cz makes it is difficult to interpret relative differences between frontal and occipital electrode responses, as the occipital electrodes are placed farther away from the Cz reference than the frontal electrodes. Similarly, signal occuring cortically near the Cz reference might only appear as though it is occipitally distributed in this montage. It is common in EEG research to remontage the data to an averaged common reference in order to better interpret the scalp distributions. As the electrode coverage was sparse for some subjects, this could be challenging, and this reviewer does not feel that it is necessary to do this analysis step, or even to drastically rewrite the body of the paper. We only request that some discussion, however brief, is included in the discussion section or the methods that recommend more dense electrode coverage in the future to better interpret scalp distributions and potential meso-scale sources.

      We have added a discussion of this as suggested.

      This is just a suggestion. The authors are encouraged to analyse (and report) time-frequency power and phase locking for these mismatch responses, as is common in much of the literature (see Roach et al 2008 Schizophrenia Bulletin). This is not to say that doing so will yield insights into oscillations per se, but converting the data to the time-frequency domain provides another perspective that has some advantages. fosters translations to rodent models, as ERP peaks do not map well between species, but e.g. delta-theta power does (see Lee et al 2018 Neuropsychopharmacology; Javitt et all 2018 Schizophrenia research; Gallimore et al 2023 Cereb Ctx). Further, ERP peaks can be influenced by the actual neuroanatomy of an individual (especially for quantifying V1 responses). Time frequency analyses may aid in interpreting the "early negative deflection with a peak latency of 48 ms " finding as well. As it stands, the report is complete, and it would be acceptable if the authors chose to save this type of analysis for a future publication.

      We have added this as suggested.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      The authors have addressed most of my concerns by providing additional analyses, partly based on new data. The volume conduction issue is partly addressed based on the result showing latency differences, however to confidently assign responses to visual regions, one would need to perform recordings with a larger number of electrodes, sufficient to perform source localization. Nevertheless, the manuscript is now more solid than the previous version.

      We have now added this point to the Discussion.

      Reviewer #3 (Recommendations for the authors):

      The reviewer appreciates that the authors have carried out time-frequency analyses, and are ok with them leaving this out of this paper.

      We have now added this to the manuscript.

      Finally, in response to the participant quote "are you printing this? hi mom!" - this reviewer concedes that it does not significantly detract from the report, and, in the interest of amusement and joy, would abide its reinstatement.

      We greatly appreciate the reviewers entertaining our attempts at humor but will leave it out as originally suggested.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors wanted to better understand how the various septin-associated kinases contribute to septin organization and function in budding yeast. This question has been recently addressed by similar kinds of studies but there are still some open questions, particularly as regards to what extent the kinases may interact with and/or modify components of the contractile ring that drives cytokinesis.

      Strengths:

      This study uses sensitive imaging with good temporal and spatial resolution to monitor the localization of various proteins in living cells. Particularly informative is the use of a GFP/GFP-binding-protein "tethering" approach to ask if the requirement for one protein can be bypassed by physically tethering another protein to a third protein. Results from a yeast two-hybrid assay for measuring protein-protein interactions in vivo are buttressed by direct in vitro binding assays using purified proteins, which is important given the likelihood of "bridging" interactions between yeast proteins in the two-hybrid approach. The authors' conclusions are quite well supported by the data.

      Weaknesses:

      A control for non-specific binding is missing from the in vitro binding assay. The figures suffer sometimes from the very small text in the labels, which obscures understanding. Ultimately, while the study provides some interesting and novel insights, we still don't understand which phosphorylation events on which proteins are important for the events occurring at the molecular level, so the advance in knowledge is somewhat incremental.

      We thank the reviewer for highlighting the strengths of our imaging pipelines and protein-protein interaction data. We have now included appropriate controls for the in vitro binding assays, which demonstrate that the observed interactions are specific (Fig. 2H). We have also revised all figures to improve clarity, including increasing font sizes to enhance visibility across panels. We agree that mapping the specific phosphorylation sites regulated by these septin kinases would provide valuable mechanistic insights. However, only a few studies have addressed this direction so far (Mortensen et al., 2002; Asano et al., 2006 and Marquardt et al., 2024) [1-3]. The current study focuses on the interplay among septin-associated kinases and their role in regulation of the actomyosin machinery (AMR). In this context, we highlight several key findings:

      (i) a molecular link between the septin kinase network and AMR through physical interaction between the KA1 domain of Gin4 and F-BAR domain of Hof1 (Fig. 2F-2H and S3H);

      (ii) a kinase-independent role for Gin4 in coordinating septin organization and AMR dynamics (Fig. 3A-3E, S3F & S3G, S3I and 4F-4H);

      (iii) a novel role for Hsl1 in regulating septins and the AMR downstream of Gin4 and Elm1, potentially through plasma-membrane binding (Fig. 4I-4K, 5A-5F, 8A-8G and S9H-S9J); and

      (iv) crosstalk between Gin4 and Hsl1 that is independent of their role in the morphogenetic checkpoint (Fig. 6A-6D and S6A-S6C).

      We have clarified this scope in the Discussion section and explicitly stated that mapping these phosphorylation sites will be an important direction for future work (Lines 623-626).

      Reviewer #2 (Public review):

      Summary:

      In this paper, Bhojappa et al. provide insights into the function of septin-related kinases Elm1, Gin4, Hsl1, and Kcc4 in septin organization and actomyosin ring (AMR) structure and constriction. Their findings are both corroborative of and complementary to previous related studies.

      First, the authors provide a comparative analysis of the dynamic localization of these kinases at the bud neck, as well as a comparative analysis of defects in septin localization, splitting dynamics, AMR constriction rates, and cell morphology in kinase deficient cells. They find that septin localization and splitting kinetics, as well as AMR constriction rates, are significantly perturbed in elm1∆ and gin4∆ mutants but remain largely unaffected in hsl1∆ and kcc4∆. A similar trend is observed in terms of cell morphology and viability.

      Next, the authors focus on elm1∆ and gin4∆ cells, demonstrating that the residence time of the F-BAR protein Hof1 is significantly increased and defective in these mutants. Using yeast two-hybrid (Y2H) and in vitro binding assays, they show that the KA1 domain of Gin4 interacts with the F-BAR domain of Hof1, which may explain the cytokinesis-related functions of Elm1 and Gin4. Supporting this, they find that Gin4's role in septin localization, AMR constriction kinetics, and Hof1 bud neck localization is kinase-independent.

      The authors then conduct a series of artificial tethering experiments given their bud neck localization is mostly interdependent. They first demonstrate that artificially tethering Gin4 to the bud neck rescues the morphology defects of elm1∆ cells, with the strongest rescue observed when Gin4 was forced to interact with Hsl1-an effect that was also kinase-independent. Additionally, artificial tethering of Hsl1 to the bud neck restores the morphology of elm1∆ cells in a KA1 domain-dependent manner, suggesting that Hsl1 functions downstream of Elm1 to maintain normal cell morphology. Consistently, artificial tethering of Elm1 to the bud neck in gin4∆ cells rescues morphology defects, as well as defects in Myo1 localization and AMR constriction, but only in the presence of full-length Hsl1. The rescue fails in the absence of Hsl1 or when using a version of Hsl1 lacking the KA1 domain, which supports the role of Hsl1 downstream to Elm1 in cytokinesis.

      Strengths:

      Altogether, this study offers valuable insights into the mode of cytokinesis regulation mediated by the septin-related kinases, mainly Elm1, Gin4, and Hsl1, and would be an important contribution to the field of septins and cytokinesis after addressing current weaknesses.

      We thank the reviewer for the detailed summary and for highlighting the novel findings of our study.

      Weaknesses:

      (1) When assessing rescue of the elm1∆ phenotype, it needs to become clearer whether only morphology or also cytokinesis and septin organization are rescued.

      To clarify the extent of rescue observed in elm1Δ cells, we extended our analysis beyond morphological parameters by quantifying septin organization and AMR constriction dynamics. These analyses now show that artificial tethering partially restores septin organization and AMR constriction kinetics in elm1Δ cells in addition to improving cell morphology. These results are now described in detail in Fig. 5 and lines 373-406.

      (2) The quantification of the microscopy data does not always match up with the example images, and it's not always clear how the authors quantitatively analyzed their data.

      We revised the manuscript to clearly outline the quantification methods used for microscopy data analysis and specified the statistical tests, number of cells analyzed, and number of experimental replicates in the figure legends. We also clarified the criteria used for phenotype scoring and quantification in the Materials and Methods section. In addition, We replaced representative images where necessary to accurately reflect the quantified data throughout the revised manuscript.

      (3) The forced tethering data are key to the paper, but the lack of a summarizing table makes it difficult to grasp the full picture.

      We agree with the reviewer and have now included a new summary table (Table 1) that compiles the results of all artificial tethering experiments presented in this study, including the percentage of rescue in cellular morphology observed upon forced tethering of these kinases to the bud neck, thereby providing a clearer overview of these experiments.

      (4) Novel results and those confirming earlier results could be better distinguished.

      We have improved the overall clarity of the manuscript to distinguish novel findings from the results that corroborate previous studies, and have cited the appropriate literature throughout the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      The study by Bhojappa et al. brings new and interesting elements about the stability of the septin ring and the crosstalk between septin and actomyosin ring assemblies. The study focuses on the four kinases associated with the septin ring, Elm1p, Gin4p, Hsl1p, and Kcc4p. Elm1 and Gin4 show strong knock-out phenotypes, whereas Hsl1p and Kcc4p show weak knock-out phenotypes. The Elm1p/Kccp1p and Gin4p/Hsl1p pairs show similar timing at the bud neck. While these kinases share redundant functions, Gin4 appears to have a unique interaction with the BAR domain protein Hof1, revealing a novel direct interaction between the septin and actomyosin rings. Interestingly, the kinase activity of Gin4 is not required for its role in septin organisation and AMR constriction. The last part of the manuscript shows an original protein tethering protocol used to show that Hsl1 and its membrane binding ability are required for phenotype rescue of gin4null cells.

      Strengths:

      The combination of genetics, cell imaging, and biochemical characterization of proteinprotein interactions is attractive.

      We thank the reviewer for recognizing the significance of our findings and for the helpful suggestions.

      Weaknesses:

      (1) Imaging and data analysis is the main weakness of this manuscript. The authors must avoid manual counting and selection when easy analysis software can be used to limit bias. Instead of presenting unclear statistics of "percentage phenotypes", they need to define clear metrics to offer meaningful phenotype analysis.

      We agree that improving the quantitative rigour of the image analysis is essential for this study. Accordingly, we implemented a semi-automated image analysis workflow in the revised manuscript that defines reproducible metrics, such as aspect ratio, and reduces reliance on subjective phenotypic scoring. The inclusion of these parametric measurements enables clearer and more objective comparison of the rescued phenotypes.

      (2) This manuscript examines a very complex mechanism with four kinases of overlapping function using new data and existing literature. A clearer picture/model at the end of the manuscript that synthesizes the current knowledge would be beneficial:

      We incorporated a new representative model (Fig. 9) that integrates current knowledge in the field with our findings. This model highlights crosstalk among Elm1, Gin4, and Hsl1 as a key mechanism coordinating septin architectural transitions with AMR constriction during cytokinesis and is discussed in lines 520-536 of the revised manuscript.

      We sincerely thank all the reviewers for their insightful comments. We incorporated new results in Fig. S3A, S3B, S4A-S4F, 5A-5F, 6A-6D, S6A-S6C, 8E, 8F, S10B and 9, along with Table 1 summarizing the artificial tethering experiments in the revised manuscript. We believe that these revisions have improved the rigor of our analyses and enhanced the overall clarity of the manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 70: " abnormal cytokinesis defects": this language is redundant, either "abnormal" or "defects" would suffice:

      We thank the reviewer for identifying the redundant wording. We have rephrased the final paragraph of the Introduction section.

      Please refer to line numbers 86-97.

      (2) Currently, the final paragraph of the Introduction is an extensive, detailed summary of the results. This is unnecessary, as the Abstract and the Results sections summarize the results. Better would be a short statement of the questions addressed in the manuscript:

      We thank the reviewer for this suggestion. We have revised the final paragraph of the Introduction to remove the detailed summary of results and instead outline the key questions addressed in this study. The revised paragraph now emphasizes the knowledge gap regarding how septin-associated kinases regulate septin organization and coordinate cytokinesis. Please refer to line numbers 86-97.

      (3) In the Results, the wording of the heading associated with the first section is confusing. "and their defects during the cell cycle": it is not expected that the kinases themselves will have defects; defects may be observed in cells upon mutation of the kinases, but that is not clear from this wording:

      We thank the reviewer for this comment. The heading has been revised from “and their defects during the cell cycle” to “and defects associated with their deletions” for improved clarity.

      Please refer to line numbers 99-100.

      (4) Lines 107-109 (" they may act as molecular signals for the transition and a trigger for crosstalk of cytokinesis") are redundant with earlier lines 104-105 (" suggesting them to be a possible trigger for septin remodelling"):

      We have rephrased the text to improve clarity and avoid redundancy.

      Please refer to line number 123-125.

      (5) Lines 121-122: "Septin-associated kinases are believed to play an essential regulatory role": is the role essential, or is it regulatory? Since these kinases are not individually essential for cytokinesis, "essential" doesn't seem appropriate here:

      This has been corrected in the revised manuscript.

      (6) Figure 1 panels C and F: the font size is exceedingly small for the labels and should be greatly increased. This is true for multiple panels in Figures 2-5 and the Supplemental Figures as well:

      We thank the reviewer for highlighting this issue. We have significantly increased the font sizes of the text and the x- and y-axis labels across all figures in the manuscript.

      (7) The timing of mitotic spindle breakdown was used as a timepoint for comparison but it is not made clear in the manuscript how this was determined. Presumably, the mRuby2-Tub1 marker was visualized and the timepoint when the mitotic spindle separated into two discrete entities was called the "breakpoint" timepoint, but it would be important to better describe (and ideally show an example) of how that timepoint was determined. The only images I find with labelled tubulin are Tub1-GFP and these do not show spindle breakdown:

      We thank the reviewer for raising this point. We used GFP-Tub1 (pAFS125-GFPTUB1) and mRuby2-Tub1 (pHIS3p:mRuby2-Tub1+3′UTR::URA3) plasmids to visualize spindle dynamics across the cell cycle, and in both cases defined the spindle breakpoint as time zero. As suggested, we have now included time-lapse images of Cdc3-mCherry and GFP-Tub1 in wild-type cells in Fig. S2D to illustrate the spindle breakpoint event used for temporal alignment. Corresponding changes have also been made in the Materials and Methods section to explicitly describe this analysis.

      Please refer to line numbers 760-761.

      (8) Line 165: "F-BAR protein Hof1, which senses and induces membrane curvature" and lines 170-172 "the F-BAR protein Hof1, which is known to be associated with septin hourglass and transit to AMR during split ring trigger". It is awkward to introduce the same protein twice, in two different ways, within a few lines of each:

      We thank the reviewer for noting the redundant description. We have rephrased the text to introduce the F-BAR protein Hof1 in a more concise and streamlined manner while retaining the relevant functional information.

      Please refer to line numbers 206-208.

      (9) Throughout, it would be helpful to introduce more line breaks and organize the Results into smaller paragraphs:

      We thank the reviewer for this suggestion. We have revised the layout of the Results section and introduced additional line breaks to improve readability.

      (10) This is somewhat of a personal preference, but in the interest of transparency (and with the understanding that the 0.05 value is entirely arbitrary), would authors be willing to show actual P values in the figure panels rather than "**" or "ns", for example? Some readers may wish to apply a different standard of significance than 0.05, and not showing the P values makes this impossible. Furthermore, some readers (like this one) may interpret differently a P value of 0.044 versus "*" or 0.051 vs "ns":

      We agree with the reviewer’s suggestion. To improve the transparency of the quantitative analyses, we have modified the graphs across all main and supplementary figures to display the “actual p-values” and specified the corresponding statistical tests in the figure legends, with significance indicated by asterisks. An example is provided in the attached image showing the residence time of Inn1-mNG in wild-type and gin4Δ cells complemented with kinase-active and kinase-dead Gin4 constructs (Fig. 3E in the revised manuscript). For this analysis, significance was assessed using the nonparametric Kruskal-Wallis statistical test.

      (11) It is written that the authors "performed a Yeast Two Hybrid screen to find novel interacting partners at the bud neck" but I do not find evidence anywhere of a "screen" being performed, i.e., an unbiased search of many proteins to find a few interactors. Instead, it appears that the authors performed a two-hybrid assay to visualize interactions between a small, specific set of proteins. "Screen" should be replaced with "assay", as is currently the case in the Methods section. It is probably also valuable to point out in the text that using a yeast two-hybrid assay to assess interactions between yeast proteins has the caveat that any interactions observed could be indirect, as they may be "bridged" by endogenous yeast proteins.

      We thank the reviewer for raising this point. Our Yeast Two-Hybrid experiments were performed using a small subset of septin-associated proteins rather than as an unbiased screen. Accordingly, we have replaced the term “screen” with “assay” throughout the revised manuscript and in the Materials and Methods section.

      Please refer to line 241.

      We also agree with the limitations inherent to this assay and have now explicitly stated it in the revised manuscript, please see lines 248-255.

      (12) Panel 2H: I do not understand what the middle (as opposed to the top and bottom) blot segment represents. It is labelled "anti-HIS", like the one above it, but it is not associated with any molecular weight/ladder marker and I do not know what other species in the binding reaction in that lane would be recognized by the anti-HIS antibody. Perhaps the top segment is an "input" sample, and below it (in the middle segment) is what was bound to the beads. The figure legend is uninformative in this regard. Also, panel I in this figure is unnecessary to show, assuming that the bands shown in H are what I think they are. The blot makes the point without the need for quantification.

      As suggested by the reviewer, we have added molecular weight markers for each blot panel showing the input and bead-bound fractions. The figure legend has also been updated accordingly to clearly describe the different blot segments and experimental conditions.

      Please refer to figure legend 2H.

      In addition, as suggested by the reviewer, we have removed the quantification graph corresponding to the in-vitro binding assay from the revised manuscript.

      (13) The results in Figure 2H demonstrate that the Gin4-KA1 fragment is not non-specifically "sticky", because it does not bind GST alone, but there is no demonstration that binding by the Hof1 fragment is specific because there is no equivalent negative control for binding:

      We thank the reviewer for this suggestion. We repeated the in-vitro binding assay using 6His-bdSUMO as a negative control alongside 6His-bdSUMO-Gin4<sup>KA1</sup> to demonstrate binding specificity of the Hof1 fragment. The Hof1 N-terminal F-BAR fragment did not pull down the control 6His-bdSUMO fragment but specifically pulled down 6His-bdSUMO-Gin4<sup>KA1</sup> under identical experimental conditions (Fig. 2H), confirming the specificity of the interaction between Hof1 F-BAR domain and Gin4KA1.

      The corresponding text and results have been updated in the revised manuscript.

      Please refer to Fig. 2H and lines 259-263.

      (14) Line 257-258: "Localisation via Hsl1 is necessary to rescue the morphological defects exhibited by Δelm1 cells partially": what does "partially" refer to here? To the rescue, or the defects?

      The term “partially” refers to the extent of rescue. Specifically, elongated cell morphology was rescued in 63.75% of the elm1Δ cell population, rather than in all cells, upon artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP.

      (15) Lines 292-293: "can restore the morphological defects": this wording is unclear. "Restore" means "return to a former condition", which in this case would be normal cellular morphology, not defective cellular morphology. Similarly, see line 304: "While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering": presumably the proper localization was restored, not the mislocalization:

      In Lines 292-293, by phrase “can restore the morphological defects” was intended to indicate rescue of the elongated/clumped morphology associated with gin4Δ cells upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP. We have now rephrased this sentence as: “can rescue the elongated/clumped phenotype exhibited by gin4Δ cells”.

      Please refer to line numbers 481-482.

      Similarly, in Line 304, the statement “While Myo1-3xmCherry mislocalisation was restored upon Elm1-GFP tethering” referred to rescue of the Myo1 mislocalization phenotype observed in gin4Δ cells. We have rephrased this sentence as: “However, Myo1-3xmCherry localization was restored to the bud neck upon artificial tethering of Elm1-GFP in gin4Δ cells”.

      Please refer to lines 506-508 in the revised manuscript.

      (16) Lines 295-296: "We find that Elm1-GFP tethering via Shs1-GBP, Bud4-GBP, and Hsl1-GBP" this should be "or", not "and":

      Thank you for pointing this out. We have now corrected the text accordingly.

      Please refer to line number 485.

      (17) Lines 311 and 312 refer to "Inn1-3xmCherry lifetime" but previously Inn1 residence time was measured. Since fluorescence lifetime is a distinct kind of measurement/ assay, it seems important to clarify here what kind of experimental data are being referred to:

      The term “Inn1-3xmCherry lifetime” was intended to describe the residence time of Inn1 at the cell division site, defined as the interval between the initial appearance of the Inn1 fluorescence signal and its complete disappearance during cytokinesis. For clarity and consistency, we have replaced the term “lifetime” with “residence time” in the revised manuscript.

      Please refer to the line numbers 511 and 512.

      (18) Discussion: "Septins are considered as the fourth cytoskeletal elements due to their extensive structural and functional diversity." This sentence is confusing, as it seems to imply that what defines a protein as being "cytoskeletal" is structural and functional diversity rather than anything to do with forming filaments, etc. The rest of this first paragraph of the Discussion also sounds like a summary of the background, is quite redundant with the Introduction, and should be shortened:

      We thank the reviewer for this suggestion. We have rewritten the first paragraph of the Discussion to improve clarity, reduce redundancy with the Introduction, and better emphasize the main findings of the study.

      Please refer to the line numbers 538-552 in the Discussion section.

      (19) Lines 365-366: "Gin4 and Hof1 are synthetic lethal": this should be revised to "gin4∆ and hof1∆ are synthetic lethal". This is a good place to point out that standard yeast nomenclature inserts the ∆ symbol after the gene name, not before (as is done in E. coli genetics, for example):

      We thank the reviewer for this suggestion. We have replaced “Gin4 and Hof1 are synthetic lethal” with “gin4Δ and hof1Δ are synthetic lethal” in the revised manuscript.

      Please refer to line number 576.

      In addition, we have now consistently placed the Δ symbol after the gene throughout the manuscript in accordance with standard yeast nomenclature.

      (20) Line 399: "We also performed an extensive GFP-GBP screens": again, here "screen" implies that a large collection of genes/proteins were assayed, perhaps in an unbiased way, which does not accurately portray what was actually done, which was an extensive tethering study using GFP-GBP:

      We thank the reviewer for this suggestion. We have replaced the term “GFP-GBP tethering screen” with “GFP-GBP tethering assay” throughout the revised manuscript. In addition, we have included a brief description of the specificity and functionality of GBP nanobody and its application in the GFP-GBP tethering strategy, extensively used in this study.

      Please refer to line numbers 326-336.

      Reviewer #2 (Recommendations for the authors):

      Major points:

      (1) Analysis of morphological defects of elm1∆ does not directly reflect the defects in cytokinesis and septin organization. For example, the deletion of SWE1 rescues the morphological defects of elm1∆ cells but not the cytokinesis or septin mislocalization (Bouquin et al., 2000). Considering this, the authors should address whether artificial tethering of Hsl1 to Gin4 in elm1∆ cells rescues the septin and cytokinesis defects or just the morphology. Is the role of Hsl1 in cytokinesis dependent on its role in the morphogenesis checkpoint? Can the authors comment on how much the defects observed by Hsl1 tethering to the bud neck may be a result of bypassing the morphogenesis checkpoint?:

      We thank the reviewer for this important point. We performed time-lapse imaging of Cdc3-mCherry in strains where Gin4-GFP partially rescued the elongated phenotype of elm1Δ cells (63.75%) when tethered to the bud neck via Hsl1-GBP. Under these conditions, 64.29% of cells showed rescue of Cdc3-mCherry mislocalization, and Gin4 localization itself was restored to the bud neck in 57.85% of cells. We also examined Myo1-ymScarletI dynamics, while 76.92% of untethered elm1Δ cells displayed Myo1 mislocalization to the bud cortex, this was reduced to 11.36% upon Gin4-GFP tethering via Hsl1-GBP. Together, these results indicate that the morphological rescue observed in elm1Δ cells is accompanied by restoration of normal septin organization and AMR dynamics.

      Previous work (Bouquin et al., 2000) [4] showed that Swe1 deletion rescues cell elongation in elm1Δ cells without restoring septin organization. Consistent with this, we found that 60.66% of elm1Δ swe1Δ cells exhibited a round morphology, but tethering of Gin4-GFP to the bud neck via Hsl1-GBP in elm1Δ swe1Δ background did not further enhance morphological rescue. These results suggest that the rescue of cellular morphology observed in our tethering experiments may, atleast in part, depend on Hsl1-mediated regulation of the morphogenesis checkpoint.

      Importantly, despite the lack of additional morphological rescue, a clear restoration of septin localization was observed when Gin4-GFP was tethered to the bud neck via Hsl1-GBP in elm1Δ swe1Δ cells. Overall, these results suggest that while Hsl1-dependent morphogenesis checkpoint regulation may contribute to cell shape rescue, the restoration of septin organization is independent of Hsl1’s function in morphogenesis checkpoint and instead reflects a direct requirement for Gin4 and Hsl1 at the bud neck.

      Please refer to Figures 5, 6, and S6 of the revised manuscript for these additional data.

      (2) As suggested by the authors, the interaction of the Gin4-KA1 domain with the FBAR domain of Hof1 may explain the cytokinesis-related functions of Gin4. As an orthogonal approach, how does KA1 domain deletion of Gin4 affect cytokinesis and Hof1 bud neck localization?

      We thank the reviewer for this suggestion. We first examined the bud neck localization of Gin4-ka1Δ-GFP in comparison with full-length Gin4-GFP. We observed that the localization kinetics of Gin4-ka1Δ-GFP were significantly altered relative to the full-length protein, with reduced recruitment and earlier removal from the bud neck. We also analyzed the localization kinetics of Hof1-mNG in both gin4-ka1Δ and gin4Δ cells. Our results show that Hof1-mNG displays increased residence time and altered accumulation kinetics during cytokinesis in both genetic backgrounds. Thus, loss of the KA1 domain phenocopies loss of the full-length Gin4 and is consistent with disruption of the physical interaction between Gin4 and Hof1.

      Please refer to Figure S4 for these results in the revised manuscript.

      (3) The authors state that "Elm1 and Kcc4 were present at lower abundance at the bud neck (Fig S1A-D) compared to the higher abundance of Gin4 and Hsl1, as observed in their fluorescence intensities (Figures S1B-C) ". This is not evident in the figures. The authors should show a quantification of how they judged abundance at the bud neck:

      We thank the reviewer for this question. Quantification of septin kinase fluorescence intensity at the bud neck was performed using the established protocol for measuring protein accumulation kinetics described by Okada et. al. 2020 [5]. Time-lapse imaging for kinetic analysis of septin-associated kinases during bud emergence shown in Fig. S1A-S1D was carried out using a point-scanning confocal microscope with a 100×oilimmersion objective. Different laser intensities were required because the fluorescence signals of Kcc4 and Elm1 were comparatively weak and not readily detectable above cellular background under the imaging conditions used for Gin4 and Hsl1. The images shown in Fig. S1A-S1D are therefore displayed using differential contrast settings to facilitate visualization. We have now explicitly clarified this in the figure legend.

      A more direct comparison of septin kinase abundance at the bud neck is now provided in Fig. S1E-F, where localization kinetics during the HDR transition/septin remodelling stage were captured using a laser-scanning spinning-disk microscope under similar imaging conditions. We have additionally included raw fluorescence intensity profiles during the HDR transition to better illustrate the relative abundance of these kinases at the bud neck during cytokinesis.

      Please refer to Figures S1E and S1F in the revised manuscript for the updated images and quantitative analyses.

      (4) How did the authors determine G1 and M-phase in the experiments shown in Figures S1A-D? Can the authors mark these phases on the timelapse images? Also, How do the authors explain the different behaviour of Cdc3 in graphs S1A-D among different strains?

      We thank the reviewer for this comment. Cell cycle stages were initially inferred based on bud size, where kinase accumulation at the bud neck correspond to bud emergence (small bud, G1), and kinase disappearance coincided with septin splitting (large bud, M phase). However, we agree that accurate assignment of cell cycle stages would require specific cell cycle markers. To avoid confusion, we have removed the cellcycle-specific stage assignments from the Results section and describe the kinetics relative to t=0 (bud emergence).

      The differential dynamics observed in the Cdc3-mCherry kinetic profiles likely reflect heterogeneity within the the cellular population. To address this, we combined the normalized fluorescence intensity profiles of Cdc3-mCherry from the strains expressing GFP-tagged septin kinases during bud emergence and have included this data as reference (Author response image 1).

      Author response image 1.

      Plot showing spatiotemporal kinetics of Cdc3-mCherry in strains expressing either Elm1-GFP, or Gin4-GFP, or Hsl1-GFP, or Kcc4-GFP.

      (5) Forced tethering based experiments are one of the key sets of experiments for this work, but it is difficult to have a comprehensive understanding of all the data considering how large the data set is and how dispersed it is in the supplemental and main figures (Figures 3-4-5 and Figures S4-S5-S6). It would be helpful to provide a table summarizing the tested forced tethering’s and the phenotypic outcome in the tested yeast strains (Wt/mutant):

      We thank the reviewer for recognising the extensive dataset generated from the artificial tethering experiments and for suggesting the inclusion of a summary table. We have now added Table 1, which summarizes the proteins used in the GFP-GBP artificial tethering experiments, their genetic backgrounds, the total number of cells quantified across three independent replicates, and the phenotypic outcomes associated with bud neck tethering under each condition.

      Please refer to Table 1 and lines 342, 361, 451, 453, 456 and 492 in the revised manuscript.

      (6) In the introduction section, it would help the reader to provide more information on already known molecular roles of septin-associated kinases in septin organization and AMR. Later in the results section (i.e. Figure S1, S2, and S6), the authors extensively explain and show data that independently corroborate some earlier findings, which makes it difficult for the reader to distinguish novel findings from the repeated findings. I suggest shortening the text for the corroborative results, which will help to put more emphasis on their novel findings:

      We thank the reviewer for this suggestion. We have now included the canonical roles of these four septin-associated kinases in the Introduction section of the revised manuscript.

      Please refer to line numbers 79-85.

      We have also revised sections describing corroborative findings and explicitly cited previous studies wherever relevant in the Results section to better distinguish previously established observations from the novel findings presented in this work.

      Other minor comments:

      (1) In Figure 1D-1F, also show the data for hls1∆ and kcc4∆ - which are shown in S3AC in the current version:

      We have now included the Inn1-mNG residence time in hsl1Δ and kcc4Δ cells, alongside elm1Δ and gin4Δ in Fig. 1E.

      Please refer to Fig. 1E in the revised manuscript.

      (2) The authors should be more careful in interpreting their negative Y2H data in Figure 2F.

      We have now explicitly discussed the caveats and inherent limitations of the Yeast Two-Hybrid assay in the manuscript.

      Please refer to line numbers 248-255.

      (3) Please provide quantification for Figure 5A, Figure S6I:

      We have added quantitative analyses showing rescue of Cdc3-mCherry and Myo13xmCherry mislocalization in gin4Δ and gin4Δ hsl1Δ strains upon artificial tethering of Elm1-GFP to the bud neck via Shs1-GBP.

      Please refer to Figures 8E, 8F, and S10B in the revised manuscript.

      (4) In Figure S6: label is missing "∆" in front of hsl1:

      We thank the reviewer for pointing out this error. We have corrected the labels accordingly.

      (5) As common consensus on yeast gene nomenclature, I suggest the use of "gene∆" instead of "∆gene":

      We thank the reviewer for this suggestion. We have now consistently placed the Δ symbol after deleted gene names throughout the manuscript in accordance with standard yeast nomenclature.

      (6) Lines (535-536): min(distribution) and max(distribution) in the formula is confusing. Clarify it or if possible use "minimum value", "maximum value" instead:

      We thank the reviewer for this suggestion. We have replaced the term “distribution” with “value” in the formula for protein accumulation kinetics analysis.

      Please refer to the updated formula in the Materials and Methods section (Lines 752753).

      (7) In line 86, "Dynamics of Septin-associated kinases and their defects during the cell cycle": Change the title as it is not clear what is meant by "their defects" given the discussed results under this title:

      We thank the reviewer for this suggestion. We have revised the section heading from “and their defects during the cell cycle” to “and defects associated with their deletions”.

      Please refer to line numbers 99-100.

      (8) On the Hof1-mNG image (Fig2C), show the line used for the line scan profile. Additionally, a similar line-scan profile could be useful in Figure S3G:

      As suggested by Reviewer 3, we removed the line-scan analysis from Figure 2 in the revised manuscript because Hof1 ring organization showed substantial heterogeneity across cells, making line-scan analysis difficult to interpret reliably.

      Reviewer #3 (Recommendations for the authors):

      Major points:

      (1) The % phenotype units are terrible. With no explanation, we do not really know whether they represent the percentage of cells that have a particular phenotype, or whether they correspond to a metric that measures some deviation between normal and extreme phenotypes. I would strongly recommend using precise quantitative metrics systematically (i.e. intensities, aspect ratios, division times, etc.) to properly quantify phenotypes:

      We thank the reviewer for suggesting the inclusion of precise quantitative metrics to assess phenotypic differences in the GFP-GBP tethering experiments. In the revised manuscript, we adopted quantification workflows that have been extensively validated and widely used in the literature, including those reported by Marquardt et al., 2024 (Fig. 7B and 7D) [2] from the Bi Lab. In response to the reviewer’s suggestion, we have now incorporated additional quantitative measurements, including cell area and aspect ratio (defined as the ratio of the cell’s major axis to the minor axis), for the experimental datasets presented in the manuscript.

      In the main figures, we now include aspect ratio quantification, while additional parameters are provided for the reviewer’s reference. We also quantified the fluorescence intensity of tethered proteins at the large bud neck and present these data together with the aspect ratio analysis in Figure 4 for elm1Δ cells in which Gin4GFP is artificially tethered to the bud neck via Hsl1-GBP. These quantitative analysis corroborates our qualitative observations and further strengthens our conclusions. Please refer to Figures 4D and 4E as representative examples.

      We have also changed the y-axis labels throughout the revised manuscript. For example, the y-axis in Fig. 7B is now labelled as “Cells exhibiting round morphology (%)”. Please refer to Fig. 4H, 4K, 7C, 8D, S5G, S6C, S7C, S9G and S9J for inclusion of aspect ratio quantification. Statistical analyses for the represented graphs were performed using Kruskal-Wallis nonparametric test, (N=3, n>150 cells/strain) (*: p<0.05, **: p<0.01, ****: p<0.0001, ns: p>0.05).

      Author response image 2.

      (2) Some quantitative analyses were performed manually where simple automated analysis should be performed to provide unbiased, accurate quantification:

      We fully agree with the reviewer that automated image analysis approaches, such as segmentation-based methods, are generally preferred for minimizing bias in morphological quantification. However, elm1Δ and gin4Δ cells exhibit severe phenotypes, including pronounced elongation and clumping, which makes reliable automated segmentation technically challenging for accurate quantification of parameters such as aspect ratio and cell area. For this reason, we used manual annotation for these analyses, as this approach enabled accurate delineation of individual cell and reliable measurements of morphological parameters such as cell area and size across the datasets despite being more time-consuming.

      Please refer lines 769-775 in the Materials and Methods section.

      (3) The "tethering" data also lack clear quantification. The authors should properly quantify the average intensity of Hsl1-GFP at the bud neck in each condition and correlate the results with cell aspect ratios or any other relevant parameters. For example, when comparing elm1null and elm1null Bud4-GBP with elm1null Kcc4-GBP, the visual impression is that as much Hsl1-GFP protein is recruited to the bud neck, whereas the phenotypes are dramatically different:

      We thank the reviewer for pointing this out. We have revised the image representation to facilitate clearer interpretation of the tethering experiments. In addition, we performed the key GFP-GBP tethering experiments using GBP-ymScarletI constructs, allowing direct visualisation of both the GFP-tagged protein and the GBP-tagged partner at the bud neck following tethering. We have included quantitative analyses of cellular morphology, raw fluorescence intensities of GFP-tagged proteins at the large bud neck, and the corresponding aspect ratio measurements for these updated datasets (Author response images 3, 4, 5). We have also included a summary table (Author response table 1) compiling these quantitative results for easier comparision.

      (4) I am very confused by Figure 3 which shows normal localization of Gin4-GFP in elm1null cells and seems to contradict other claims in the manuscript. This is very problematic for the interpretation of most of the "tethering" data:

      We understand the reviewer’s concern and have replaced the representative images of Gin4-GFP in elm1Δ cells in Figure 4B. Although Gin4-GFP is initially recruited to the presumptive bud neck during bud emergence in elm1Δ cells, it subsequently becomes mislocalized to the bud cortex during early cell cycle stages, resulting in reduced bud neck localization. As the cell cycle progresses, the Gin4-GFP signal at the bud neck decreases substantially in elm1Δ cells while remaining stable in wild-type cells until its departure prior to septin HDR remodelling (Fig. S5A-S5D).

      (5) Figure 6 is neither explained in the text nor in its legend. Could the authors explain the model and offer a comprehensive picture of the current knowledge?

      We have simplified the representative model to more clearly distinguish previously established knowledge from the findings presented in this study. Based on our results, We propose that Hsl1 functions both downstream of and in coordination with Elm1 and Gin4 to regulate septin stability and the timely execution of cytokinesis. Deletion of Elm1 disrupts the normal localization and crosstalk between Gin4 and Hsl1 at the bud neck, leading to septin mislocalization and misregulation of AMR dynamics, thereby revealing a previously uncharacterized role for Hsl1 in cytokinesis. The Results section has also been updated to reflect the revised model.

      Please refer to Figure 9 and lines 520-536.

      (6) Could the authors provide information about the double/triple mutant kinase phenotypes to clarify the overlap of functions among them?:

      Barral et al., 1999 [6] reported that individual deletions of Hsl1 and Gin4 result in mild cytokinetic defects, whereas deletion of Kcc4 does not produce any striking phenotype compared to wild-type cells. In contrast, the hsl1Δ gin4Δ kcc4Δ triple mutant remains viable but exhibits severe morphological abnormalities, including branched chains of elongated cells with defective cell separation. These mutants also display aberrant septin organization at the bud neck, characterized by irregular patch-like structures. Analysis of double mutants (hsl1Δ gin4Δ, gin4Δ kcc4Δ, and hsl1Δ kcc4Δ) revealed intermediate phenotypes between the corresponding single and triple mutants, with the hsl1Δ gin4Δ combination showing the strongest defects. Together, these findings suggest that the Nim1-related kinases function redundantly to regulate Swe1 activity and maintain septin architecture at the bud neck.

      Further supporting this model, Bouquin et al., 2000 [4] demonstrated that Elm1 operates independently of the Nim1-related kinases in controlling septin organization. The hsl1Δ gin4Δ kcc4Δ elm1Δ quadruple mutant exhibits severe septin localization defects and strong growth defects, in contrast to the elm1Δ single mutant, which primarily displays septin mislocalization from the bud neck to the bud cortex. These findings indicate that the combined activity of these kinases is essential for proper septin anchorage at the division plane and for assembly of the septin ring.

      Importantly, the progressively stronger phenotypes observed in double, triple and quadruple mutants also suggest that these kinases retain partially specialized functions at the bud neck. Consistent with this framework, our results support a model in which Nim1-related kinases function redundantly to regulate septin architecture and cytokinesis, likely through modulation of the AMR machinery. Because the localization of these kinases appears interdependent, as reported previously (Marquardt et al., 2020; Marquardt et al., 2024) [2,7] and corroborated by our findings, interpretation of mutant phenotypes remains complex and future studies will be required to delineate their individual contributions more precisely.

      Minor points:

      (1) Abbreviations are not defined in the manuscript.

      We have now expanded and defined all abbreviations throughout the manuscript.

      (2) Some of the writing in the figures is too small. Please make sure that a minimal size of letters/numbers is respected:

      We thank the reviewer for raising this issue. We have enlarged the figure labels and axis labels throughout the revised manuscript to improve readability.

      (3) Figure 2D. I am not sure that the line scans bring any useful information as the rings in mutant cells are quite heterogenous. This panel is also not cited in the text. Please make sure that every panel is cited at least once:

      We agree with the reviewer regarding the heterogeneity observed in the Hof1 ring organization in mutant cells and have therefore removed the line-scan analysis from the revised manuscript.

      (4) Knocked-out genes are written incorrectly. Please use the usual yeast nomenclature:

      We thank the reviewer for this suggestion. We have corrected the nomenclature for all deleted genes throughout the manuscript in accordance with standard yeast nomenclature.

      Author response image 3.

      Artificial tethering of Gin4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Gin4-GFP via Shs1-GBP-ymScarletI, Hsl1-GBP-ymScarletI, Bud4-GBPymScarletI and Bni5-GBP-ymScarletI in elm1Δ cells. DC*=Differential contrast. Scale bar5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple-comparison test (**: p<0.01, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=397, elm1Δ: n=408, elm1Δ-Shs1-GBPymScarletI: n=467, elm1Δ-Hsl1-GBP-ymScarletI: n=462, elm1Δ-Bud4-GBP-ymScarletI: n=401 and elm1Δ-Bni5-GBP-ymScarletI: n=317 cells). (C) Quantification of aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, n>170 cells/strain). (D) Graph depicting the raw fluorescence intensity of Gin4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=166, elm1Δ: n=177, elm1Δ-Shs1-GBP-YmScarletI: n=176, elm1Δ-Hsl1-GBPymScarletI: n=188, elm1Δ-Bud4-GBP-ymScarletI: n=185 and elm1Δ-Bni5-GBP-ymScarletI: n=151 cells).

      Author response image 4.

      Artificial tethering of Hsl1-GFP to the bud neck via septins or Nim1-related kinases rescues cellular morphology in elm1Δ cells. (A) Representative images showing the relocalization of Hsl1-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI and Gin4GBP-ymScarletI. Scale bar-5µm. (B) Bar graph representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), one-way ANOVA Tukey’s multiple comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=508, elm1Δ: n=418, elm1ΔShs1-GBP-ymScarletI: n=535 and elm1Δ-Gin4-GBP-ymScarletI: n=482 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (*: p<0.05, ****: p<0.0001), (N=3, n>165 cells/strain). (D) Quantification of raw fluorescence intensity of Hsl1-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=156, elm1Δ: n=166, elm1Δ-Shs1-GBP-ymScarletI: n=169 and elm1Δ-Gin4-GBP-ymScarletI: n=165 cells.

      Author response image 5.

      Targeted localization of Kcc4-GFP to the bud neck via Hsl1-GBP-ymScarletI rescues cellular morphology in elm1Δ cells. (A) Representative images showing artificial tethering of Kcc4-GFP to the bud neck in elm1Δ cells via Shs1-GBP-ymScarletI, Hsl1-GBPymScarletI and Gin4-GBP-ymScarletI. Scale bar-5µm. (B) Quantitative analysis representing the percentage of cells exhibiting round morphology in the indicated strains shown in (A), oneway ANOVA Tukey’s multiple-comparison test (****: p<0.0001, ns: p>0.05), (N=3, wildtype: n=443, elm1Δ: n=489, elm1Δ-Shs1-GBP-ymScarletI: n=312, elm1Δ-Hsl1-GBP-ymScarletI: n=563 and elm1Δ-Gin4-GBP-ymScarletI: n=337 cells). (C) Quantification of the aspect ratios in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (****: p<0.0001, ns: p>0.05), (N=3, n>165 cells/strain). (D) Quantification for the raw fluorescence intensity of Kcc4-GFP at the large bud in the indicated strains shown in (A), Kruskal-Wallis nonparametric statistical test (**: p<0.01, ***: p<0.001, ****: p<0.0001), (N=3, wildtype: n=163, elm1Δ: n=168 elm1Δ-Shs1-GBP-ymScarletI: n=161, elm1Δ-Hsl1-GBP-ymScarletI: n=172 and elm1Δ-Gin4-GBP-ymScarletI: n=150 cells).

      Author response table 1.

      Summary table showing rescue of elongated morphology in elm1Δ cells upon forced recruitment of Nim1-related kinases via septins or its related kinases tagged with GBPymScarletI.

      Additional changes:

      The graph in Fig. S3D (revised preprint) has been updated to reflect a slight increase in the Chs2-mNG residence time in both the elm1Δ and gin4Δ strains, whereas our previous version indicated a delay only in the gin4Δ strain. Because the elm1Δ strain exhibited a more pronounced phenotype than the gin4Δ strain, we re-examined the analysis. The revised results show that the residence time of Chs2 during cytokinesis is modestly prolonged by approximately 2 minutes in both backgrounds. Accordingly, the graph and statistical analyses have been updated.

      References:

      (1) Mortensen, E.M., McDonald, H., Yates, J., and Kellogg, D.R. (2002). Cell Cycle-dependent Assembly of a Gin4-Septin Complex. Molecular Biology of the Cell 13, 2091-2105. 10.1091/mbc.01-10-0500.

      (2) Marquardt, J., Chen, X., and Bi, E. (2024). Reciprocal regulation by Elm1 and Gin4 controls septin hourglass assembly and remodeling. J Cell Biol 223. 10.1083/jcb.202308143.

      (3) Asano, S., Park, J.E., Yu, L.R., Zhou, M., Sakchaisri, K., Park, C.J., Kang, Y.H., Thorner, J., Veenstra, T.D., and Lee, K.S. (2006). Direct phosphorylation and activation of a Nim1-related kinase Gin4 by Elm1 in budding yeast. J Biol Chem 281, 2709027098. 10.1074/jbc.M601483200.

      (4) Bouquin, N., Barral, Y., Courbeyrette, R., Blondel, M., Snyder, M., and Mann, C. (2000). Regulation of cytokinesis by the Elm1 protein kinase in Saccharomyces cerevisiae. Journal of Cell Science 113, 1435-1445. 10.1242/jcs.113.8.1435.

      (5) Okada, H., MacTaggart, B., and Bi, E. (2021). Analysis of local protein accumulation kinetics by live-cell imaging in yeast systems. STAR Protoc 2, 100733. 10.1016/j.xpro.2021.100733.

      (6) Barral, Y., Parra, M., Bidlingmaier, S., and Snyder, M. (1999 Jan 15). Nim1-related kinases coordinate cell cycle progression with the organization of the peripheral cytoskeleton in yeast. Genes & Development 13. 10.1101/gad.13.2.176.

      (7) Marquardt, J., Yao, L.L., Okada, H., Svitkina, T., and Bi, E. (2020). The LKB1-like Kinase Elm1 Controls Septin Hourglass Assembly and Stability by Regulating Filament Pairing. Curr Biol 30, 2386-2394 e2384. 10.1016/j.cub.2020.04.035.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The extent to which P. falciparum liver stage parasites export proteins into the host cell is unclear. Most blood-stage exported proteins tested in liver stages were not exported. An exception is LISP2, which is exported in P. berghei but not P. falciparum liver stages. While the machinery for export is present in liver stages, efforts to demonstrate export have so far been mostly unsuccessful. Parasite proteins exported during the liver stage could be presented by MHC and thereby become the target of immune control, an incentive to study liver stage export and identify proteins exported during this stage. However, particularly for P. falciparum, it is very difficult to study liver stages.

      This work studies LSA3 in P. falciparum blood and liver stages. The authors show that this protein is exported into the host cell in blood stages, but in liver stages, no or only very little export was detected. A disruption of LSA3 reduced liver stage load in a humanized mouse model, indicating this protein contributes to efficient development of the parasites in the liver.

      The paper also studies the localization of LSA3 in blood stages and uses a known inhibitor to show that it is processed by plasmepsin 5, a protease important for protein trafficking. The work also shows that LSA3 is not needed for passage through the mosquito.

      Strengths:

      The main strength of this work is the use of the humanized mouse model to study liver stages of P. falciparum, which is technically challenging and requires specialized facilities. The biochemical analysis of LSA3 localization and processing by plasmepsin 5 is thorough and mostly overcame adverse issues such as a cross-reactive antibody and the negative influence of the GFP-tag on LSA3 trafficking. The mosquito stage analysis is also notable, as these kinds of studies are difficult with P. falciparum. However, there was no evidence for a function of LSA3 in mosquito stages.

      We thank the reviewer for their perspective on the strengths of the study.

      Weaknesses:

      The cross-reactivity of the antibody, together with the co-infection strategy, prevents reliable assessment of LSA3 localization in liver stages. Despite this, it seems LSA3 is not exported in liver stages, and the paper does not bring us closer to the original goal of finding an exported liver stage protein.

      While the localization analysis in blood stages is well done and thorough, the advance is somewhat limited. LSA3 may be in structures like J dots, but this hypothesis was not tested. Although parasites with a disrupted LSA3 were generated, the function of this protein was not explored. Given that a previous publication found some inhibitory effect of LSA3 antibodies on blood stage growth, a comparison of the growth of the LSA3 disruption clones with the parent would have been very welcome and easy to do. At this point, LSA3 is one more of many proteins exported in blood stages for which the function remains unclear.

      It might be possible to refine some of the conclusions. The impact on liver stage development is interesting, but which phase of the liver stage is affected, and the phenotype remains largely unknown. The co-infection (WT together with LSA3 mutant) has the advantage of a direct comparison of the mutant with the control in the same liver, but complicates phenotypic analysis if the LSA3 antibody is also cross-reactive in liver stages. This issue adds a question mark to the shown localization and precludes phenotypic comparisons. The authors write that they do not know if the cross-reactive protein is expressed at that stage. But this should be immediately evident from the mixed WT/mutant infection. If all cells are positive for LSA3, there is a cross-reaction. If about half of the cells are negative, there isn't. In the latter case, the localization shown in the paper is indeed LSA3, and morphological differences between WT and LSA3 disruption could be assessed without additional experiments.

      We thank the reviewer for their comments. While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Significance:

      The conclusion from the paper that "our study presents just the second PEXEL protein so far identified as important for normal P. falciparum liver-stage development and confirms the hypothesized potential of exported proteins as malaria vaccine candidates" is partially misleading. Neither LISP2 nor LSA3 seems to be exported in P. falciparum liver stages, and we can't confirm the potential of vaccines with proteins exported in this stage. LSA3 is still important and may still be the target of the immune response, but based on this work, probably not due to export in liver stages.

      We thank the reviewer for the comment. We would like to emphasize the possibility that proteins localized at the PVM may be considered exported ‘if’ part or all of the protein (eg, a domain) faces the host cell lumen from the hepatocyte. We have not shown this to be the case for LSA3 or LISP2 but that possibility remains open. Nonetheless, LISP2 is exported (by P. berghei liver stages) and LSA3 is exported (by P. falciparum blood stages); both are exported proteins.

      Reviewer #2 (Public review):

      Summary:

      Immunogenic Plasmodium falciparum proteins that could be targeted to prevent parasite development in the liver are of significant interest for novel anti-malarial vaccine development. In this study, McConville et al evaluate the trafficking and functional importance of LSA3, a protein expressed in the blood and liver stages and previously shown to provide protection in immunized chimpanzees. LSA3 contains a PEXEL motif, but the authors have previously shown that this protein does not appear to be exported beyond the PVM in the liver stage (McConville et al, PNAS 2024). However, LSA3 trafficking and functional importance have not been comprehensively evaluated across stages. In the present study, the authors find that blood stage LSA3 undergoes PEXEL processing, and a portion of the protein is exported into the erythrocyte, where it localizes to punctate structures distinct from Maurer's clefts. Using a knockout mutant, LSA3 is shown to be dispensable for blood and mosquito stages but important to liver-stage development. Collectively, these results validate LSA3 as a liver-stage target and place it among several other PEXEL proteins that display differential trafficking beyond the PVM in the erythrocyte but not the hepatocyte.

      Strengths:

      The authors present a thorough analysis of LSA3 trafficking in the blood stage. PEXEL processing by Plasmepsin 5 is clearly demonstrated through a combination of mini LSA3-GFP reporters and Plasmepsin 5 inhibitors. Importantly, an LSA3 knockout mutant is used to show that the LSA3-C anti-sera also react with additional, unidentified parasite proteins in the blood stage. Nonetheless, comparison between the WT and KO parasites clearly indicates that a portion of LSA3 is exported into the erythrocyte, which is further supported by protease-protection assays with fractionated iRBCs. This contrasts with the liver stage, where LSA3 does not appear to traffic beyond the PVM, similar to what has been observed for other PEXEL proteins in the rodent malaria model.

      This study provides the first direct analysis of LSA3 function by reverse genetics, showing this protein is important for liver stage development in chimeric human liver mice. Several PEXEL proteins in P. berghei have been shown to be exported into the host cell in the blood stage, but do not appear to cross the PVM in the liver stage. These observations reinforce that even without detectable export into the hepatocyte, PEXEL proteins play critical roles during liver stage development.

      We thank the reviewer for their feedback regarding the strengths of the paper. 

      Weaknesses:

      A previous study reported that anti-LSA3 antibodies inhibit blood-stage growth, suggesting a role for LSA3 during erythrocyte infection. While the authors carefully evaluate the LSA3 mutant in mosquito and liver stages, the impact on blood stage fitness is not tested. While the knockout shows LSA3 is not essential in the blood stage, its importance during erythrocyte infection remains unclear.

      The authors previously reported that anti-LSA3-C signal in the liver stage localizes within the parasite and at the parasite periphery but is not exported into the hepatocyte. In the present study, it is shown that anti-LSA3-C reacts with other parasite proteins beyond LSA3 in the blood stage, and this may also occur in the liver stage. However, since liver-stage IFAs were only performed on samples co-infected with both WT and ∆LSA3 parasites, non-specific anti-LSA3C reactivity at this stage could not be determined, and the localization of LSA3 in the liver stage remains somewhat unclear.

      We thank the reviewer for their comments. The phenotype of the NF54 DLSA3 mutant generated in this study at the blood stage was underway (by an independent lab in collaboration with us) and we are happy to disclose that the outcomes were recently published (May 2026) in an accompanying manuscript (PMID: 41135800). While the LSA3-C antibody may cross-react with another parasite protein(s) in addition to binding LSA3 itself, we observed no strong evidence that this antibody localized beyond the liver-stage PVM, indicating that LSA3 is likely not targeted to the host cell compartment. We cannot exclude the possibility that a domain of LSA3 faces the hepatocyte lumen from this membrane and thus may be considered exported though follow-up studies are required (and are very challenging) to answer it. We completely agree that independently infected humanized mice would be helpful to address further remaining questions around the localization and temporal phenotype for LSA3 essentiality, which again will require follow up studies. In the present study, we intended to address whether LSA3 is important functionally, as this had not been reported.

      Reviewer #3 (Public review):

      Summary:

      This manuscript provides a comprehensive characterization of the Plasmodium falciparum protein LSA3, combining biochemical, genetic, and in vivo approaches. The authors convincingly demonstrate that LSA3 is expressed during liver stage infection and that disruption of the gene leads to a modest but reproducible reduction in liver stage parasite load in humanized mice.

      Strengths:

      Their biochemical and cell biological analysis of blood stages provides strong evidence that LSA3 is exported to the infected erythrocyte, and the detailed analysis of its PEXEL motif processing is well executed.

      We thank the reviewer for their comments.

      Weaknesses:

      The study suggests LSA3 as one of only two known P. falciparum PEXEL proteins contributing to this stage, although there is no evidence for the export beyond the vacuolar membrane. Several key conclusions, particularly regarding antibody specificity, localization in liver stage parasites, and the interpretation of the phenotypic data, are not fully supported by the current experiments.

      We understand the reviewer’s points. We agree that there is no evidence provided that LSA3 is targeted beyond the PVM; whether any of the protein faces the hepatocyte cytosol is unknown (and challenging to conduct) but this possibility remains plausible. LISP2- and LSA3deficient liver stages are less fit than parental controls and thus we stand by the conclusion that they are the two so far identified P. falciparum PEXEL proteins that are important for liver-stage development.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Line 163 says: "Altogether, this demonstrates that LSA3 is important but not critical for blood stage growth of P. falciparum": this is based on the cited Morita et al., 2017. However, previously LSA3 was considered dispensable based on a knock out in 3D7 (Maier et al., 2008; PMID: 18614010). Given that the authors generated a mutant for this work, it would be straightforward to test growth and clarify the importance of LSA3 in blood stages. If important, the analysis of the location and transport of LSA3 in blood stages would immediately become more relevant.  Maybe the data for this is already in the paper: the number of stage V gams was similar between mutant and control (Figure 4A). If this was calculated from the total number of asexual starting parasitemia, it includes blood stage growth, and it can be assumed that there is no growth defect in the mutant in the blood stages. If the number of stage 5 gams was calculated from the number of committed schizonts/rings, nothing can be said about blood stage growth, and asexual blood stage growth should be tested in specific experiments.

      We thank the reviewer for raising the function of LSA3 in blood stages and agree it was an obvious omission, though for good reason - a separate, collaborative study was underway. While this eLife preprint was in revision, our accompanying manuscript on the blood stage was published, showing the characterization of our NF54 DLSA3 mutant during blood-stage growth (PMID:41135800). The findings are now summarized and the citation included in the revised version of this preprint.

      Manuscript line 105: "although, notably, functional characterization of lsa3 deletion mutants has not yet been reported to confirm an important function": at least in blood stages, it was reported to be dispensable, see above. The corresponding study (Maier et al., 2008, PMID: 18614010) could be cited in that context. 

      The citation of PMID18614010 and 39913589 have now been added and we thank the reviewer.

      (2) Some questions central to the conclusions of this paper remain because it was unclear whether the serum did indeed detect LSA3 in the liver or not. It would be easy to check if all cells from the WT/Mutant mix experiment show LSA3 signal (this would mean it cross-reacts) or if only about half are positive (the mutants would be negative if there is no cross-reaction). This would be important to mention for Figure 5 because, at present, it is not known that what is labeled by the LSA3-C antibody in these images is (only) LSA3. 

      We thank the reviewer for this point and completely understand. We did check this via microscopy of liver sections co-infected with LSA3 mutant and control liver-stage parasites as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5-day post-infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (3) It is also unclear which parasites were imaged in Figure 5. The text of the results states that NF54 liver stages were used, but later: "As we employed a co-infection strategy to assess the essentiality of LSA3 versus NF54 in mice, we could not perform IFAs on individually infected mice in this study to validate the specificity of LSA3-C at the liver-stage". The legend says NF54 sporozoites on day 5 post-infection were used. I suspect it was a WT/mutant mix, in which case the above applies, and in the absence of cross-reactivity, half of the cells should be LSA3-C negative. If this is not the case, the localization in the liver becomes dubious.

      We apologize for the confusion and have corrected this. In Figure 5, we utilized liver sections from NF54-infected humanized mice that were stored at -80 C from a previously published study (McConville et al, PNAS 2024). Ideally, we would validate the specificity of LSA3 antibodies at the liver-stage using liver sections containing only DLSA3 parasites however the number of mice available was limited and the samples available to us also contained the Control line for qRT-PCR analyses (the co-infection strategy). As mentioned above, we couldn’t distinguish between these two strains by IFA at the time point analysed and this precluded us unequivocally validating the LSA3-C specificity in the liver-stage; however it cannot be excluded that the signal observed at the PVM is indeed LSA3. We are currently focusing research efforts on obtaining more humanised mice to answer this.

      Minor:

      (1) Introduction: Before the part on the PEXEL motifs, there are almost no references; please add references for all statements.

      We have added references.

      (2) Figure 1B is unclear regarding which part of the gene was deleted. The system used would permit a complete gene deletion, but the homology flanks seem to be within LSA3. If parts of the gene are left, the 75 kDa on the western blots might be a degradation product arising from both the truncated and the full-length protein. Please clarify in the sketch exactly where the homology flanks are, with respect to the start and stop of the gene. 

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (3) Line 161: Replace was with were.

      Corrected.

      (4) Figure 2, 224: Why do the authors think LSA3 must be in the luminal leaflet of the PVM as opposed to the outer leaflet of the plasma membrane?

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte.

      (5) Line 245: GFP core "derived from digestion of the reporter in the food vacuole, which confirmed it was secreted from the parasite". I wonder if the amount of GFP "core" really can be used as evidence for secretion, and its amount can be compared between experiments. Did the author quantify this for the full-length protein to get a proportion per sample?

      Use of GFP core to measure defects in P. falciparum GFP reporter secretion has been described previously (for example PMID:23387285 and 35906227). The comparison the reviewer asked for is an interesting and important question: however the control would be to compare the ratio of GFP core to uncleaved in the control lanes as well, which is not possible to do since the full-length protein is digested by plasmepsin V in the native PEXEL versions of the experiments (mLSA3-GFP in the first blot, Vehicle in the second blot) leaving no full-length protein to compare to. It stands to reason that inhibition of N-terminal processing results in less protein removal from the membrane (ER or COPII vesicle or PM) resulting in less secretion out of the parasite for retrograde transport to the food vacuole with cytostomal vacuoles (analogous to plasmepsin II). In the food vacuole, the chimeras are in normal cases digested by proteases back to the GFP core that is resistant to cleavage and evident as GFP core on the immunoblots (PMID:10775264 and 14709539 and 19055692 and 20130643). 

      (6) Figure 3 has the word plasmid in two lanes. In Figure 3E, amend the labelling of the blots.

      We apologize for the formatting error in converting the figures to PDF during the original submission and thank the reviewer for the suggestion. This has now been corrected.

      (7) Lines 266/271/284: "live IFAs", live immunofluorescence assay. Does this mean an antibody was given to living   parasites?

      The correct term is live microscopy and this has been corrected.

      (8) Does Figure 6A fit with the data in Figure 6B? It seems 6B has a milder phenotype than 6A.

      We thank the reviewer for the question. Yes the data directly correspond to each other and are represented in two ways: Panel A shows the qRT-PCR raw data for liver load of each parasite strain per humanized mouse using a scientific scale on the y-axis. Panel B shows that magnitude of the DLSA3 defect as a percentage of the total liver load per mouse:

      % total parasite liver load  = ( strain 1 or strain 2 liver load ) x100

      sum of strain 1 + strain 2 liver loads

      The intent of showing both data is to convey the correct magnitude of the difference in two ways to assist the reader in understanding the true defect, both are accurate and both are statistically significant. In revision we detected mislabelling of humanized mouse 2 and 3 in the original graphs that has now been corrected and we sincerely thank the reviewer for helping us identify this error.

      (9) Line 482: Please add references for this debate. 

      These have been added.

      Reviewer #2 (Recommendations for the authors):

      Major Comments: 

      (1) In general, the authors have taken care not to overstate conclusions from their study. Nonetheless, while not technically inaccurate, the title might misleadingly suggest LSA3 is exported in the liver stage (this was my initial impression on reading it until I looked at the data). I suggest the authors revise the title to avoid confusion by clarifying that export was only observed in the blood stage.

      We sincerely appreciate the reviewer’s point. As this article was posted as a preprint that has now been cited several times, we have carefully weighed the comment and in the end decided to retain the current title for the above reason.

      (2) While the ability to generate the ∆LSA3 parasites clearly shows that the protein is not essential in the blood stage, the impact on parasite fitness is never tested but simply assumed (for instance, in lines 163-164: "...this demonstrates that LSA3 is important...for blood-stage growth..."). Do the ∆LSA3 parasites have a fitness defect in the blood stage consistent with the previous GIA data that would support this claim? Since the rabbit anti-LSA3-C antibodies produced by Morita et al did not have GIA activity against the blood stage, it is possible that the GIA observed with the human and mouse antibodies might have been due to reactivity with a different protein. If ∆LSA3 does cause a fitness defect, it would be interesting to know if the endogenous GFP-tagged line, which alters protein trafficking/membrane association, also produces this effect.

      We agree with the reviewer and would like to clarify that this omission was not intended to create confusion but was by design, due to a separate collaborative study that was underway to address such questions. While this eLife preprint was in revision, our accompanying manuscript on characterising NF54 DLSA3 at the blood stage was published (PMID:41135800). The findings are now summarized and the citation included in the revised version of this eLife preprint. In sum, LSA3 is not critical for erythrocyte invasion but its deletion perturbs the rate and efficiency of merozoite invasion, at the step(s) of resealing of the PVM/host cell, resulting in aberrant accole forms that protrude from the infected erythrocyte.

      (2) Figure 1D: While the images are compelling and I don't doubt the claim that LSA3 is exported in the blood stage (also supported by the fractionation/Pk experiments), the authors should provide quantification of the difference in exported signal between the WT and ∆LSA3 parasites in these IFAs to rigorously support this conclusion. Also, please include details about how many independent experiments are represented by the microscopy data throughout the manuscript (Figures 1, 2, 3, and 5).

      We understand the reviewer’s request and wish to indicate that the export signal was absent in all cells infected with DLSA3 that was imaged. The microscopy performed was from n=2-3 experiments except for Figure 5 which was from n=1 humanized mouse per time point in which multiple EEFs from the liver were imaged. This has been indicated in the figure legends. 

      (3) Careful inspection of the z-series images in Figure 5A shows that most of the LSA3-C signal seen outside the PVM (beyond the boundary delineated by EXP1) is closely associated with DAPI puncta, suggesting these are merozoites. Together with the prominent gap in the EXP1 signal, this suggests the schizont has already ruptured. Thus, anti-LSA3-C signal beyond the PV seems best explained as coming from merozoites or other material released by PV rupture, not from export across the PVM, and this should be added to the text in place of comments about localization to PV extensions or potential export (lines 358-359, 422-423).

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have added the comment as requested.

      Minor Comments:

      (1) The authors may want to denote the disordered repeat region in the LSA3 schematic in Figure 1A that is mentioned in the text.

      We have added the residue boundaries of the predicted domain from AlphaFold into both the schematic and the text and included a link to the LSA3 pages in PlasmoDB and

      AlphaFold in the Methods section.

      (2) The authors use rabbit anti-LSA3-C antibodies previously generated by Morita et al. These polyclonal antibodies were raised against a recombinant C-terminal region of LSA3 (residues 750-1433), but the schematic in Figure 1A indicates the antibodies recognize a smaller region between residues 1154-1433. Please adjust the figure accordingly, or if this is not the same antiLSA3-C antibody reported by Morita, please provide details about its production.

      The figure is corrected.

      (3) The authors use Alphafold to identify a region of LSA3 with similarity to the substrate binding domain of DnaK, but the data is not shown. Please include the Alphafold prediction in supplementary figures and provide information about how the predicted structural homology was determined.

      We have added a link to the AlphaFold page for PF3D7_0220000 in the methods.

      (4) The schematic in Figure 1B indicates that the DHFR cassette was inserted at an internal site within the lsa3 gene. If this is the case, it seems possible that an N-terminal portion of the protein is still expressed, but I was unable to find details about the boundaries of the homology flanks to determine the precise insertion site. Please clarify the knockout strategy and indicate the specific insertion site.

      The LSA3 gene was disrupted using the flanks as shown. The DHFR selection cassette comprises its own promoter and terminator such that insertion into the coding sequence completely disrupts expression of the protein thereafter, including the C-terminus within which the LSA3-C antibody binds. The new LSA3-T antibody described in our recently published accompanying manuscript that binds more N-terminally than LSA3-C also does not label the truncated protein. The original 5’ and 3’ flanks used for integration of the disrupted LSA3 allele by double cross-over recombination were then looped out into the original knockout plasmid and this was negatively selected against using exogenous 5-fluorocytidine (5-FC) via the suicide gene cassette CDUP (cytosine deaminase and uracil phosphoribosyl transferase that also contains a 5’ promoter and 3’UTR terminating element) in the construct. These features should provide clarification and have now been indicated in the figure and legend.

      (5) Line 162: I think this should read "antibodies that react with LSA3 were...".

      Corrected.

      (6) Figure 1D: The merge with the transmitted light channel is missing for the third panel in the ∆LSA3 IFAs. Also, please define the scale bar length in the legend.

      Corrected.

      (7) Lines 744-746: The IFA fixation panel order description (top, bottom) in the Figure 2A legend is reversed from what is shown in the actual figure. Also, please define the scale bar length. 

      Corrected.

      (8) Lines 184-186: Since the fractionation/PK protection assays suggest most of LSA3 is in the PV, it would be interesting to know if the strong peripheral/PV signal observed in the PFA-fixed IFAs in Figure 2A is also present in the ∆LSA3 parasites, or is this non-specific? 

      Thank you for the suggestion. We agree this would be an interesting result to know but do not have the capacity at the present time.

      (9) Lines 219-225: It is unclear to me why these results are interpreted to suggest that the majority of LSA3 is peripherally associated with the luminal leaflet of the PVM. Wouldn't an integral membrane configuration in the PVM (with the C-terminus facing the host cytosol) or PPM (with the C-terminus facing the parasite cytosol) also account for the data? Adding a carbonate extraction would help clarify this point.

      Several pieces of evidence combined led us to this conclusion in Figure 2B. i) if LSA3 was on the outer PVM leaflet, it would be substantially degraded in the EQT Pellet + PK fraction but a substantial population remained insensitive to PK, indicating much of the total protein pool was protected by the PVM (and possibly the parasite membrane; PM), ii) yet saponin, which leaves the PM intact, allowed PK to access and almost completely degrade LSA3 (see Saponin Pellet + PK), indicating that a substantial population of LSA3-C is located inside the boundary of the PVM, and this is membrane associated as saponin did not liberate it, rather, it remained in the Saponin Pellet before PK was added, iii) the TX-100 Super fraction confirmed LSA3 is membrane associated, as more is present in the TX-100 Super than the Saponin Super fractions, iv) if LSA3 was inside the PM, the Saponin Pellet fraction should be resistant to PK (as was the case for the cross-reactive band indicated with a red asterisk) but LSA3 (green asterisk) in the Saponin Pellet was PK sensitive. Altogether, our best conclusion from these data is that LSA3 is likely to be PVM-associated with the LSA-C-binding domain facing internal to the PV, and a fraction is also exported beyond the PVM into the erythrocyte. If the question is whether LSA3 is an integral PVM protein, we agree that use of carbonate in the future would answer that question.

      (10) Figures 3D and E: There are some problems with some of the text wrapping in these panels.

      We apologise, this was a formatting issue as the manuscript was converted to PDF.

      We have corrected this error.

      (11) Line 422-423: In fact, the Z-sections shown in Figure 5 appear to indicate that the LSA3-C signal is predominantly located within the parasite, not at the PVM.

      We do appreciate the reviewer’s careful eye and caution and are in complete agreement. We have corrected the final conclusion to be more accommodating of this.

      (12) Lines 468-470: Since cross reactivity of anti-LSA3-C is substantial in the blood stage but was not defined in the liver stage by analysis of unmixed infections, how do the authors know that they were not observing ∆LSA3 parasites in their IFAs? I think what they mean here is that parasites lacking anti-LSA3-C reactivity were not observed, which is an important distinction.

      The reviewer is correct and this has been corrected.

      (13) Lines 478-479: The authors should also mention that the P. berghei PEXEL proteins evaluated in Fougere et al are exported in the blood stage, similar to LSA3. Moreover, other studies have shown something similar for additional endogenous PEXEL proteins or reporters in P. berghei (PMIDs 22329949, 26347246, 34956312).

      We have added the additional text regarding export into the infected erythrocyte and the reference to IBIS1.

      (14) Line 491: The data here don't support that LSA3 is "required" for liver stage development, only that it is important to it. Since the authors have not defined the cross-reactivity of anti-LSA3C in unmixed infections, it is not clear that ∆LSA3 parasites are arrested early in the liver stage, only that they show a reduced number of genome copies relative to the parental control. 

      We have amended the sentence to “required for normal liver stage development”.

      (15) Line 530: I think NGF54 should be NF54.

      Corrected.

      Reviewer #3 (Recommendations for the authors):

      (1) Antibody specificity in liver stage IFA experiments:

      The specificity of the anti-LSA3 antiserum (LSA3-C) used in liver stage IFA is not fully convincing. While the KO parasites were used effectively to validate specificity in blood stages, the same is not true for liver stages. 

      (a) It is essential to repeat IFA with ΔLSA3 parasites in liver stage infections to determine whether the observed PVM staining is truly specific.

      We appreciate the reviewer’s point, however at a cost of over $5000 per humanized mouse, we do not have the capacity to conduct this experiment at the present time. We highlight that, as the blood stage IFAs confirmed the specificity of LSA3-C for LSA3, the possibility remains open that LSA3 is specifically recognized at the PVM.

      (b) If the antibody is the same polyclonal serum used in Morita et al. (2017), why did the authors not employ a monoclonal antibody, which they presumably have access to and which would provide greater specificity? 

      We have included new data confirming that LSA3 is exported using LSA3-T, in addition to LSA3-C.

      (c) Given that rabbit antisera often show non-specific staining at the PVM in liver stage parasites, co-localization with PVM markers is not sufficient. Inclusion of the ΔLSA3 parasites in liver stage IFA is critical. It will also show whether there is any cross-reaction of the antiserum in liver stage parasites, as seen by IFA for blood stage parasites. 

      We thank the reviewer for their feedback.

      (d) To validate the serum further, the authors should infect HC-04 cells in vitro with GFP-LSA3 parasites and stain with LSA3-C to confirm overlap between the tagged protein and the antibody signal.

      We thank the reviewer for their feedback.

      (e) For higher-resolution co-localization, expansion microscopy - now commonly used even in malaria research - would substantially improve the analysis. 

      We thank the reviewer for their feedback.

      (2) The localization of LSA3 in this study differs notably from Morita et al. 2017, who reported localization to dense granules in merozoites and staining in ring-stage parasites at the PVM. 

      (a) The authors confirm DG localization, but they do not examine ring-stage parasites. They should include the IFA of ring stages to clarify whether they can replicate the previous findings.

      We thank the reviewer for their feedback.

      (b) Additionally, the differences in Western blot banding patterns between the two studies should be addressed. Do the authors have an explanation for these discrepancies? 

      We thank the reviewer for their feedback.

      (3) The authors report a ~40% reduction in liver parasite load using qPCR, which is statistically significant. However, this phenotype is modest and should not be interpreted as showing that LSA3 is essential.

      (a) Please avoid terms like "required" or "essential" and instead describe the protein as "contributing to normal development" or "influencing fitness."

      We have used the term “required for normal liver stage development”.

      (b) Since the authors generated liver sections, they should take advantage of these to quantify the number and size of liver stage parasites, which would help determine whether the phenotype reflects fewer infected cells or reduced parasite growth.

      We did check this via microscopy of liver sections, but all mice were co-infected with LSA3 mutant and control liver-stage parasites, as we shared the reviewers line of enquiry. Unfortunately we could not detect parasites without LSA3 signal at the 5 day post infection time point. This type of analysis does sound straightforward on paper but in reality is more challenging owing to several factors i) identifying sufficient individual parasites in an entire liver by microscopy can be challenging and variable from lobe to lobe and mouse to mouse, ii) the number of parasites required for a meaningful statistical analysis is increased due to coinfection of the liver (see Figure 4B as an illustration of this), iii) day 5 is a rather late liver-stage time point and so if there was a growth defect the defective parasites may be very small or sparse, iv) we cannot exclude that the LSA3 antibody may cross-react at the liver-stage, v) definitive conclusions are thus challenging and we feel require individual co-infections to be clear in the future. Nonetheless, the detailed qRT-PCR analyses identify a significant reduction in DLSA3 parasite liver load on day 5, indicating this protein is important for the human malaria parasite’s growth within human hepatocytes.

      (c) It would also be valuable to include IFA from singly infected ΔLSA3 livers (rather than co-infected), and possibly at earlier timepoints, to identify the developmental window affected.

      We agree it would be valuable.

      (4) The manuscript suggests that LSA3 may be exported beyond the PVM into the hepatocyte, based on a small number of peripheral puncta.

      (a) This claim is not convincingly supported by the data. The punctate signals shown in Figure 5 are weak and may rather reflect PVM extensions or TVN. In fact, one punctum even overlaps with the DAPI signal (figure 5, middle panel), which raises further doubt about the localization.

      We appreciate the reviewer’s careful eye and caution and have added the comment regarding DAPI.

      (b) Given the lack of KO controls in these liver stage IFAs, the authors should not describe LSA3 as "exported beyond the PVM". The language should be revised to reflect that the protein localizes predominantly to the PVM, and any extra-PVM signal remains unconfirmed and could be non-specific. 

      (c) This is especially important given the well-known tendency of rabbit antisera to produce background PVM staining in liver stage parasites. 

      Corrected.

      (e) In an earlier report (McConville et al, 2024, PNAS), they clearly state that LSA3 is NOT exported beyond the PVM. Actually, the staining in the previous report looks quite different from the images provided for Figure 5. The authors might wish to comment on this. 

      We thank the reviewer for their feedback.

      Minor comments:

      In some sections, the manuscript uses "exported" to refer to trafficking to the PVM. This terminology should be used more carefully and consistently, since "export" often implies translocation into the host cytosol

      We understand that export involves a protein localizing within the host cell and so protrusion through the PVM may also be considered exported, however, we have not confirmed this for LSA3 in liver stages.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We greatly appreciate the reviewers for their efforts in reviewing our manuscript. We highlight that the key contributions of our paper are to provide a framework for calibration and validation of high-fidelity cardiac electromechanical models based on a diverse compilation of clinical datasets, and that we provide one example of such an evaluation of our own baseline electromechanical model. The comments raised by the reviewers were chiefly focused on the second goal, which is our specific model and the outcomes of the evaluation process, rather than on the evaluation framework itself. As such, we have made improvements to our model implementation and to provide additional confidence in our specific modelling framework through this review process. Specifically, we have strengthened the verification component of this evaluation, provided additional quantitative measures, and included a more in-depth discussion of the remaining limitations in our modelling framework. We hope that the updated version of the manuscript and our efforts to improve it are well-received by our reviewers and editors, as well as by the modelling and simulation community at large.

      eLife Assessment

      This is a potentially important study that explores the relevant range of parameter values for calibration and validation of cardiac electromechanics in ventricular models. Although much of the work presented is solid, the evidence provided to support the authors' key scientific claims is incomplete, especially as it relates to the emphasis on standardized validation and verification approaches. Notably, the level of model personalization presented in this work falls short of the threshold for what could reasonably be called a "digital twin", even by the relatively relaxed standards that have emerged in computational physiology and related fields in recent years.

      We appreciate the eLife assessment for identifying the potential importance of our study. Regarding the threshold for 'digital twin', we note that a cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning, which is a goal that to our knowledge no published electromechanical study has simultaneously fulfilled. It is for this reason that the community refers to the 'digital twin vision' rather than its realisation, and we adopt this framing consistently throughout the manuscript.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers within a single study. In our revision, we have clarified that the model evaluation presented here is an example application of that framework, which was designed not to certify a model as complete, but to provide a transparent audit of current capability that identifies where confidence is established and where further development is needed. We have updated the title and language throughout the manuscript to reflect this framing consistently.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The study by Wang et al. investigates cardiac electromechanical modeling and simulation techniques, focusing on the calibration and validation of ventricular models according to ASME V&V40 standards. The researchers aim to calibrate model parameters to align with key biomarkers such as QRS duration and left ventricular ejection fraction, and validate the model against independent measurements such as displacement and strain metrics. The authors also examine the impact of parameter variations on deformation, ejection fraction, strains, and other biomarkers. The overarching aim of the study is to give "credibility to the underlying computational electromechanics framework" and to "pave the way towards credible cardiac electromechanical Digital Twins."

      Strengths:

      (1) The study presents a solid validation strategy for cardiac models based on independent data.

      (2) It integrates electrophysiological, mechanical, and hemodynamic biomarkers for sensitivity analysis and calibration.

      Weaknesses and Limitations:

      (1) Model Assumptions: The study employs simplified modeling assumptions that are not state-of-the-art, e.g.,

      (a) Isotropic scaling of the mesh to generate an unloaded reference geometry.

      (b) Simple afterload and preload models that fail to produce physiological results.

      (c) Simplified epicardial boundary conditions.

      While our model was able to broadly achieve physiological behaviour based on the calibration and validation datasets, it also contains several simplifications that can be expanded with more sophisticated techniques to allow explorations in specific areas. We have added a dedicated Limitations subsection to the Discussion section of the manuscript to address these and to provide references to relevant studies.

      (2) Numerical Framework:

      (a) The mesh resolution and/or the numerical framework used for the mechanical part appears to suffer from known numerical artifacts (locking effects), leading to overly stiff or inaccurate behavior in finite element analysis. This results in an artificially stiff response to deformation, which is compensated by setting active contraction to ten times the value reported in the literature. The authors attribute this to limitations in using ex vivo tissue measurements to represent in vivo function, although similar issues were not observed in previous works.

      We thank the reviewer for raising this point and have investigated it carefully. We have added a verification section as well as discussions to the manuscript to more comprehensively discuss this point. In short, through various tests against benchmark (Land) and comparing stress-strain curves in cube simulations with the same mesh resolution, we could not identify evidence of volumetric locking effects. The elevation in contractile force was also necessary in a simplified ellipsoid version of the model in a previous publication [ref 7, Levrero-Florencio, et al. 2020]. We note that these benchmarks were performed in the incompressible transversely isotropic regime; whether analogous locking effects exist in the dynamic orthotropic active contraction framework used in the full biventricular simulations remains an open question, which we have identified as a priority for future benchmarking, for example against the Arostica et al. 2025 benchmark.

      We agree that the explanation of this as ex vivo vs in vivo difference in contractile force is too simple, and other contributing factors are better understood through comparison with similar studies in the field. Strocchi et al. (2023) used a four-chamber model with explicit atrial mechanics, and in her history matching varied Tref within +- 33-55% of a reference value of 120-150 kPa, targeting a peak active tension of 160 +- 15 kPa, which was a considerably more modest adjustment than applied here, likely reflecting differences in model geometry, pericardial constraint, and circulatory model between the two studies. Gerach et al. (2021, Mathematics) applied manual parameter adjustments informed by in vivo active tension measurements of 120 – 150 kPa and achieved ejection fractions of approximately 63%; however, they reported that systolic pressures in both ventricles were too high for a healthy heart, and similarly reported elevated peak ejection rates compared to MRI measurements, a difficulty we also encountered. Notably, Gerach et al., report that atrial contraction contributes approximately 11-13% of end-diastolic volume, which in a biventricular-only model would directly reduce the achievable LVEF and necessitate compensating adjustments to active tension. Zingaro et al. (2024, Journal of Computational Physics), using an alternative active tension model (RDQ20), similarly found it necessary to increase contractility parameter (a_XB) to achieve sufficient ejection, and explicitly report that no single parameter configuration simultaneously achieved physiological peak ejection rate and LVEF, a fundamental tension we also encountered. Together, these comparisons suggest that the elevated Tref in our model most likely reflects a combination of the absence of atrial filling, simplifications in pericardial constraint, and the lack of poroelastic behaviour, rather than volumetric locking alone. Additional investigations are needed as explained in the manuscript, and these factors are identified as open priorities for future development within our framework.

      We added a comparison with these three studies to the discussion section of the manuscript and have updated the limitations section to reflect this more nuanced account of the factors contributing to the elevated active tension scaling. Further work will be required to address these points.

      (b) Further, the authors employ the monodomain model for the simulation of the electrical excitation and relaxation on a relatively coarse grid with an approximate edge length of 1mm. This resolution is known to be insufficient for reliable results in organ-scale electrophysiology modeling.

      Our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached, as performed in Camps et al. (2024) (ref 33) using the tuneCV tool in monoAlg3D (https://github.com/rsachetto/MonoAlg3D_C/tree/master/scripts/tuneCV), which is similar to the tool in openCARP (described here: https://opencarp.org/documentation/examples/02_ep_tissue/03a_study_prep_tunecv).

      While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for simulations of ECGs in this study. We have noted this in the methods section.

      (3) Geometrical model and digital twin: The geometrical model, taken from a public cohort and calibrated to an ECG of another individual along with population-averaged values from a databank (UK Biobank), and unrelated measurements from surgical procedures, can hardly be considered a digital twin. Further, validation of the model was then performed against data from yet another cohort.

      We thank the reviewer for this point and welcome the opportunity to clarify our dataset choices. The use of multiple data sources was a deliberate methodological decision. While an ideal dataset for electromechanical model evaluation would combine full biventricular geometry, 12-lead ECG, invasive pressure measurements, and myocardial strain data from a single individual, no such dataset currently exists in the public domain, and acquiring it routinely would be impractical in clinical settings. Multi-source integration therefore reflects the realistic deployment scenario for future clinical translation of these tools.

      The specific choice of geometry was principled: the mesh associated with the ECG dataset that was available to us was truncated at the base due to the clinical acquisition protocol, which would have prevented physiologically realistic basal boundary conditions. The female Rodero geometry we chose provides full ventricular coverage and was selected on that basis.

      Demonstrating that a coherent, systematically evaluated framework can be constructed from compiled multi-modal data is itself a contribution because it makes the tools accessible to the wider community without requiring a single ideally acquired dataset.

      (4) Calibration procedure: There are apparent flaws in the calibration procedure, or it is not described in sufficient detail. The authors dedicate significant effort to motivating parameter ranges, but in the end they use mostly other parameters for the calibration process, aiming to maximize left ventricular ejection fraction. It is not clear whether the chosen parameters result in, e.g., physiological calcium traces or calibrated parameters that are within physiological ranges.

      Thank you for raising this point, which we have now clarified in the manuscript. The parameters that were chosen for the calibration process were based on the results of the sensitivity analyses.

      In addition, we have supplemented results Figure 1 with a subfigure F showing that the calcium transient and action potential durations fall within physiological ranges after calibration.

      (5) Goodness of fits, e.g., a direct comparison of the measured and the simulated ECG, are not provided to assess calibration quality.

      The calibrated model achieves QRS duration of 89 ms and QT interval of 360 ms, both of which fall within the healthy reference ranges compiled in Table 2, providing a biomarker-level assessment of calibration quality (Figure 1A). A full quantitative goodness of fit analysis of the simulated ECG morphology was performed following the methodology of Camps et al. [52], in which the same beat-averaged ECG was processed; we direct the reader to that work for full details rather than reproducing the analysis here.

      (6) Due to these limitations and weaknesses, the authors fall short of achieving some of their goals, particularly establishing credibility for the underlying computational framework and in reproducing healthy pressure-volume loops, and in achieving physiological simulations while using physiological or reported ranges for the calibrated parameters.

      For example, a key physiological requirement is that the right and left ventricular stroke volumes are approximately equal in a heart beating at a limit cycle, as the blood pumped by the right ventricle into the pulmonary circulation must match the amount pumped by the left ventricle into the systemic circulation. This balance is not achieved in this study.

      We thank the reviewer for identifying the stroke volume imbalance. We acknowledge the physiological requirement that, in a steady-state limit cycle, the right ventricular stroke volume must approximately equal the left ventricular stroke volume. However, since our model does not explicitly prescribe volumes, to achieve this, we would need to either explicitly tune active tension for the left and right ventricles separately, such as done in https://www.frontiersin.org/journals/physiology/articles/10.3389/fphys.2021.716597/full or develop a more sophisticated circulatory model and employ a multistep procedure that sequentially tunes circulatory dynamics, passive mechanics, and active contraction, such as done in https://www.biorxiv.org/content/10.64898/2025.12.11.693778v1.full. Both of which are beyond the scope of this paper.

      We note that, despite the absence of explicit RV calibration, the RV volumetric measures and pressures remain within physiological ranges, suggesting that the coupled biventricular mechanics are broadly plausible. As such, we have noted this limitation in our discussion section, and sign-posted to other studies where the stroke volume match is achieved.

      (7) The conclusive claim that "the study paves the way towards credible electromechanical cardiac Digital Twins" is not supported. The model exhibits non-physiological behavior, requires unsupported parameter alterations (such as a 10-fold active stress scaling), and does not represent a digital twin, as model data are drawn from various unrelated, non-patient-specific sources.

      We thank the reviewer for this comment, which gives us the opportunity to clarify our use of the term 'digital twin'. A cardiac digital twin is envisioned as a patient-specific computational model of the heart, personalised from multi-modal clinical data and continuously updated to support diagnosis, prognosis, and treatment planning. This is a transformative goal for precision cardiology that the field is actively working towards, with credible, systematically validated electromechanical models as its essential foundation. To our knowledge, no published study in cardiac electromechanical modelling has simultaneously fulfilled all three requirements, and it is for this reason that the community often refers to the 'digital twin vision' rather than its realisation.

      The primary contribution of this manuscript is the framework: a systematic application of ASME V&V40 standards to a fully coupled electromechanical model, spanning electrical, mechanical, and haemodynamic biomarkers in a single study. The model evaluation presented here is an example application of that framework. Importantly, the framework is not designed to certify a model as complete, but to provide a transparent audit of current capability by identifying where confidence is established and where further development is needed. In this sense, the limitations surfaced through this evaluation are themselves a contribution: they define open problems and priorities for the field.

      We have added a definition of the digital twin concept and the roadmap towards its realisation to the introduction and have updated the language throughout the manuscript to consistently reflect the distinction between the framework contribution and the model evaluation. We maintain that this transparent approach represents a meaningful step towards the digital twin vision.

      The specific limitations of the current model implementation are addressed in detail in the relevant sections of this response and in the updated manuscript, where we have substantially strengthened the verification and discussion components.

      Conclusion:

      Overall, this reviewer considers that the study requires a major revision, including improvements in numerical methods, modeling choices, and checks for physiological behavior. Nevertheless, the provided tables with averaged values from the UK Biobank and the presented validation strategy could be valuable to the research community.

      Reviewer #2 (Public review):

      The authors present an interesting study on calibrating and validating a biventricular cardiac electromechanical model. This is an important contribution, but some questions remain about the quantitative validation and verification aspects of the study.

      Major comments:

      (1) The title and paper stress the importance of validation on several occasions. However, the actual validation performed is limited to the section in lines 427-439. Furthermore, it is entirely qualitative, making assessing the model's quality difficult. Most of the paper is focused on sensitivity analysis, which is also interesting but unrelated to validation. Can you include a quantitative comparison with deformation biomarkers? E.g., spatially quantify strain differences between simulation and in vivo data, or overlay the current configuration of the geometry with MRI in various views, and calculate a displacement error norm.

      We thank the reviewer for this comment.

      We have strengthened the quantitative aspect of the validation by reporting the peak simulated strain values for each component and comparing them against the physiological ranges compiled in Table 2. Specifically, the simulated peak strains were: E_ff ≈ -0.20, E_cc ≈ -0.15, E_rr ≈ +0.15, and E_ll ≈ -0.23. These show broad agreement with the in vivo reference ranges from Moulin et al. (2021), noting that the reference ranges are derived from a cohort of 30 subjects and therefore represent a relatively narrow population sample. Shortening strains (fibre and circumferential) are in good agreement, while radial strain is underestimated. We have noted this as a limitation. We have also indicated that a further validation would include a fully quantitative spatial comparison, such as a displacement error norm or voxel-wise strain difference map. This would require access to the raw image data and patient-specific geometry registration, which is beyond the scope of the current study.

      (2) You mention the ASME V&V40 standards throughout your paper. Yet, you only address the "second V" validation, ignoring the "first V" verification. How did you ensure that your computational models are implemented correctly?

      Thank you for raising this point. We have now included a section on model verification to the manuscript at where we perform benchmarking simulation using the Land (2015) passive inflation benchmark. We also provide a mesh subdivision analysis of the final calibrated model. Additional verifications and previous sensitivity analyses using the same numerical scheme with idealised ellipsoid geometries are also referenced in the verification section, to provide additionally confidence.

      (3) All parameters discussed in this publication are physical parameters. What is the sensitivity of your model outputs concerning computational parameters?

      Numerical analyses for the Alya solver used in this study has previously been published in works including Levrero et al (2021), which performed sensitivity analyses in a truncated ellipsoid geometry, and Santiago et al (2018), which demonstrated mesh convergence in a cantilever. We have updated the manuscript to point the reader to these studies.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major concerns:

      (1) Active stress scaling:

      The initial value for T_ref appears to be 120kPa * 10, which would be ten times the literature value fitted to human contraction data. Additionally, Table 2 lists a range of [1200-2400], which is 10 to 20 times the literature value.

      This discrepancy suggests that other model parameters, model assumptions, or the numerical scheme may be inadequate. In contrast, similar calibrations using comparable models (ToRORd-Land) in other works, such as Strocchi et al. [29], yielded T_ref values close to the literature value.

      We thank the reviewer for this comment. As discussed in our response to the public review comment 3a, the elevated T_ref scaling warrants explanation.

      We note that Strocchi et al. use a four-chamber geometry include atrial mechanics and a different pericardial constraint, any of which could contribute to differences in the required T_ref scaling. The elevated scaling in our model likely reflects a combination of factors including the absence of poro-elastic behaviour, simplifications in pericardial constraint, and the lack of atrial mechanics, rather than volumetric locking alone. We have added text to the discussion acknowledging this more explicitly and have flagged planned additional benchmarking of the dynamic orthotropic scheme as future work.

      (2) Non-physiological results, see Figure 1:

      In a healthy heart, RV stroke volume should approx. match LV stroke volume. This is clearly not the case in Figure 1B, where the RV EF is also notably low at 35%.

      Consequently, the study fails to reproduce healthy pressure-volume loops, undermining its claim to create a credible cardiac electromechanical digital twin. Hence, also the "Question of interest" posed in line 206 must be answered with a clear "No".

      Matching stroke volumes should be a primary calibration goal.

      We thank the reviewer for this comment. We agree that stroke volume balance is an important physiological criterion, and we have added it explicitly to the framework criteria in the updated manuscript, noting that our current model evaluation does not satisfy it. This is precisely the kind of transparent appraisal the V&V40 framework is designed to produce: a systematic accounting of which criteria are met and which require further development. A framework that only gets applied to models that pass all criteria would be selection-biased and less informative to the community.

      However, we respectfully disagree that the question of interest must be answered with a clear 'No'. We draw the reviewer's attention to the quantities of interest defined in the paper, which are predominantly left ventricular biomarkers, reflecting the intended scope of the calibration framework. The framework successfully reproduces these defined quantities of interest, and the LV pressure-volume loops, strain, volumes and ejection fraction are all within physiological ranges and well-matched to reference data. These quantities of interest were selected based on their clinical implications in cardiac diseases, as detailed in Table 3.

      We agree that stroke volume balance is an important physiological requirement for a fully calibrated biventricular model, and we have strengthened the future work and limitations section accordingly.

      (3) Inadequate numerical framework:

      (a) Monodomain model: The geometries from Rodero et al. [25] have an average edge length of 1mm. It is known that such a coarse resolution leads to inaccurate EP results. It is not mentioned if the authors refined that geometry to an appropriate resolution or used an Eikonal model to mitigate this issue.

      As explained earlier, our ECG simulations are robust against coarse mesh resolutions since we use an Eikonal solution to prescribe the activation times on the endocardial surface, as performed in Camps et al. (2024), and we tune the diffusivity parameters in the model such that the correct conduction velocities are reached. We have added a figure in the appendix of this manuscript to show that by increasing the mesh resolution by one subdivision, we get virtually identical ECG simulations. While this mesh resolution may not be sufficient for simulations of more complex behaviour, such as re-entry and fibrillation patterns, it is sufficient for this study. We have noted this in the methods section.

      (b) Material law:

      - recent publications show that an unsplit deformation gradient for the anisotropic contribution is beneficial to reduce locking effects, see, e.g., Gueltekin et al. Computational Mechanics 63, no. 3 (2019): 443-53. https://doi.org/10.1007/s00466-018-1602-9.

      - K_ct is a penalty parameter to enforce some degree of incompressibility. Results are highly dependent on the grid size and the finite element formulation due to locking effects.

      As the authors write: "In our simulations, we saw that the LVEF was strongly sensitive to changes in the incompressibility of the tissue (Kct), such that an increase in compressibility of the myocardial tissue helped to increase LVEF." Which exactly points to the issue of locking effects.

      So an option would be to use a finer grid or a more adequate numerical scheme with quadratic finite elements, as eg. in [5] Fedele et al., or [6] Gerach et al,. or stabilized elements as in Karabelas et al. CMAME 394 (2022) https://doi.org/10.1016/j.cma.2022.114887.

      Overall, this does not point to "limitations in using ex vivo tissue measurements to represent in vivo function" but to limitations in the numerical setup. In fact, with an adequate numerical scheme, the simulations should be largely insensitive to the choice of this penalty parameter K_ct. See, e.g., Karabelas et al. above, where the authors varied K_ct from 650kPa to infinity (representing an incompressible material), and there is no visible influence on the PV loops.

      We investigated this point using the Alya solver, and we found that the mesh resolution did not alter the LVEF, and our benchmark simulations against Land (2015) did not show the existence of the volumetric locking issue that the reviewer refers to. It is possible, however, that such an effect exists in the elastodynamic orthotropic framework but not in the incompressible and transversely isotropic framework that the Land (2015) benchmarks were set up in. Future analyses could focus on performing additional benchmarking against more recent elastodynamic benchmarks, such as presented in Arostica (2025). We have updated the limitations text in our manuscript to reflect this and to cite relevant literature on this issue.

      (4) Boundary conditions:

      "This was a simplified version of the method [28], which uses an exponential decay formulation at the 'edge' of the pericardial constraint rather than a step function": I don't really see this in the cited work [28] which gives a spatially varying Robin-type boundary condition at the whole epicardium (i.e. regional scaling of normal springs stiffness based on image-derived motion from CT images) and not only at the edge.

      This is motivated by the fact that the pericardial tissue is in contact with various organs of different material properties. Not using spatially varying pericardial parameters is a limitation that might lead to non-physiological deformations, see also Pfaller et al. Biomechanics and Modeling in Mechanobiology 18 (2019): 503-29. https://doi.org/10.1007/s10237-018-1098-4.

      We thank the reviewer for this point and we have corrected the manuscript accordingly. To clarify: our implementation applies a uniform Robin spring constraint along the majority of the epicardial surface with zero constraint at the base, which is conceptually similar to Strocchi et al. [28]. The key difference is that Strocchi et al. use a smooth gradient transition from uniform constraint to zero constraint near the base, whereas our implementation uses an abrupt step transition. We acknowledge that a smooth spatially varying transition would more accurately represent the frictionless pericardial contact and have noted this as a limitation in the manuscript with reference to Pfaller et al. [41].

      Also check:

      - line 117: Gamma_valve_epi is introduced but not used. Was there any boundary condition defined on this valve plane?

      - the third equation, maybe (0,T] missing.

      - line 120: epicardium instead of endocardium.

      These errors have been corrected in the updated manuscript. No boundary conditions were applied on the epicardial surface of the valve plugs, the reference to gamma_valve_epi has been removed.

      (5) Reference geometry:

      The choice to scale the mesh to a lower volume for the unloading procedure seems questionable. This approach does not ensure that the reloaded mesh aligns with the mesh derived from image data. As a result, the geometry used for the simulations is no longer truly patient-specific.

      This mismatch is a significant limitation, as there are established methods available to achieve a proper unloaded configuration, as, e.g., in

      Marx et al. Journal of Computational Physics 463 (2022): 111266. https://doi.org/10.1016/j.jcp.2022.111266, and

      Regazzoni et al. Journal of Computational Physics 457 (2022): 111083. https://doi.org/10.1016/j.jcp.2022.111083.

      As our study aimed at creating a framework for calibration and validation in data-scarce scenarios such as it is often the case in the clinical context, using a compilation of multi-modal data from difference sources, rather than a specific method of personalisation, we did not feel it appropriate to invest significant energy to identify a patient-specific resting geometry, but rather felt that it was important for the resting geometry to fall within population values in terms of diastasis volume. We have clarified this issue in the manuscript and softened claims to Digital Twins in this study. The limitation has been addressed in the updated manuscript, and future work could further address this point.

      (6) Calibration procedure:

      There are apparent flaws in the calibration procedure, or it is not described in sufficient detail.

      We thank the reviewer for raising this point and we have substantially revised the calibration description in the manuscript to clarify the rationale behind each step.

      (a) Step 1: "Sample..." Why? kws and Cal50 are not the most significant parameters in the sensitivity analysis. Kct is a penalty parameter dependent on the numerical framework as described above; "ejection pressure threshold" was never mentioned, is it "P ejection LV" in Table 2? Aiming just for the highest LVEF might neglect non-physiological responses to parameter changes.

      While kws and Cal50 are not the single most significant parameters for LVEF in isolation, they were grouped in Step 1 because they affect both LVEF and peak systolic pressure simultaneously through cross-bridge cycling rate and residual active tension, making it necessary to sample them jointly rather than sequentially. Kct was included because myocardial compressibility affects wall thickening and therefore stroke volume. The ejection pressure threshold is P_ejection_LV in Table 1 and has now been described explicitly in the methods section. Regarding the concern about non-physiological responses: the action potential duration and active tension were monitored throughout calibration and verified to remain within physiological ranges, as now noted in the manuscript.

      (b) Step 2: As systolic pressure is directly dependent on arterial resistance for a 2-element Windkessel model, a uniform sampling approach might not be the best choice here.

      We acknowledge that uniform sampling may not be the most efficient approach for Step 2. However, since arterial resistance influences not only peak systolic pressure but also stroke volume and therefore LVEF, a more targeted approach focusing solely on pressure matching could compromise the LVEF achieved in previous steps. Uniform sampling allowed us to select the value that best balanced both quantities simultaneously.

      (c) Step 3: The authors mention in line 527: "A four-fold increase in GCaL caused an eight-fold increase in cellular active tension peak". An increase in active tension peak results in higher LVEF. So this step is likely to yield the upper boundary of the GCaL interval.

      The reviewer is correct that Step 3 tends to yield a high GCaL value. This was intentional — GCaL was used as a last resort to achieve physiological LVEF after Steps 1 and 2, since the model consistently undershot the target. The upper boundary of the sampled GCaL interval corresponds to a two-fold increase, which remains within the physiological variability bounds applied in previous studies. The resulting action potential duration was verified to remain within physiological ranges.

      (d) Step 4: Why again k_ws? It is not the most significant parameter in the SA.

      kws was resampled in Step 4 not to increase LVEF further, but to specifically target peak ejection rate and dP/dtmax, which were not adequately matched after Step 3. kws is the dominant parameter affecting these ejection dynamics biomarkers in the sensitivity analysis. Resampling at this stage allowed fine-tuning of ejection dynamics while maintaining the LVEF achieved in previous steps.

      (e) Step 5: As far as I can tell, the "diastolic volume change parameter" was mentioned the first time here.

      The diastolic volume change parameter C_pLAV has now been described in the methods section in the Phase 5 passive filling description, where it appears as the inverse of the penalty term controlling the rate of return to diastasis volume in the left ventricle.

      The whole calibration procedure seems to aim for the highest LVEF, and final values of the calibration parameters are not given.

      We note that the calibration procedure does not aim solely for the highest LVEF. As described above, the sequential strategy targets multiple quantities of interest in order of clinical importance: LVEF, peak systolic pressure, peak ejection rate, and peak filling rate, with each step designed to improve a specific subset of biomarkers without compromising those already matched. The final calibrated parameter values are reported in Figure 1F of the revised manuscript.

      (7) Novel features in this paper are actually scarce. A way more advanced calibration strategy with a whole heart model, emulators, and also the ToRORd-Land model was already presented in the study by Strocchi et al. [29]. The calibration to ECGs was presented by some of the same authors in Camps et al. [15], and the analysis of cellular effects was already published in several studies by the same group and in other publications, e.g., by the groups of Severi et al.

      The systematic compilation of credibility criteria spanning ECG morphology, pressure-volume characteristics, strain and displacement represents a novel contribution in itself, providing the field with a reusable evaluation framework. Furthermore, the present study is designed to yield mechanistic insight into how parameters at different scales influence both electrical and mechanical outputs simultaneously. This goal was not tackled in previous publications, which covered individual components, including ECG calibration in Camps et al. [15] and global sensitivity analysis with whole-heart models in Strocchi et al. [29] with no ECG consideration.

      Thus, the work by Camps et al. on ECG calibration was purely electrophysiological and did not investigate the influence of mechanical or haemodynamic parameters on ECG morphology in a fully coupled electromechanical framework. While the effect of mechanical parameters on ECG has been explored by others (e.g. Favino, 2016), this has not previously been examined alongside the relative importance of cellular, mechanical and haemodynamic parameters on pressure-volume characteristics within a single coupled framework. While Strocchi et al. present an emulation strategy, they did not address ECG biomarkers. This distinction is now stated explicitly in the introduction, where we position the present study relative to Camps et al. and Strocchi et al.

      (8) How could the calcium sensitivity Cal50 have such a drastic effect on diastolic function, i.e., filling and end-diastolic volume? As far as I understand from the description, the simulation starts with Phase 0 (loading), Phase 1 (atrial filling), and then in Phase 2, electrical activation ensues and active contraction develops, see also the section starting in line 165. Based on this description, I would expect the end-diastolic volumes to be identical across all Cal50 values. Or are the PV loops shown actually limit cycles established over simulations with multiple beats? This point wasn't explicitly clarified in the manuscript.

      Calcium sensitivity (Cal50) affects not only systolic active tension development but also diastolic residual active tension, i.e. the degree to which the muscle remains partially activated at end diastole. Higher Cal50 values increase this residual tone, effectively stiffening the myocardium during diastolic filling and reducing end-diastolic volume. This mechanism is well established as a contributor to diastolic dysfunction in heart failure [88]. We have clarified this in the manuscript and also clarified that the PV loops shown are single-beat simulations, not limit cycles, with the end-diastolic volume determined by the prescribed filling pressure alongside the passive and residual active stiffness of the myocardium.

      Minor concerns:

      (9) Line 29: The values provided: LVEF of 51%, EDV of 110 mL, and ESV of 50 mL are inconsistent. If these values are all related to the LV, the calculated LVEF should be approximately 54.55%, not 51%.

      The values quoted in the original abstract were rounded approximation, this has been corrected to report EDV=105 mL and ESV=51 mL, which are consistent with the simulated LVEF of 51%.

      (10) "Electromechanical cardiac Digital Twins have had broad applicability ..."

      Many of the cited works here are not true "Digital Twins" but rather static, non-patient-specific models of cardiac electromechanics. In some cases, the geometry may be derived from patient data, but this alone does not qualify the model as a digital twin.

      This sentence in the introduction has been rephrased as ‘Electromechanical cardiac models have had broad applicability...’. Furthermore, as stated earlier, we have removed explicit claims of Digital Twin from the paper while retaining the fact that this study provides a significant step towards rigorous credibility assessment of the high-fidelity electromechanical models that make Digital Twin construction possible.

      (11) While in the abstract and the conclusion, the authors mention "uncertainty quantification", it is mostly a sensitivity analysis that was performed in the paper.

      We have updated the text to say ‘sensitivity analysis’ where appropriate in the abstract, results, and conclusion, and replaced ‘uncertainty ranges’ with ‘variability ranges’ throughout. However, since the sensitivity analyses were performed over biologically informed ranges derived from population variability in the literature, the results are informative about how uncertainty in model inputs propagates to uncertainty in simulated biomarkers. We have therefore retained the framing of sensitivity analysis as a first step towards uncertainty quantification in the abstract and conclusion, and have added a clarifying sentence to the methods to this effect.

      (12) Line 98: As far as I can tell, the conduction velocity assigned to the endocardial surface - intended to mimic the Purkinje fiber network - is never specified. In the section beginning at line 294, only the transmural conduction velocities are reported.

      The endocardial conduction velocity has been specified in the methods section: Purkinje-myocardial junctions were modelled using a fast endocardial activation layer with isotropic conduction velocity of 300 cm/s.

      (13) Line 198, Table1:

      (a) "21/02/2025 11:09:00 AM" on two occasions is maybe not intended

      This has been removed.

      (b) For easing up comparisons, units should be consistent between the initial value and the literature ranges, e.g., PV control parameters, heart rate.

      Units have been made consistent between the initial values and literature ranges throughout Table 1.

      (14) Line 232: "... have already been used to calibrate and validation ...".

      This has been corrected.

      (15) Line 279, Table 2: This table of variability ranges is not entirely clear and could be improved:

      Table 2 has been combined with Table 1 such that the variability ranges sit next to the literature values, for ease of comparison.

      (a) "21/02/2025 11:09:00 AM" is maybe not intended.

      This has been removed.

      (b) use of units should be improved; sometimes it's given in the first column, sometimes in the second column (arterial resistance, compliance), then for k_epi it should be either kPa or kPa/cm.

      Units have been made consistent and the units for k_epi has been added in Table 1.

      (c) units should also be consistent throughout the paper, e.g. in Figure 1 E arterial resistance is Barye.ms/mL while in Table 2 it is mmHg.ms/mL.

      Barye has been removed and replaced by corresponding kPa values throughout the manuscript. This was in the original manuscript since the Alya simulation software were in units of cm, s, g, Barye.

      (d) it is also not clear how variability ranges were chosen; e.g., for arterial compliance,e literature ranges are 0.2-2.73 while the chosen range is [0.1,0.2].

      The previous ranges were chosen to achieve better LVEF. We have now updated the variability ranges to be purely based on literature values and updated the sensitivity analysis results. The ranges are now presented in Table 1 alongside the literature values for ease of comparison.

      (e) For Kct, the initial value in Table 1 is 5000kPa, the literature values are between 10 and 3333, and then the variability range is [10,500]? I guess there is a typo in one of these values.

      This has been corrected in the new Table 1.

      (f) Table 1 and 2 are in parts redundant.

      Table 1 and 2 have been combined into a single new Table 1.

      (15) Line 290: It should be uvc_l for the longitudinal coordinate.

      This has been corrected.

      (16) Line 388, Table 3, regarding values for pressure volume from reference [49]:

      (a) the number of participants is 800, including males and females; not only females, see also Table 12 https://jcmr-online.biomedcentral.com/articles/10.1186/s12968-017-0327-9/tables/12

      (b) why using female values here while having mixed sex for most of the others? Because the model is female?

      The reviewer is correct that reference [49] reports values from a mixed-sex cohort of approximately 800 participants. We used the female-specific values from Table 12 of that reference because the biventricular mesh used in this study was derived from a female subject, making sex-matched reference values the most appropriate comparison. This has been clarified in the manuscript.

      (17) Figure 5: What is Jup; why did you choose 0.93 x Jup as reference? Also in Figure 4, why did you use 0.93 x GCal as a reference?

      J_up refers to the SERCA<sup2+</sup> reuptake current, which has been relabelled as SERCA throughout the manuscript for consistency. The reference value of 0.93× was used because the sensitivity analysis sampled parameters uniformly between 50% and 200% of baseline using a fixed number of samples, and no sample fell exactly at 1.0×. The closest sampled value was 0.93×, which was therefore used as the reference. This has been clarified in the figure caption.

      (18) Line 377: The link to the GitHub repository does not work.

      This link has now been made publicly available.

      (19) Line 397, Table 3: for the sake of completeness, all abbreviations should be included: e.g., SVL, ESP, EDV, ESV are not included.

      This has been written out in full in the new Table 2.

      (20) Tick marks in many figures are not readable, e.g., Figure 5 and all the Figures in the appendix.

      Tick mark sizes and line widths have been increased across Figures 4, 5, and all appendix figures. The figures have been replotted and updated in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      Minor Comments:

      (1) The provided GitHub link https://github.com/jennyhelyanwe/Alya_input_setup/ does not work, potentially because the repository is private. It would be nice to see the repository during the review.

      This link has now been made publicly available.

      (2) Table 3: Can you include the simulation outputs obtained for validation (with an error indication)? This would summarize the validation that's currently spread out over the results section.

      A new Table 3 has been added to the manuscript under the validation section, summarising the simulated values for all deformation and strain biomarkers alongside their reference ranges. The calibration and validate datasets are now reported separately in Tables 2 and 3, respectively.

      (3) Figure 1: Add axis labels to all plots.

      Axis labels have been added to all subplots in Figure 1 in the revised manuscript. Simulated pseudo-ECG amplitudes are normalised and therefore dimensionless.

      (4) Figure 2: Simulated and in vivo strains with exactly the same axes (size, range, ticks) and add grid lines to enable a comparison. Add the mean values of each in the other plot.

      The revised Figure 2 now includes the median in vivo strain values from Moulin et al. overlaid as a red dashed reference line on the simulation panels, enabling direct visual comparison. The simulated mean could not be overlaid on the in vivo panels as the original Moulin et al. figure data are not publicly available for replotting. Exact axis matching was not applied as this would cause some simulated curves to fall outside the visible range, obscuring the model behaviour.

      (5) Figure 3: The thickness (relative importance of the connections) is impossible to see in this plot. Instead of having gray background connections, remove them entirely below a certain threshold. Make the differences in thickness more pronounced or introduce a continuous color scale for the magnitude of the positive or negative correlation. Alternatively, you could rank the parameters from least to most important in each subfigure A-D and/or provide some numeric values.

      Figure 3 has been updated. All non-significant connections (|r| < 0.6 or p > 0.05) have been removed entirely, and gray lines have been removed in each subfigure, making the significant relationships clearer. A continuous blue-to-red colour scale has been applied to indicate the direction of correlation (blue: negative, red: positive), with line thickness proportional to the magnitude of the r-value.

      (6) Figure 3 and Table 2: Why were material parameters b, bf, bs, and bfs omitted from this study (but included a, af, as, and afs)?

      The b parameters (b, bf, bs, bfs) appear in the exponent of the Holzapfel-Ogden constitutive law and are strongly coupled to the a parameters (a, af, as, afs), which carry units of kPa. In practice, the b parameters can only be reliably identified from ex vivo multiaxial stretch experiments, whereas the a parameters can be estimated from clinical imaging data. Since our study focuses on calibration and validation in a clinical data setting, we included only the a parameters in the sensitivity analysis, consistent with previous personalisation studies.

      (7) Figure 4: What do the dotted lines represent?

      The dotted lines in Figure 4E highlight the increased longitudinal shortening with increasing GCaL, showing the basal plane moving towards the apex while the apical position remains unchanged due to the pericardial constraint. This has been clarified in the figure caption.

      (8) Figures 4, 5, A2-45: Can you use a continuous color scale (e.g., from blue to red) for low to high parameter uncertainty?

      A continuous blue-to-red colour scale has been applied to Figures 4, 5, and all appendix figures A2–A6, where blue indicates the lowest parameter value and red indicates the highest. A colour bar has been added to each figure for reference.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Wang et al., recorded concurrent EEG-fMRI in 107 participants during nocturnal NREM sleep to investigate brain activity and connectivity related to slow oscillations (SO), sleep spindles, and in particular their co-occurrence. The authors found SO-spindle coupling to be correlated with increased thalamic and hippocampal activity, and with increased functional connectivity from the hippocampus to the thalamus and from the thalamus to the neocortex, especially the medial prefrontal cortex (mPFC). They concluded the brain-wide activation pattern to resemble episodic memory processing, but to be dissociated from task-related processing and suggest that the thalamus plays a crucial role in coordinating the hippocampal-cortical dialogue during sleep.

      The paper offers an impressively large and highly valuable dataset that provides the opportunity for gaining important new insights into the network substrate involved in SOs, spindles, and their coupling.

      Thank you for this encouraging assessment. We appreciate your recognition of the value of the dataset and of the questions it allows us to address. Below, we respond to each of your points directly and revise the manuscript accordingly.

      Comments on revisions:

      Re 1: The revised introduction now cites a couple of papers but discusses them only very superficially, lumping together several studies with very different key results. This is still not very informative for the reader and does not sufficiently acknowledge previously published work. Here are two examples to illustrate this:

      (a) "These studies have generally reported that slow oscillations are associated with widespread cortical and subcortical BOLD changes, whereas spindles elicit activation in the thalamus, as well as in several cortical and paralimbic regions." Several studies even showed e.g., a clear activation of the hippocampus and parahippocampal gyrus associated with spindles, not just the thalamus

      Thank you for this comment. We agree that our previous sentence was too broad and did not sufficiently reflect the range of findings in the sleep literature. We have therefore rewritten the Introduction to state explicitly that spindle-related BOLD changes have been reported not only in the thalamus, but also in cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).

      Introduction, Page 3-4, Lines 58-62

      “Consistent with this view, prior human EEG-fMRI studies have reported spindle-related activation not only in the thalamus, but also in the hippocampus and adjacent parahippocampal gyrus (Bergmann et al., 2012; Schabus et al., 2007). Spindle-related activity has also been linked to striatal engagement, suggesting a broader network that may support memory-related processing during sleep (Fogel et al., 2017).”

      Introduction, Page 4, Lines 71-78

      “Previous EEG-fMRI studies on sleep have examined both global sleep characteristics (Hale et al., 2016; Moehlman et al., 2019) and the neural correlates of specific waves, including slow oscillations and spindles. These studies have generally shown that slow oscillations are associated with widespread cortical and subcortical BOLD changes (Czisch et al., 2009; Ilhan-Bayrakcı et al., 2022; Picchioni et al., 2011), whereas spindles have been linked not only to thalamic activation but also to cortical and paralimbic regions, including the hippocampus and parahippocampal gyrus (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      Introduction, Page 5, Lines 103-106

      “This coupling was associated with increased activation in both the thalamus and hippocampus, with functional connectivity patterns suggesting thalamic coordination of hippocampal-cortical communication, in line with prior EEG-fMRI studies of spindle-related activity (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Schabus et al., 2007).”

      (b) "Although these findings provide valuable insights into the BOLD correlates of sleep rhythms, they often do not employ sophisticated temporal modeling (Huang et al., 2024) [, ...]." - previous studies have used e.g., spindle event-related regressors with individual spindle amplitudes as parametric modulators, first and second order derivatives of the HRF function, as well as PPI connectivity analyses, which I would consider rather sophisticated temporal modelling.

      We agree that several previous studies have already employed sophisticated modelling approaches, including parametric modulation, HRF derivatives, and PPI analyses (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Our intention was not to suggest that such methods are absent from the literature.

      Rather, we aimed to highlight that most prior work has focused on modelling individual SO or spindle events, whereas explicit modelling of their temporal interaction (e.g., SO-spindle coupling) has been less commonly addressed. We have revised the sentence to clarify this point more precisely.

      Introduction, Page 4 Lines 78-82

      “Although these findings provide important insight into the BOLD correlates of sleep rhythms, most previous studies have focused on individual oscillatory events rather than explicitly modelling their temporal interaction (Bergmann et al., 2012; Caporro et al., 2012; Fogel et al., 2017; Picchioni et al., 2011). Only a few recent studies have begun to examine coupling between rhythms directly, for example Huang et al. (2024).”

      Re 4+9: The short overall recordings in some subjects on the one hand and the large number of spindles and SOs detected in N1 sleep stages are still highly concerning, in fact even more so, now that the actual numbers have been provided in the Supplementary Tables. Either the sleep staging or the detection of SO and spindle events must be incorrect. I understand that for specific EEG analysis and fMRI modelling purposes sometimes slightly different thresholds are used as compared to clinical sleep staging, but several parameters here are alarmingly off.

      (a) Given that proper NREM sleep (N2+N3) is the relevant stage for the analyses conducted in this paper, some of the N2+N3 durations are very short (eg 7-8 min) while those subjects' results have the same impact on the group level analyses as those with >100 min of N2+N3. Either subjects with very little relevant data (not overall recording time but N2+N3 time) should be excluded or weighting subject data for the group analyses according to the amount od contributed data should be done.

      Thank you for the suggestion. It is true that participants with very little N2/3 sleep could contribute noisier subject-level estimates to the group analysis. We therefore checked this directly. Only three participants contributed less than 10 min of N2/3 sleep, and excluding them did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant after exclusion, t<sub>(103)</sub> =2.50, p = 0.0071, compared with t<sub>(106)</sub> = 2.50, p = 0.0070 in the full sample. We have added this control analysis to the Results so that the robustness of the group findings is explicit in the manuscript.

      Results, Page 11-12, Lines 238-250

      To ensure the results were not driven by individual differences or parameter selection, we conducted a series of control analyses. First, we excluded participants with less than 10 minutes of N2/3 sleep. Only three participants met this criterion, and their exclusion did not change the main results. For example, hippocampal activation during SO-spindle coupling remained significant (t<sub>(103)</sub> = 2.50, p = 0.0071), comparable to the full sample (t<sub>(106)</sub> = 2.50, p = 0.0070). Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st-80th percentile; Fig. S6). Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).

      (b) The authors argue that the SO and spindle detection algorithms are valid since widely used and that they were developed for N2+N3 stages, which is why they will also detect events in other stages: "While, because the detection methods for SO and spindle are based on percentiles, this method will always detect a certain number of events when used for other stages (N1 and REM) sleep data, but the differences between these events and those detected in stage N23 remain unclear." I do agree that with very liberal thresholds, also SO and spindle vents may be detected in other stages, but it shouldn't be that many. If the percentiles of amplitude thresholds were defined based on properly scored N2+N3 stages only, very few events should be detected (erroneously!) in N1, as the occurrence of K-complexes (isolated SOs) and spindles per definition makes it N2, and during REM sleep only very few spindles and SOs are allowed to occur, without scoring it NREM instead. For the first subject (just as example, but with similar numbers for the rest of the sample), reveals as many as 60 SOs and 31 spindles within 8 min of N1 sleep (Table S2) as well as 13 SOs and 7 spindles within 2 min of REM sleep (Table S4). These numbers are completely unrealistic and question the correctness of the sleep staging as well as the physiological relevance of the EEG graphoelements identified as SO and spindles. It also completely undermines the interpretability of the respective event regressors for the fMRI analyses.

      (c) Likely, given the large numbers of coupled SO-spindle events and the apparently very low amplitude criteria for event identification, also the number of SO-spindle couplings is likely severely overestimated.

      We thank the reviewer for raising this important point. We agree with you that the original stage-wise percentile thresholding could inflate the apparent number of SOs and spindles outside N2/3 sleep. In the original analysis, the thresholds were estimated separately within each sleep stage. As you point out, this procedure can force the detector to label a relatively large number of events in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 SOs or spindles. We have therefore revised the detection procedure. Following your concern and Reviewer 2’s suggestion, the SO and spindle thresholds are now defined only from N2/3 sleep within each participant, where SOs and spindles are most abundant and physiologically expected to occur. These fixed N2/3-derived thresholds were then applied unchanged to N1 and REM for descriptive reporting. This avoids the artificial normalisation of event detection across sleep stages that can arise when each stage has its own percentile threshold. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We agree with you that the remaining detections in N1 and REM should not be treated as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We therefore report them only as descriptive detector outputs obtained under a fixed N2/3-derived threshold (see Table S2, S4 in the revised manuscript). We do not use them to support any physiological claim about SO-spindle coupling in N1 or REM.

      This point is also important for the fMRI analyses. You are right that inflated N1 or REM detections would undermine the interpretability of event regressors if those detections entered the EEG-informed fMRI models. They did not. All EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep, where SOs, spindles, and their coupling are physiologically expected and where the detection thresholds were defined. Thus, the central fMRI event regressors were based only on N2/3 events, not on detections from N1 or REM.

      We also agree with you that the absolute number of detected SO-spindle couplings depends on the chosen detection threshold. For this reason, we tested whether the main EEG-fMRI result depended on the specific detector setting. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We therefore do not argue that the detector provides a uniquely correct absolute count of SOs, spindles, or coupling events in every sleep stage. Our conclusion is more specific. The main N2/3 EEG-fMRI finding is robust across a reasonable range of SO detection thresholds, detections in N1 and REM are reported only descriptively, and the physiological interpretation of SO-spindle coupling is restricted to N2/3 sleep.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Re 10: The rationale for using a lateralized frontal electrode (F3) for both SO (should have been at least bilateral or central) and spindle detection (should have been a centro-parietal electrode) is not convincing. Other EEG-fMRI spindle or SO papers have used a number of frontal (SO) or centro-parietal (spindles) electrodes averaged or even approaches including all EEG electrodes. Searching events with low thresholds at suboptimal recording sites does not dot this highly valuable dataset justice.

      We thank the reviewer for this important comment. We agree that this choice is more sensitive to frontal SOs than to the centro-parietal fast spindle component. Our choice of F3 was driven by the practical constraints of prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to FCz can have reduced signal contrast relative to the reference, and electrodes near the vertex are also more vulnerable to prolonged pressure against the MRI head coil when participants sleep supine for several hours. In this setting, frontal electrodes provided more stable signal quality across the recording. Because the EEG events were used primarily as temporal markers for fMRI modelling, our priority was to obtain reliable event timing during N2/3 sleep rather than to estimate the full scalp topography of SOs and spindles.

      We also agree with your concern that this valuable dataset would ideally be analysed with multichannel detection strategies. To test whether the main result depended on the single lateralised F3 site, we repeated the main EEG-informed fMRI analysis using Fz, a midline frontal electrode. The result was unchanged. Hippocampal activation during SO-spindle coupling remained significant when events were detected from Fz, t<sub>(106)</sub> = 2.47, p = 0.0076, closely matching the original F3-based result, t<sub>(106)</sub> = 2.50, p = 0.0070. This control analysis does not remove the limitation that centro-parietal fast spindles may be underrepresented, and we do not claim that it does. It does show, however, that the main hippocampal fMRI finding is not driven by idiosyncratic detections from one lateralised frontal electrode. We have made this clearer in the revised manuscript.

      Finally, your concern about low thresholds is also important. As described in our response above, the revised analysis now defines SO and spindle thresholds from N2/3 sleep and applies these thresholds uniformly for descriptive comparisons across stages. We also tested the robustness of the hippocampal fMRI result across SO detection thresholds, and the effect remained significant across the 71st to 80th percentile range. We have therefore narrowed the interpretation in the revised manuscript. The main EEG-fMRI result reflects BOLD activity associated with frontal-channel-detected SO-spindle coupling during N2/3 sleep, rather than a full multichannel characterisation of all SO and spindle topographies.

      Results, Page 12, Lines 247-250

      “Third, to test whether the results depended on the use of a single lateralised frontal electrode, we repeated the EEG-informed fMRI GLM using events detected from Fz. Hippocampal activation during SO-spindle coupling again remained significant (t<sub>(106)</sub> = 2.47, p = 0.0076), closely matching the original F3-based result (t<sub>(106)</sub> = 2.50, p = 0.0070).”

      Discussion, Page 18, Lines 380-394

      “Third, sleep oscillation detection was based on a single frontal electrode. This choice improved signal stability and event timing in the prolonged simultaneous EEG-fMRI setting, but it did not exploit the full multichannel EEG information and cannot characterise the full spatial distribution of SOs and spindles. In particular, F3-based detection may be more sensitive to frontal SOs and frontal sigma activity than to the centro-parietal fast spindle component. We therefore interpret the EEG-informed fMRI results as reflecting BOLD activity associated with frontal-channel-detected SOs, spindles, and their coupling during N2/3 sleep. Future studies using multichannel or source-informed detection strategies, with separate treatment of slow and fast spindles, will be better suited to capture the spatial dynamics of these sleep oscillations. Fourth, the use of large anatomical ROIs may mask subregional contributions of specific thalamic nuclei or hippocampal subfields. Finally, without a memory task, we cannot establish a direct behavioral link between sleep-rhythm-locked activation and memory consolidation. Future studies combining ultra-high-field fMRI or iEEG with cognitive tasks, as well as multichannel or source-informed detection strategies that separately characterize slow and fast spindles, will be better suited to refine our understanding of subregional network dynamics and the functional significance of sleep oscillations.”

      Methods, Page 25, Lines 556-566

      “It is worth noting that the primary aim of EEG rhythm detection was to identify reliable event times for EEG-informed fMRI modelling. Detection was performed on the F3 electrode because this channel provided stable signal quality during prolonged nocturnal EEG-fMRI recordings. In our MR-compatible EEG setup, FCz was used as the online reference. Central electrodes close to this reference, and electrodes near the vertex that were in prolonged contact with the head coil during supine sleep, were more susceptible to reduced signal contrast, impedance drift, and pressure-related degradation of electrode-scalp contact. We therefore used F3 as a pragmatic choice to maximize reliable event timing in N2/3 sleep. This choice was not intended to characterise the full scalp topography of SOs or spindles, and it may underrepresent the centro-parietal fast spindle component. As a sensitivity analysis, we repeated the main EEG-informed fMRI GLM using Fz, a midline frontal electrode, with the same detection and modelling procedure.”

      Re 7: It is not clear to me why/how larger voxels would reduce susceptibility-related distortions and partial volume effects. Usually, the opposite is true. This should be elaborated.

      What we meant was that we chose a relatively large voxel size to preserve signal-to-noise ratio and whole-brain coverage within a feasible repetition time for a long overnight EEG-fMRI protocol. This choice is useful for maintaining BOLD sensitivity in sleep recordings, where head motion, physiological noise, and participant comfort are major practical constraints. We agree that it may not be accurate to describe it as reducing susceptibility-related distortion or partial volume effects.

      We have rewritten the Methods to state this trade-off directly. The voxel size of 3.5 × 3.5 × 4.2 mm<sup>3</sup> allowed whole-brain coverage with a TR of 2000 ms, which was important for modelling sleep-rhythm-related BOLD responses across the whole brain during prolonged nocturnal recordings. A smaller voxel size would have improved spatial specificity, but would also have required either a longer TR, reduced brain coverage, or lower SNR, none of which would have been ideal for the present EEG-fMRI sleep design. We now explicitly acknowledge the cost of this choice.

      Methods, Page 21 Lines 453-463

      “For the functional scans, whole-brain images were acquired using a T2*-weighted gradient echo-planar imaging (EPI) sequence sensitive to the BOLD contrast. The sequence parameters were as follows: 33 slices in interleaved ascending order, TR = 2000 ms, TE = 30 ms, voxel size = 3.5 × 3.5 × 4.2 mm<sup>3</sup>, FA = 90°, matrix = 64 × 64, gap = 0.7 mm. A relatively large voxel size was chosen to preserve signal-to-noise ratio while maintaining whole-brain coverage within a feasible repetition time. This compromise was important for the prolonged overnight EEG-fMRI sleep protocol, where head motion, physiological noise, participant comfort, and sustained acquisition stability are substantial practical constraints (Bodurka et al., 2007; Laufs et al., 2008). A smaller voxel size would have improved spatial specificity, but would have required either a longer repetition time, reduced brain coverage, or lower signal-to-noise ratio.”

      Reviewer #2 (Public review):

      In this study, Wang and colleagues aimed to explore brain-wide activation patterns associated with NREM sleep oscillations, including slow oscillations (SOs), spindles, and SO-spindle coupling events. Their findings reveal that SO-spindle events corresponded with increased activation in both the thalamus and hippocampus. Additionally, they observed that SO-spindle coupling was linked to heightened functional connectivity from the hippocampus to the thalamus, and from the thalamus to the medial prefrontal cortex-three key regions involved in memory consolidation and episodic memory processes.

      This study's findings are timely and highly relevant to the field. The authors' extensive data collection, involving 107 participants sleeping in an fMRI while undergoing simultaneous EEG recording, deserves special recognition. If shared, this unique dataset could lead to further valuable insights.

      Thank you for this encouraging assessment. We appreciate your recognition of the effort involved in collecting this simultaneous EEG-fMRI sleep dataset. Below, we respond directly to your remaining concern.

      Comments on revisions:

      The authors' efforts in revising the manuscript and addressing the reviewers' comments are certainly commendable. However, I remain concerned about potential issues in detecting sleep-related oscillations (SOs, spindles, and consequently coupled SO-spindle events), which may arise due to suboptimal parameter selection or inaccurate sleep staging, potentially impacting all subsequent analyses.

      A review of Supplementary Tables 1-4 reveals an unusually high number of detected SOs and spindles during sleep stage N1 and REM sleep. While the authors correctly note that a percentile-based detection approach will always identify a certain number of events across sleep stages, the particularly high counts in N1 and REM are concerning. To mitigate the limitations of this method, the authors could have performed event detection independently of sleep stages (i.e., across the entire dataset for each participant) and subsequently assigned the detected events to the corresponding sleep stages. If the event counts in N1 and REM remained disproportionately high, this would indicate a fundamental issue with the detection procedure.

      In the previous version, thresholds were estimated separately within each sleep stage. As you point out, this can force the detector to identify a relatively large number of SOs and spindles in N1 and REM, even when those waveforms should not be interpreted as canonical N2/3 events.

      We have therefore revised the detection procedure so that event detection is no longer based on separate stage-wise thresholds. Following the logic of your suggestion, we first defined a fixed threshold for each participant and then assigned the detected events to their corresponding sleep stages afterwards. We used N2/3 sleep to define the SO and spindle thresholds because this is the stage in which these events are physiologically expected and most reliably observed. These same N2/3-derived thresholds were then applied unchanged to N1 and REM. This avoids the circularity of forcing a percentile-defined number of detections within each sleep stage.

      With this revised procedure, detections outside N2/3 are clearly lower than those in N2/3. The mean densities are 2.95 SOs/min, 2.71 spindles/min, and 0.75 coupling events/min in N1, and 2.07 SOs/min, 1.81 spindles/min, and 0.43 coupling events/min in REM. We also agree with you that the remaining detections in N1 and REM should not be interpreted as physiological equivalents of canonical N2/3 SOs, spindles, or SO-spindle complexes. We now state this explicitly in the manuscript. They are reported only as descriptive detector outputs under the fixed N2/3-derived criterion. And we have revised all relevant sections of the manuscript, including “[Results, Page 6-7 Lines 134-148]; [Fig. 1e]; [Results, Page 9 Lines 175-191]; [Fig. 2b]; [Methods, Page 25-27, Lines 567-604]; [Fig. S2-S4]; [Table S2, S4].”

      We also would like to clarify our sleep staging procedure. The sleep staging was first performed using an established automated algorithm, YASA toolkit (Vallat & Walker, 2021), and then manually reviewed by two sleep experts. More importantly for the central results, all EEG-informed fMRI GLM and PPI analyses were restricted to N2/3 sleep. Thus, N1 and REM detections did not enter the event regressors used for the main fMRI analyses and do not affect the interpretation of the hippocampal or thalamic findings.

      Finally, we agree that the absolute number of SO-spindle coupling events depends on the detection threshold. We therefore tested whether the main fMRI result depended on the specific SO threshold. Hippocampal activation during SO-spindle coupling remained significant when the SO detection threshold was varied between the 71st and 80th percentiles, as shown in Fig. S6. We have made this clearer in the revised manuscript.

      Results, Page 6-7 Lines 134-148

      “Each sleep stage is characterised by distinct spectral properties and rhythmic waveforms, serving as physiological markers (Fig. 1c). Because SO and spindle detection relies on amplitude-based percentile thresholds, we avoided estimating separate thresholds within each sleep stage. Instead, for each participant, the SO and spindle thresholds were defined from N2/3 sleep only, where these rhythms are most abundant and physiologically expected, and the same fixed thresholds were then applied to N1 and REM for descriptive comparison.”

      “Under this fixed N2/3-derived thresholding, detected SOs and spindles were larger and more frequent in N2/3 than in N1 or REM. SO and spindle amplitudes were significantly higher during N2/3 sleep (SO: 25.59 ± 1.49 μV; spindle: 7.39 ± 0.27 μV) than during N1 (SO: 20.15 ± 2.32 μV; spindle: 5.23 ± 0.27 μV) and REM sleep (SO: 19.84 ± 1.22 μV; spindle: 5.60 ± 0.22 μV; all p < 1e-4; Fig. 1e, Fig. S2). The corresponding event densities showed the same pattern, with 9.64 ± 0.25 SOs/min and 4.19 ± 0.10 spindles/min in N2/3, compared with 2.95 ± 0.16 SOs/min and 2.71 ± 0.14 spindles/min in N1, and 2.07 ± 0.17 SOs/min and 1.81 ± 0.14 spindles/min in REM (all p < 1e-4). We therefore report detections in N1 and REM only as descriptive outputs of the detector under a fixed N2/3-derived criterion, rather than as physiological equivalents of canonical N2/3 SOs or spindles.”

      Fig. 1 legend, Page 8, Line 166-172

      “e, Amplitudes (μV) of detected SOs (left) and spindles (right) across sleep stages. SO and spindle detection thresholds were defined from N2/3 sleep within each participant and then applied unchanged to N1 and REM for descriptive comparison. Detections in N1 and REM should therefore be interpreted as detector outputs under this fixed N2/3-derived criterion. The SO amplitudes were measured from the 0.16-1.25 Hz filtered EEG data, and spindle amplitudes were measured from the 12-16 Hz filtered EEG data. Each dot represents an individual participant. Error bars indicate SEM. *** p < 0.001.”

      Results, Page 9 Lines 175-191

      “SO-spindle coupling is considered important for sleep-dependent memory consolidation. In the current study, using the same N2/3-derived detection thresholds described above, we found that SO-spindle coupling occurred most frequently during N2/3 sleep (2.46 ± 0.06 events/min). Coupling density was significantly lower in N1 (0.75 ± 0.05 events/min, t<sub>(106)</sub> = 23.54, p < 1e-4) and REM sleep (0.43 ± 0.04 events/min, t<sub>(106)</sub> = 31.24, p < 1e-4; Fig. 2b, Table S2-S4), consistent with the expected predominance of SO-spindle coupling in NREM sleep (Ngo et al., 2013; Staresina et al., 2015). As with the individual SO and spindle detections, coupling events detected in N1 and REM were retained only for descriptive stage-wise reporting (see Table S2, S4). They were not used to support physiological claims about SO-spindle coupling in these stages, and they were not entered into the EEG-informed fMRI analyses. All subsequent fMRI GLM and PPI analyses were restricted to N2/3 sleep.”

      “After extracting all N2/3 EEG epochs in which SO-spindle coupling occurred, we analysed their spectral and phase characteristics. The spindles were most likely to occur slightly before the UP-state peak of SOs (Fig. 2a, e), aligning with results from both animal studies (Maingret et al., 2016) and human research (Staresina et al., 2015). In our data, this pattern was consistent across subjects (Fig. 2d, Rayleigh test: z = 9.51, p < 1e-4), with the peak of the spindle aligned at an SO phase of −41.61 ± 0.86° (the SO UP-state peak is 0°).”

      Fig. 2 legend, Page 10, Line 202-205

      “b, SO-spindle coupling density across sleep stages, using SO and spindle detections obtained with fixed N2/3-derived thresholds. Coupling events in N1 and REM are shown only for descriptive comparison. The EEG-informed fMRI analyses used N2/3 coupling events only.”

      Results, Page 11-12, Lines 242-247

      “Second, because the absolute number of detected SO-spindle coupling events depends on the SO detection threshold, we examined whether the main EEG-fMRI results were sensitive to this parameter. To this end, we varied the SO percentile threshold and reconstructed the EEG-informed GLM at each level. Hippocampal activation during SO-spindle coupling remained significant across a range of thresholds (71st - 80th percentile; Fig. S6).”

      Methods, Page 25-26, Lines 567-575

      “Detection of SOs. Data were first bandpass-filtered between 0.16 and 1.25 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). After identifying all positive-to-negative zero crossings, potential SOs were defined based on the interval between consecutive zero crossings, ranging from 0.8 s to 3 s. For each potential SO, we calculated the amplitude range as the peak minus the trough. For each participant, the amplitude threshold was defined as the 75th percentile of candidate SO amplitude ranges observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. Only candidates exceeding this threshold were labelled as SOs, following previous work (Schreiner et al., 2021).”

      Methods, Page 26, Lines 576-583

      “Detection of sleep spindles. Detection of sleep spindles. Data were bandpass-filtered between 12 and 16 Hz (Butterworth filter, order 3, bidirectional filtering for zero phase). The root mean square (RMS) of the filtered signal was computed with a 200 ms sliding time window. For each participant, the spindle threshold was defined as the 75th percentile of RMS values observed during N2/3 sleep. This fixed N2/3-derived threshold was then applied unchanged across the recording for descriptive stage-wise summaries. Detected events were assigned to N1, N2/3 or REM according to the sleep-stage label at the event time. RMS segments exceeding this threshold for 0.5 s to 3 s were identified as spindles (Staresina et al., 2015).”

      Methods, Page 26, Lines 584-591

      “Detection of SO-spindle couplings. From the detected SOs and spindles, we identified the peak time of each spindle. Within each SO interval, we checked whether a spindle peak occurred; if so, that SO was labelled as an SO-spindle coupling event. For descriptive stage-wise summaries, coupling events were assigned to the sleep stage of the corresponding SO trough. For every SO-spindle coupling event, an epoch was created time-locked to the SO trough as the central reference, following Schreiner et al. (2021). We extracted data in a [−4 s to 4 s] window around this point, forming the epoch for each coupling event. For the EEG-informed fMRI analyses, only SO, spindle and SO-spindle coupling events detected during N2/3 sleep were used.”

      Methods, Page 26-27, Lines 592-604

      “The detection procedures described above were developed primarily for N2 and N3 sleep, where SOs, spindles and their coupling are physiologically expected and most reliably observed (Hahn et al., 2020; Helfrich et al., 2019; Helfrich et al., 2018; Ngo, Fell, & Staresina, 2020; Schreiner et al., 2022; Schreiner et al., 2021; Staresina et al., 2015; Staresina et al., 2023). Because percentile-based thresholds can otherwise force the detector to label events in every sleep stage, we did not estimate separate thresholds within N1 or REM. Instead, for each participant, all SO and spindle thresholds were defined from N2/3 sleep and then applied uniformly across the recording. Tables S1 and S3 report detailed statistical information on sleep rhythm and N2/3 events detection. The N1 and REM events detection reported in Tables S2 and S4, and illustrated in Fig. S2-S4, should therefore be interpreted as descriptive detector outputs under this fixed N2/3-derived criterion, rather than as evidence for canonical N2/3 SOs, spindles or physiological SO-spindle complexes in those stages. These detections were not used in the EEG-informed fMRI GLM or PPI analyses, which were restricted to N2/3 sleep.”

      Reviewer #3 (Public review):

      Summary:

      Wang et al., examined the brain activity patterns during sleep, especially when locked to those canonical sleep rhythms such as SO, spindle, and their coupling. Analyzing data from a large sample, the authors found significant coupling between spindles and SOs, particularly during the up-state of the SO. Moreover, the authors examined the patterns of whole-brain activity locked to these sleep rhythms. The authors next investigated the functional connectivity analyses, and found enhanced connectivity between the hippocampus and the thalamus and the medial PFC. These results reinforced the theoretical model of sleep-dependent memory consolidation, such that SO-spindle coupling is conducive for systems-level memory reactivation and consolidation.

      Strengths:

      There are obvious strengths in this work, including the large sample size, state-of-the-art neuroimaging and neural oscillation analyses, and the richness of results. The results now inform hemodynamic neural activity that coincided with SO-spindle couplings.

      Weaknesses:

      My earlier comments were about the inability to make inferences on memory given the lack of memory tasks, and the weakness in using the open-ended cognitive state decoding.

      Comments on revisions:

      The current revision has addressed these major concerns. The authors expanded discussions regarding the theoretical implications of the work in a more nuanced manner.

      Thank you for taking the time to re-evaluate the manuscript. We are pleased that the revised Discussion now reads as more nuanced, especially in relation to the limits of the memory-related interpretation. Your earlier comments helped us sharpen both the claims and the framing, and we are grateful for that.

      References:

      Bergmann, T. O., Mölle, M., Diedrichs, J., Born, J., & Siebner, H. R. (2012). Sleep spindle-related reactivation of category-specific cortical regions after learning face-scene associations. Neuroimage, 59(3), 2733-2742.

      Bodurka, J., Ye, F., Petridou, N., Murphy, K., & Bandettini, P. A. (2007). Mapping the MRI voxel volume in which thermal noise matches physiological noise—implications for fMRI. Neuroimage, 34(2), 542-549.

      Caporro, M., Haneef, Z., Yeh, H. J., Lenartowicz, A., Buttinelli, C., Parvizi, J., & Stern, J. M. (2012). Functional MRI of sleep spindles and K-complexes. Clinical neurophysiology, 123(2), 303-309.

      Czisch, M., Wehrle, R., Stiegler, A., Peters, H., Andrade, K., Holsboer, F., & Sämann, P. G. (2009). Acoustic oddball during NREM sleep: a combined EEG/fMRI study. PloS one, 4(8), e6749.

      Fogel, S., Albouy, G., King, B. R., Lungu, O., Vien, C., Bore, A., Pinsard, B., Benali, H., Carrier, J., & Doyon, J. (2017). Reactivation or transformation? Motor memory consolidation associated with cerebral activation time-locked to sleep spindles. PloS one, 12(4), e0174755.

      Hahn, M. A., Heib, D., Schabus, M., Hoedlmoser, K., & Helfrich, R. F. (2020). Slow oscillation-spindle coupling predicts enhanced memory formation from childhood to adolescence. Elife, 9, e53730.

      Hale, J. R., White, T. P., Mayhew, S. D., Wilson, R. S., Rollings, D. T., Khalsa, S., Arvanitis, T. N., & Bagshaw, A. P. (2016). Altered thalamocortical and intra-thalamic functional connectivity during light sleep compared with wake. Neuroimage, 125, 657-667.

      Helfrich, R. F., Lendner, J. D., Mander, B. A., Guillen, H., Paff, M., Mnatsakanyan, L., Vadera, S., Walker, M. P., Lin, J. J., & Knight, R. T. (2019). Bidirectional prefrontal-hippocampal dynamics organize information transfer during sleep in humans. Nature Communications, 10(1), 3572.

      Helfrich, R. F., Mander, B. A., Jagust, W. J., Knight, R. T., & Walker, M. P. (2018). Old brains come uncoupled in sleep: slow wave-spindle synchrony, brain atrophy, and forgetting. Neuron, 97(1), 221-230. e224.

      Huang, Q., Xiao, Z., Yu, Q., Luo, Y., Xu, J., Qu, Y., Dolan, R., Behrens, T., & Liu, Y. (2024). Replay-triggered brain-wide activation in humans. Nature Communications, 15(1), 7185.

      Ilhan-Bayrakcı, M., Cabral-Calderin, Y., Bergmann, T. O., Tüscher, O., & Stroh, A. (2022). Individual slow wave events give rise to macroscopic fMRI signatures and drive the strength of the BOLD signal in human resting-state EEG-fMRI recordings. Cerebral Cortex, 32(21), 4782-4796.

      Laufs, H., Daunizeau, J., Carmichael, D. W., & Kleinschmidt, A. (2008). Recent advances in recording electrophysiological data simultaneously with magnetic resonance imaging. Neuroimage, 40(2), 515-528.

      Maingret, N., Girardeau, G., Todorova, R., Goutierre, M., & Zugaro, M. (2016). Hippocampo-cortical coupling mediates memory consolidation during sleep. Nature Neuroscience, 19(7), 959-964.

      Moehlman, T. M., de Zwart, J. A., Chappel-Farley, M. G., Liu, X., McClain, I. B., Chang, C., Mandelkow, H., Özbay, P. S., Johnson, N. L., & Bieber, R. E. (2019). All-night functional magnetic resonance imaging sleep studies. Journal of neuroscience methods, 316, 83-98.

      Ngo, H.-V., Fell, J., & Staresina, B. (2020). Sleep spindles mediate hippocampal-neocortical coupling during long-duration ripples. Elife, 9, e57011.

      Ngo, H. V., Martinetz, T., Born, J., & Molle, M. (2013). Auditory closed-loop stimulation of the sleep slow oscillation enhances memory. Neuron, 78(3), 545-553.

      Picchioni, D., Horovitz, S. G., Fukunaga, M., Carr, W. S., Meltzer, J. A., Balkin, T. J., Duyn, J. H., & Braun, A. R. (2011). Infraslow EEG oscillations organize large-scale cortical–subcortical interactions during sleep: a combined EEG/fMRI study. Brain research, 1374, 63-72.

      Schabus, M., Dang-Vu, T. T., Albouy, G., Balteau, E., Boly, M., Carrier, J., Darsaud, A., Degueldre, C., Desseilles, M., & Gais, S. (2007). Hemodynamic cerebral correlates of sleep spindles during human non-rapid eye movement sleep. Proceedings of the National Academy of Sciences, 104(32), 13164-13169.

      Schreiner, T., Kaufmann, E., Noachtar, S., Mehrkens, J.-H., & Staudigl, T. (2022). The human thalamus orchestrates neocortical oscillations during NREM sleep. Nature Communications, 13(1), 5231.

      Schreiner, T., Petzka, M., Staudigl, T., & Staresina, B. P. (2021). Endogenous memory reactivation during sleep in humans is clocked by slow oscillation-spindle complexes. Nature Communications, 12(1), 3112.

      Staresina, B. P., Bergmann, T. O., Bonnefond, M., van der Meij, R., Jensen, O., Deuker, L., Elger, C. E., Axmacher, N., & Fell, J. (2015). Hierarchical nesting of slow oscillations, spindles and ripples in the human hippocampus during sleep. Nature Neuroscience, 18(11), 1679-1686.

      Staresina, B. P., Niediek, J., Borger, V., Surges, R., & Mormann, F. (2023). How coupled slow oscillations, spindles and ripples coordinate neuronal processing and communication during human sleep. Nature Neuroscience, 1-9.

      Vallat, R., & Walker, M. P. (2021). An open-source, high-performance tool for automated sleep staging. Elife, 10.

    1. Author response:

      The following is the authors’ response to the original reviews

      eLife Assessment

      This valuable study reports that the ALDH-abundant cells display stem cell properties and may play a key role in the endometrial epithelial development in the mouse. The data supporting the main conclusion are solid, although further improvements are needed to strengthen the conclusions. This work will be of great interest to reproductive biologists and biomedical researchers working on women's reproductive health.

      We thank the reviewers and editor for their critical reading and assessment of our manuscript. We carefully considered each of the points raised by the reviewers. In this document and in the edited manuscript and figures, we have carefully addressed each of the comments and requested modifications. In light of these changes, we expect that you will find that the manuscript has improved.

      We indicate our responses to the reviewers below in blue font and highlight the changes in the manuscript using the line numbers corresponding to the tracked version of the revised document.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript by Tang et al. characterizes the expression dynamics and functional roles of aldehyde dehydrogenase 1 activity in uterine physiology. Using a combination of in vivo lineage tracing and cell ablation coupled with organoid culture, the authors propose that Aldh1a1 lineage-marked cells contribute to uterine gland development and epithelial regeneration. The descriptive data will be of interest to reproductive biologists and clinicians and will build on established hypotheses in the field. The manuscript is well written and scientifically sound; however, several experimental limitations and interpretation caveats should be addressed.

      We thank the reviewer for their comments and expert assessment of our paper.

      (1) The methods surrounding the passage number and duration of culture following sorting prior to transcriptomic profiling should be clarified in the figure legends. Related to this, the representative images in Figures 1D and 1E do not appear consistent with the quantification presented in Figures 1F-H and should be reconciled.

      Thanks for this comment. We have now clarified this in the Figure 1 legend as follows,

      LINES 1026-1029: “Organoid formation assay performed immediately after luminal epithelial cell isolation and by plating equal numbers of viable ALDH<sup>LO</sup> (D) and ALDH<sup>HI</sup> (E) epithelial cells. ALDH<sup>LO</sup> and ALDH<sup>HI</sup> organoids were cultured for two weeks and passaged once prior to the organoid formation assays and transcriptomic analyses.”

      Regarding the second comment, we recognize that the images we showed may not have been the most representative of our quantification. As such, we replaced them with the organoid images so that they better reflect the quantification outlined in Figure 1F-H.

      (2) The conclusion that ALDH1A1+ cells are enriched in populations with stem cell characteristics relies primarily on transcriptomic analysis. Protein-level co-localization should be performed to strengthen this claim.

      We thank the reviewer for this comment. Unfortunately, the antibodies for many of these stem cell markers (such as LGR5, AXIN2, and SUSD2) are not well-suited for immunostaining. Others that have been proposed in human and are amenable to immunostaining are not suitable markers for mouse endometrial stem cells (such as CDH2). We hope that by showing that ALDH1A1 is expressed in patterns that are similar to the previously published stem cell markers LGR5 and AXIN2 (i.e., throughout the epithelium in the developing uterus and subsequently enriched in the tips of the endometrial glands of adult mice), along with transcriptomic studies, we can demonstrate its utility as a marker for mouse endometrial stem cells.

      (3) The overlap of 19 genes between the data set here and AXIN2 HI data is presented as evidence of shared stemness identity, but no statistical assessment of this overlap is provided. A hypergeometric test should be performed to determine whether this overlap is greater than expected by chance.

      Thank you for this suggestion. We have performed a hypergeometric test and determined that the reported shared genes between the two datasets are greater than is expected by chance. We have updated the results section to state the following:

      Lines 137-140: "We determined that the overlap between ALDH<sup>HI</sup> and Axin2<sup>+</sup> stemness marker genes was significantly greater than expected by chance for both upregulated (21/346 genes, 1.81-fold enrichment, p = 0.0067) and downregulated (19/674 genes, 1.67-fold enrichment, p = 0.021) gene sets (hypergeometric test, universe = 23,182 genes)."

      (4) The impact of tamoxifen injection on Aldh1a1 expression should be characterized in the neonatal uterus, as tamoxifen itself has known estrogenic activity that could confound interpretation of the lineage tracing results at early postnatal timepoints.

      Although we took measures to control for this possibility by using multiple time-points and models to trace the impact of Aldh1a1<sup>+</sup> cells in development and adulthood, we recognize the importance of this comment and acknowledge that this is a limitation in the design of our study. We have included the following text to the Discussion acknowledging this point:

      Lines 433-441: “Given the well-documented impacts of tamoxifen for lineage tracing studies, it is imperative to use doses of tamoxifen that will minimize estrogenic impacts and result in off-target effects (Rios et al., 2016). This often requires administration at doses that will achieve maximal recombination of the desired gene, while ensuring that the potential deleterious impacts of tamoxifen are minimized (Chen et al., 2023; Pimeisl et al., 2013). The cre/ERT2 tamoxifen inducible model is widely used to study uterine biology where it serves as a useful tool to interrogate the spatiotemporal impact of key genes, either through inactivation or for lineage tracing. Despite its widely documented utility across many tissue types and developmental timepoints, the use of tamoxifen and its impacts on the endometrium remain a limitation of our study, which we tried to address by implementing multiple timepoints, doses, and orthogonal assays in our experimental design.”

      (4b) Related to this, while low-dose tamoxifen is shown to label individual cells within 24 hours of injection, the translation dynamics of the label following Cre-mediated recombination can require up to 72 hours. The presence of only a few labeled clones at PND8 but multiple separate clones per cross-section at later timepoints warrants discussion and may reflect labeling kinetics rather than clonal expansion.

      The reviewer raises an important point. We agree that the 72hr-translation kinetics of the cre-mediated recombination is a legitimate consideration for interpreting our data and we have added the text below to the Discussion section acknowledging this point.

      We have addressed this by adding the following text to the discussion:

      Lines 417-422: We hypothesized that the singly labeled cells observed from one day tracing experiments expanded in a clonal fashion during the various timepoints we measured. We note that the translation kinetics of the labeled cells following cre-mediated recombination may contribute to the limited labeling observed at PND8/PND15 and there is a potential for delayed labeling of cells between 24 and 72 hours of tamoxifen administration. However, the continuous increase in labeled cells at the subsequent timepoints favors our interpretation of clonal expansion as the primary explanation.

      (5) It would strengthen the in vivo ablation data to validate the degree of cell death following diphtheria toxin treatment directly. It is possible that a general decrease in cell number rather than specific loss of a stem cell population is responsible for the observed reduction in gland number and FOXA2 expression (Tongtong et al 2017).

      We agree that this is an important control to incorporate into our experimental design. To rule out this possibility, we performed immunohistochemistry of cleaved caspase 3 in the uterine tissues of DTR<sup>flox/flox</sup> and DTR<sup>flox/flox</sup>;Aldh1a1<sup>cre/ERT2</sup> mice 4 days after administration of diphtheria toxin. The results indicate similar levels of cleaved caspase 3 detection in both genotypes, suggesting that the decrease in FOXA2+ cells is not due to non-specific cell death, but rather the result of ALDH1A1<sup>+</sup> cells. These data and the following text have been added to the manuscript:

      Lines 320-324: “We determined that the decreased in FOXA2<sup>+</sup> cells in the experimental mice was not the result of non-specific DT-mediated cell death, as similar levels of cleaved caspase 3-positive cells were detected in the DT-treated control ROSA26<sup>DTR/DTR</sup> and ROSA26<sup>DTR/DTR</sup>;Aldh1a1<sup>cre/ERT2/+</sup> mice 4 days post-diphtheria toxin administration (Figure S3G-H’).”

      (6) The lineage tracing data in the postpartum endometrium demonstrate that Aldh1a1-marked cells are present during regeneration, but it remains unclear whether these cells are preferentially activated or expanded in response to tissue injury. Coupling these studies with diphtheria toxin-mediated ablation during active regeneration would more directly test the proposed regenerative role of this population.

      This is a great point and one that we would be very interested in pursuing as follow-up studies in our future work. Regretfully, due to the long generation time and experimental procedures associated with these proposed studies, we are not able to include these experiments in the current manuscript. Thus, we have changed our wording and conclusions throughout the manuscript to be less definitive in terms of the role of Aldh1a1 in regeneration, since this will be the focus of future studies.

      The contribution of stromal Aldh1a1 lineage-positive cells is underexplored in the discussion, given the lineage tracing data showing stromal labeling across multiple timepoints and its potential relevance to mesenchymal-to-epithelial transition.

      Thank you for the suggestion. We have now expanded this section in the Discussion to include the following:

      Lines 496-504: We also found ALDH1A1<sup>+</sup> stromal cells were more prevalent when tracing began in adult mice. Other studies have shown that mesenchymal cells contribute to endometrial regeneration in the postpartum phase or after induced menses through a process of MET (Cousins et al., 2014; Kirkwood et al., 2022; Li et al., 2025). Similarly, lineage tracing studies have shown that MET is an active process and contributes to epithelial cell regeneration in the post-partum phase (Huang et al., 2012; Patterson et al., 2013). Although this is an area of active investigation in the field, with some contradicting reports, it is plausible to hypothesize that endometrial tissue has the capacity to undergo wound-healing and regeneration via several mechanisms (Ang et al., 2023; Ghosh et al., 2020). The process of MET in wound healing is widely documented in other organs, such as the kidney, liver and lung, where MET is associated with depletion of the resident epithelial cell pool (Bi et al., 2012; Niayesh-Mehr et al., 2024; Zeisberg et al., 2005).

      Finally, the word 'control' may overstate the functional evidence presented. 'Contribute' may be more accurate given the partial and context-dependent nature of the phenotypes observed.

      We agree with the reviewer’s point that control may overstate the evidence that we provide in the manuscript. To reflect this, we have edited the manuscript title and text to address this suggestion.

      Reviewer #2 (Public review):

      Tang et al. investigated the contribution of Aldh1a1+ cells, as putative stem/progenitor cells, to endometrial development, maintenance during the estrous cycle, and postpartum repair in mouse models. They employed in vitro organoid formation and in vivo lineage tracing models coupled with RNA-seq to test the stem-ness of Aldh1a1+ cells. They found that mouse endometrial cells with high ALDH activity (using the ALDEFLUOR assay) formed more and larger organoids and were enriched for stem/progenitor cell gene signatures. Similar results were shown using endometrial cells from a human patient sample. Epithelial ALDH1A1 expression was shown to be hormonally regulated, becoming more restricted to the glands, a putative epithelial stem cell niche, under estrogen stimulation. Using lineage-tracing initiated postnatally/prepubertally, Aldh1a1+ epithelial cells were shown to expand, contributing to both the luminal and glandular epithelium into adulthood, whereas adult initiation of labeling showed expansion of stromal Aldh1a1+ cells but not epithelial. Postnatal ablation of single-labeled Aldh1a1+ epithelial cells resulted in impaired gland development. Lastly, Aldh1a1-lineage traced cells (adult labeled) were present during postpartum endometrial repair as were epithelial/mesenchymal transitional cells.

      This study addresses an important area of research in the field of endometrial stem/progenitor cell biology. The authors are commended for their use of multiple complementary methods, including lineage tracing, DTR-mediated cell ablation, organoid assays, and RNA-seq in mouse and human models to assess the stem-like nature of Aldh1a1+ cells. The data support the stem/progenitor phenotype of Aldh1a1+ epithelial cells during endometrial development; however, there are noted discrepancies between organoid formation assays and lineage tracing experiments regarding the stemness of Aldh1a1+ epithelial cells in adults. Specifically, organoids were generated from adult cells and demonstrated in vitro stem cell activity; however, in vivo lineage-tracing of adult cells either during the estrous cycle or postpartum repair does not show expansion of Aldh1a1+ cells, suggesting they do not have stem/progenitor activity. Additionally, the stem-ness of epithelial vs stromal Aldh1a1+ cells is confounded in the study because epithelial cells were not purified for organoid experiments, epithelial cells were not exclusively lineage-traced as stromal cells were also labeled, and mesenchymal-epithelial transition was suggested to occur during postpartum repair. The following specific comments are presented to detail these concerns:

      We thank the reviewer for their critical reading of our manuscript and constructive comments.

      (1) The statement in the brief summary, "...critical for lifelong endometrial regeneration," is not supported by the data provided.

      We have edited the brief summary to exclude this statement, it now reads as follows:

      Lines 4-5: “We uncover ALDH1A1<sup>+</sup> cells as a group of hormone sensitive stem cells contributing to endometrial development and regeneration.”

      (2) AlDH1A1 is not restricted to the endometrial epithelium, and epithelial cells were not purified by flow cytometry for experiments in Figure 1. Figure 2 clearly shows the presence of mesenchymal cells, even using the described method for enriching for epithelial cells. Therefore, contaminating mesenchymal cells with high ALDH activity may confound the experimental results in Figure 1, either through promoting epithelial cell growth or through MET. The authors should provide clear evidence of epithelial purity in organoid experiments or that mesenchymal cells are not contained in the ALDHhi population. These comments also apply to the human organoid experiments in Figure 7.

      We thank the reviewer for raising this important point. Our group has been using the enzymatic method to routinely separate epithelial from stromal cell populations from the mouse uterus (see references dating back to 2015, PMID 26721398, 28324064, 34099644). In these experiments we typically obtain >98% purity in the epithelial and stromal cell compartments, respectively. We can directly observe this purity in the immunofluorescence images shown I Author response image 1 and Author response image 2, where mouse endometrial epithelial cells and stromal cells were enzymatically separated and immunostained with E-cadherin and vimentin antibodies to detect epithelial and mesenchymal cells in both cell preparations. The images show very few contaminating epithelial and stromal cells in either cell preparation. We have observed similar results when preparing epithelial and stromal cell preparation from the human endometrium, where the epithelial cell organoids display high purity with ~100% epithelial cell expression when we perform immunostaining.

      Author response image 1.

      Purity of mouse endometrial epithelial cells obtained via enzymatic and mechanical dissociation. A-B) Shows the epithelial (A) and stromal (B) cells plated on glass coverslips and immunostained with an epithelial cell marker (cytokeratin 8, red), a stromal cell marker (vimentin, green), and DAPI.

      Author response image 2.

      Human endometrial epithelial organoids were fixed and immunostained with cytokeratin 8 (green) and DAPI. The images are typical for our epithelial cell cultures and demonstrate that all epithelial cells are CK8-positive.

      (3) Lines 186-187: Susd2 was increased in EpSC clusters, yet this is a mesenchymal stem/progenitor marker in humans. The authors should discuss the implications of this.

      We thank the reviewer for highlighting this. We have now included the following in our Discussion to address this point:

      Lines 527-532: Clustering with this population of EpSCs were Susd2<sup>+</sup> cells, which are well-characterized mesenchymal progenitors that are enriched in the perivascular regions of the human endometrium (Darzi et al., 2016; Khanmohammadi et al., 2021). The presence of Susd2<sup>+</sup> cells, while unexpected in an epithelial stem cell niche, could indicate the presence of a transitional mesenchymal or perivascular cell that is differentiating into epithelium. Evidence for both mesenchymal and Nestin2<sup>+</sup> pericytes have been recently described in the mouse endometrial epithelium (Kirkwood et al., 2022; Li et al., 2025).

      (4) In Figure 5, RFP+ epithelial cells should be quantified as in previous figures to substantiate the statement in lines 279-280, "At PPD5, the proportion of RFP+ epithelial cells had expanded relative to PPD1 and PPD3 (Figure 5E-E')." Especially because in the low mag images (C-E), RFP+ epithelial cells appear to be most abundant at PPD1 and decrease at PPD3 and PPD5, suggesting that they may not be involved in endometrial regeneration/repair (contradicting the interpretation in line 285). Further, if there is in fact a decrease over postpartum repair, then regeneration should be removed from the title of the manuscript. RFP+ stromal cells should also be quantified.

      We appreciate this reviewer’s comment and agree that as stated, the conclusion is not fully supported by the data. To address this comment, we have edited the results so that they clearly indicate the results and remove any ambiguity:

      As requested, we quantified the number of RFP+ stromal and epithelial cells during the postpartum phase and noted that RFP+ cells were prominent in the stromal compartment of the endometrium. While RFP+ epithelial were also observed during these timepoints, they were less abundant than RFP+ stromal cells. Because the number of RFP+ cells did not significantly change over the postpartum phases in neither the stromal nor epithelial compartment, we have modified our conclusion to state that ALDH1A1+ cells are transiently detected in the regenerating endometrium.

      Results:

      Lines 287-294: “By analyzing the uterine tissues near the placental detachment site, we observed that RFP positive cells were prominent in the endometrial stromal cells that were adjacent to the luminal epithelium (Figure 5C-C’, green arrows). RFP<sup>+</sup> cells were also observed in the stromal cells near the placental detachment sites at PPD1 and PPD3 (Figure 5D’-E’, red & blue arrows) and in limited luminal epithelial cells (Figure 5D”,E”). Quantification of RFP+ cells throughout these postpartum phases indicated that stromal cells had more frequent ALDH1A1<sup>+</sup> stromal cells (360 ± 103, PPD1, n=3; 217 ± 107, PPD3, n=3; 254 ± 32, PPD5, n=4) than ALDH1A1<sup>+</sup> epithelial cells in the regenerating endometrium (65 ± 65, PPD1, n=3; 20 ± 10, PPD3, n=3; 114.25 ± 39, PPD5, n=4) (Figure S4).”

      Discussion:

      Lines 512-520: “We also noted that a majority of ALDH1A1<sup>+</sup> cells were localized to the active areas of endometrial regeneration near the placental detachment sites at PPD1 with a pronounced expression in the sub-epithelial stromal cells. As regeneration progressed, we continued to observe ALDH1A1<sup>+</sup> cells in the stromal compartment within the placental detachment sites at PPD3 and PPD5, with a progressive, but not statistically significant, increase in ALDH1A1<sup>+</sup> epithelial cells. Collectively, our data demonstrate that ALDH1A1<sup>+</sup> lineage cells participate in the restoration of endometrial architecture and functional compartments in the postpartum phase, even if their direct contribution is transient. Future detailed and mechanistic studies will be necessary to fully characterize their role in this process and their long-term consequence in postpartum regeneration.”

      (5) For Figure 7F, it should be clearly stated in the main text that the results are from one patient sample and the data presented are experimental replicates, so as not to be confused with biological replicates (the same for Supplementary Figure S4). Were B and G in Figure 7 also from one patient?

      Thanks for pointing this out. We have edited the figure legends in the main text and supplemental figures to indicate this.

      Lines 336-337: “…main figures show representative results from one patient sample performed in technical replicates, with additional patient samples included in the supplement…”

      (6) Lines 425-427: "Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed a complete restriction of ALDH1A1 to the glandular crypts." In Figure 2 S' ALDH1A1+ cells are visible in the LE (the staining is lighter than in the GE but looks real), contradicting this statement.

      This is an important distinction. We have now edited this part of the manuscript to state:

      Lines 458-461: “Ovariectomized mice treated with 90-day E2 pellets, on the other hand, showed enriched ALDH1A1 in the glandular crypts with weak luminal epithelial staining, while the ovariectomized controls had strong ALDH1A1 expression throughout the luminal and glandular epithelium.”

      (7) Lines 466-467: "In cycling mice, we found sporadic cells that expressed both stromal and epithelial markers in the ALDHA1+ cells." These data are not presented.

      We apologize for the confusion, this sentence has been removed from the discussion.

      (8) These data support the role of Aldh1a1+ cells in endometrial epithelial development, but conclusions about their role in repair/regeneration should be tempered as the data are much weaker here.

      We thank the reviewer for their overall assessment. To address this point, we have thoroughly edited the appropriate areas to temper the conclusions and ensure that they are strongly supported by our data. We have also edited the manuscript’s title to reflect this.

      Reviewer #3 (Public review):

      Summary:

      Tan et al demonstrated the importance of ALDH-high cells in the epithelial development in the mouse endometrium, and these cells displayed properties of stem cells.

      We thank the reviewer for their assessment of our manuscript.

      Strengths:

      The findings are solid, supported and validated through a combination of technical methods. I appreciated this combined use of mouse and human endometrial cells to strengthen the findings. Genomic results from a single-cell sequencing dataset were informative as they depicted the different stages of the estrus cycle during the regeneration process. Verification with immunostainings with various markers made it convincing for readers to visualize the cell's location, progression, and status at different timepoints. Utilizing human endometrial cells further demonstrated that the phenomenon observed in mice can be translated to humans.

      This work will greatly advance the understanding of endometrial regeneration for reproductive biologists.

      We thank the reviewer for their expert assessment and positive comments regarding our manuscript.

      Weaknesses:

      No major weaknesses were identified by this reviewer.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) As this study evaluated Aldh1a1+ cells in both the epithelium and stroma, it is recommended that the title and abstract be revised to reflect this.

      Thank you. Both the title and abstract title have been updated to reflect this comment.

      (2) Lines 46-47 in the abstract: "Aldh1a1+ cells expanded during postnatal development, estrus cycling, and following post-partum repair." It is recommended to clarify stromal vs epithelial Aldh1a1+ cell expansion, as only the stromal cells expanded during the estrous cycle. Also, RFP+ epithelial cells were not quantified during postpartum repair and visually appear to decrease (see comment below regarding Figure 5), so this statement is misleading.

      The abstract was edited following the suggestions so that it depicts our results. Similarly, we have addressed the comment regarding Figure 5 and the interpretation of the post-partum regeneration experiments (see comment above for the full explanation and edits to the manuscript).

      (3) Lines 65-69, 186-187: the authors should be clear when describing markers of putative epithelial vs mesenchymal stem/progenitor cells in the introduction (e.g., CDH2/SSEA1/SOX9 for epithelial and SUSD2 for mesenchymal).

      Thank you for this suggestion. The Introduction section now states the following:

      Lines 70-73: “These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).”

      (4) Lines 78-80: "Studies tracing the fate, ablation, and proliferative capacity of Lgr5+ cells in the uterus identified an Lgr5+ niche that is enriched in the crypts of the glandular epithelium and promotes endometrial regeneration (Seishima et al., 2019)." This statement is incorrect regarding regeneration, as Lgr5 marks stem/progenitor cells in the developing uterus but not the adult, during which endometrial regeneration occurs. Please revise.

      Thank for you for this clarification. The statement has been revised and now reads as follows:

      Lines 82-84: “Studies tracing the fate, ablation, and proliferative capacity of Lgr5<sup>+</sup> cells in the uterus identified an Lgr5<sup>+</sup> niche that marked stem/progenitor cells in the developing uterus (Seishima et al., 2019).”

      (5) Recommend using free-form shapes to outline GE, LE, and EpSC in Fig 2C and stating in the text which clusters correspond to LE and GE (lines 167-169).

      Thank you, this has been edited in Figure 2C and in the text, which now reads as follows:

      Lines 176-177: “Clusters 5, 7, 21 and 16 were classified as glandular epithelial cells, and clusters 0, 2, 3, 10, 14, and 24 were classified as luminal epithelial cells (Figure 2C).”

      (6) "Estrus" refers to the specific stage of the "estrous" cycle. Estrus cycle is incorrect and should be estrous cycle.

      Thank you for pointing this out. It has been corrected throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Suggest increasing the font size for some of the labels in the figures.

      Thank you, we have increased font size in the figures to improve the quality.

      (2) Need to include more details of the human endometrial tissue used in this study: pathology, age, and stage of menstrual cycle.

      We thank the reviewer for the helpful suggestion. The details that are available to us have been included in Supplementary Table S5.

      (3) Line 65-67 - Different endometrial stem cell subsets - SUSD2+ reside in perivascular regions, while other markers are located in glandular epithelium, need revision.

      Thank you, this has been revised in the Introduction. The area now reads as follows:

      Lines 70-73: These markers identify specific cell types in the endometrium, with CDH2<sup>+</sup>, SSEA1<sup>+</sup>, and SOX9<sup>+</sup> cells corresponding to endometrial epithelial cells, while SUSD2<sup>+</sup> cells corresponding to mesenchymal endometrial cells enriched in the perivascular regions (Cousins et al., 2021).

      (4) Include a discussion about the interpretation of their current finding in relation to the dynamic regeneration observed in human endometrium due to menstrual bleeding/tissue breakdown, compared to the cycles of growth and regression that occur in mice.

      This is a great suggestion. We have added the following statement to our Discussion section:

      Lines 537-544: Additionally, our studies in human endometrium extend our characterization of ALDH1A1 as an adult endometrial stem cell marker and emphasize the importance of ALDH1A1+ in the regenerative potential of the endometrium. The conserved hormonal responses between human and mouse endometrium support the hypothesis that cycles of proliferation, differentiation, and regression, whether through resorption/autophagy in mice or menstrual breakdown in humans, are governed by concerted growth factor signaling and dedicated stem cell populations with the capacity to expand and differentiate across repeated cycles of repair. Our detailed studies in both mouse and human models indicate that ALDH1A1+ cells represent a dedicated cell type within the endometrium with the potential to drive repair during adulthood. Collectively, these findings advance our understanding of the mechanisms that control endometrial cycling and regeneration throughout the reproductive lifespan.

      Reference

      Ang, C.J., Skokan, T.D., and McKinley, K.L. (2023). Mechanisms of Regeneration and Fibrosis in the Endometrium. Annu Rev Cell Dev Biol 39, 197-221.

      Bi, W.R., Jin, C.X., Xu, G.T., and Yang, C.Q. (2012). Bone morphogenetic protein-7 regulates Snail signaling in carbon tetrachloride-induced fibrosis in the rat liver. Exp Ther Med 4, 1022-1026.

      Chen, M.Y., Zhao, F.L., Chu, W.L., Bai, M.R., and Zhang, D.M. (2023). A review of tamoxifen administration regimen optimization for Cre/loxp system in mouse bone study. Biomed Pharmacother 165, 115045.

      Cousins, F.L., Murray, A., Esnal, A., Gibson, D.A., Critchley, H.O., and Saunders, P.T. (2014). Evidence from a mouse model that epithelial cell migration and mesenchymal-epithelial transition contribute to rapid restoration of uterine tissue integrity during menstruation. PLoS One 9, e86378.

      Cousins, F.L., Pandoy, R., Jin, S., and Gargett, C.E. (2021). The Elusive Endometrial Epithelial Stem/Progenitor Cells. Front Cell Dev Biol 9, 640319.

      Darzi, S., Werkmeister, J.A., Deane, J.A., and Gargett, C.E. (2016). Identification and Characterization of Human Endometrial Mesenchymal Stem/Stromal Cells and Their Potential for Cellular Therapy. Stem Cells Transl Med 5, 1127-1132.

      Ghosh, A., Syed, S.M., Kumar, M., Carpenter, T.J., Teixeira, J.M., Houairia, N., Negi, S., and Tanwar, P.S. (2020). In Vivo Cell Fate Tracing Provides No Evidence for Mesenchymal to Epithelial Transition in Adult Fallopian Tube and Uterus. Cell Rep 31, 107631.

      Huang, C.C., Orvis, G.D., Wang, Y., and Behringer, R.R. (2012). Stromal-to-epithelial transition during postpartum endometrial regeneration. PLoS One 7, e44285.

      Khanmohammadi, M., Mukherjee, S., Darzi, S., Paul, K., Werkmeister, J.A., Cousins, F.L., and Gargett, C.E. (2021). Identification and characterisation of maternal perivascular SUSD2(+) placental mesenchymal stem/stromal cells. Cell Tissue Res 385, 803-815.

      Kirkwood, P.M., Gibson, D.A., Shaw, I., Dobie, R., Kelepouri, O., Henderson, N.C., and Saunders, P.T.K. (2022). Single-cell RNA sequencing and lineage tracing confirm mesenchyme to epithelial transformation (MET) contributes to repair of the endometrium at menstruation. Elife 11.

      Li, S.Y., Whiteside, S., Li, B., Sun, X., and DeFalco, T. (2025). Mesenchymal-to-epithelial transition of perivascular cells contributes to endometrial re-epithelialization. Nat Commun 16, 10174.

      Niayesh-Mehr, R., Kalantar, M., Bontempi, G., Montaldo, C., Ebrahimi, S., Allameh, A., Babaei, G., Seif, F., and Strippoli, R. (2024). The role of epithelial-mesenchymal transition in pulmonary fibrosis: lessons from idiopathic pulmonary fibrosis and COVID-19. Cell Commun Signal 22, 542.

      Patterson, A.L., Zhang, L., Arango, N.A., Teixeira, J., and Pru, J.K. (2013). Mesenchymal-to-epithelial transition contributes to endometrial regeneration following natural and artificial decidualization. Stem Cells Dev 22, 964-974.

      Pimeisl, I.M., Tanriver, Y., Daza, R.A., Vauti, F., Hevner, R.F., Arnold, H.H., and Arnold, S.J. (2013). Generation and characterization of a tamoxifen-inducible Eomes(CreER) mouse line. Genesis 51, 725-733.

      Rios, A.C., Fu, N.Y., Cursons, J., Lindeman, G.J., and Visvader, J.E. (2016). The complexities and caveats of lineage tracing in the mammary gland. Breast Cancer Res 18, 116.

      Seishima, R., Leung, C., Yada, S., Murad, K.B.A., Tan, L.T., Hajamohideen, A., Tan, S.H., Itoh, H., Murakami, K., Ishida, Y., et al. (2019). Neonatal Wnt-dependent Lgr5 positive stem cells are essential for uterine gland development. Nat Commun 10, 5378.

      Zeisberg, M., Shah, A.A., and Kalluri, R. (2005). Bone morphogenic protein-7 induces mesenchymal to epithelial transition in adult renal fibroblasts and facilitates regeneration of injured kidney. J Biol Chem 280, 8094-8100.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This paper investigates how heparan sulfate (HS) engagement functions in the cellular entry of SARS-CoV-2. A prevailing model that has been developed over the last five years by work from many laboratories using a variety of biochemical, structural, and microscopic approaches is that HS acts a co-receptor for SARS-CoV-2; its binding to SARS-CoV-2 both concentrates virus on the surface of target cells and allosterically alters the spike protein to promote an "up/open" RBD conformation that enables engagement of the proteinaceous receptor human ACE2 on the cell surface (PMID: 32970989, 35926454, 38055954, 39401361, 40548749). These two events enable plasma membrane fusion (after a cleavage event promoted by plasma membrane TMPSS2) or endocytosis and subsequent pH-dependent fusion (which requires a cathepsin L-mediated cleavage of the spike).

      The authors in this study used a series of microscopy techniques, labeled pseudoviruses and authentic SARS-CoV-2 strains, and cells lacking or expressing HS and/or hACE2 to re-examine the specific stage(s) HS and hACE2 function in the entry process. They suggest that HS mediates SARS-CoV-2 cell-surface attachment and endocytosis, and that hACE2 functions "downstream" of this to facilitate productive infection. Their results also suggest that SARS-CoV-2 binds clusters of HS molecules projecting 60-410 nm, which act as docking sites for viral attachment. Blocking HS binding with pixantrone, a drug under clinical evaluation for cancer (due to its anti-topoisomerase II activity), inhibited SARS-CoV-2 Omicron JN.1 variant from attaching to and infecting human airway cells. The authors conclude that their work establishes a revised entry paradigm in which HS clusters mediate SARS-CoV-2 attachment and endocytosis, with ACE2 acting at some stage downstream. They speculate this idea might apply broadly to other viruses known to engage HS and has translational implications for developing antiviral agents that target HS interactions.

      The strengths of the interesting and technically well-executed study include the use of multiple high-resolution microscopy modalities, the tracking of labelled viruses, the use of both pseudoviruses and authentic SARS-CoV-2, and the use of primary airway cells. Nonetheless, there are issues that need to be addressed to buttress the proposed model compared to earlier ones. These include: (a) the distinction between macropinocytosis and receptor-mediated endocytosis and what this might mean for productive SARS-CoV-2 infection; (b) the need to account for TMPRSS2 expression and plasma membrane fusion; (c) addition of genetic studies in which hACE2 is expressed in cells lacking HS; (d) an unclear picture of exactly where downstream hACE2 functions; and (e) and a need for comparative/additional study of earlier SARS-CoV-2 variants, which preferentially fuse at the plasma membrane.

      We thank the reviewer for the strong support of this manuscript. We addressed the reviewer’s concerns in the Recommendations to the authors. We did not distinguish whether the endocytic route is macropinocytosis or receptor-mediated endocytosis, because it is a separate study beyond the scope of the present work. We did not examine earlier SARS-CoV-2 variants because we considered it a study beyond the scope of the present work, but a good idea that we may work on in the future. For detail on how we address the remaining concerns, please see our response to the reviewer’s Recommendations for the authors.

      Reviewer #2 (Public review):

      In this manuscript by Han et al, the authors assess the binding of SARS-CoV-2 to heparan sulfate clusters via advanced light microscopy of viral particles. The authors claim that the SARS-CoV-2 spike (in the context of pseudovirus and in authentic virus) engages heparan sulfate clusters on the cell surface, which then promotes endocytosis and subsequent infection. The finding that HSPGs are important for SARS-CoV-2 entry in some cell types is well-described, but the authors attempt to make the claim here that HS represents an alternative "receptor" and that HS engagement is far more important than the field appreciates. The data itself appears to be of appropriate quality and would be of interest to the field, but the overly generalized conclusions lack adequate experimental support. This significantly diminishes enthusiasm for this manuscript as written. The manuscript is imprecise and far overstates the actual findings shown by the data. Additional controls would be of great benefit.

      Further, it is this reviewer's opinion that the findings do not represent a novel paradigm as claimed. HS has been well described for SARS-CoV-2 and other viruses to serve as attachment factors to promote initial virus attachment. While the manuscript provides new insight into the details of this process, the manuscript attempts to oversell this finding by applying new words rather than new molecular details. The authors would be better served by presenting a more balanced and nuanced view of their interesting data. In this reviewer's opinion, the salesmanship significantly detracts from the data and manuscript.

      We thank the reviewer for pointing out that our manuscript is of interest to the field. However, we do not think that we oversell our data. hACE2 has been widely considered the receptor (or the binding partner) that mediates SARS-CoV-2 cell-surface attachment, whereas HS is considered only an attachment factor that facilitates SARS-CoV-2 binding with hACE2 at the cell surface. In the present work, we found that HS, but not hACE2, is the cell-surface attachment receptor (or binding partner), whereas hACE2 is not essential for attachment, but acts downstream of virus endocytosis to facilitate viral genome expression. This finding suggests significant modification of the current model by replacing the attachment receptor (or binding partner) from hACE2 to HS, treating HS as a primary receptor rather than an attachment factor, and relocating the hACE2 action site from the cell surface to the endosome. For these reasons, we do not consider these statements overselling our data. However, as the reviewer suggested in his/her specific comments, we revised the manuscript to ensure that we did not overgeneralize our findings (see our responses to the reviewer’s Recommendations to the authors).

      Major Comments:

      The authors need to rigorously define a "receptor" vs an "attachment factor." They also should avoid ambiguous terms such as "receptor underlying ...attachment" and "attachment receptor" (or at least clearly define them). Much of their argument hinges on the specific definition of these terms. This reviewer would argue that a receptor is a host factor that is necessary and sufficient for active promotion of viral entry (genome release into the cytoplasm), while an attachment factor is a host factor that enhances initial viral attachment/endocytosis but is neither necessary nor sufficient. The evidence does NOT implicate HS as a receptor under this fairly textbook definition. This is proven in Figure 1 (and elsewhere) in which ACE2 is absolutely required for viral entry.

      The authors should genetically perturb HS biosynthesis in their key assays to demonstrate necessity. HS biosynthesis genes have been shown to be important for SARS-CoV-2 entry into some cells but not others (Huh7.5 cells PMID 33306959, but not in Vero cells PMID 33147444, Calu3 cells 35879413, A549 cells 33574281, and others 36597481. The authors need to discuss this important information and reconcile it with their data and model if they want to claim that HS is broadly important.

      Is targeting HS really a compelling anti-viral strategy? The data show a ~5-fold reduction, which likely won't excite a drug company. The strengths and limitations of HS targeting should be presented in a more balanced discussion. Animal data showing anti-viral activity of PIX is warranted. This would enhance this claim and also provide key evidence of a relevant role for HS in a more physiologic model.

      The authors provide little discussion of the fact that these studies rely exclusively on cell lines (which also happen to be TMPRSS2-deficient). The role of proteases in the role of HS should be tested in the cell lines and primary cells used, as protease expression is a key determinant of the site of fusion.

      The claim that "SARS-CoV2 JN.1 variant binds to heparan sulfate, not hACE2, in primary human airway cells" is extraordinary and thus requires extraordinary evidence.

      First, PIX reduces attachment by 5-fold, which is not the same as "nearly abolished." Also, anti-ACE2 "nearly abolished" entry in 7D, while PIX did not. If the authors want to make these claims, an alternative method to disrupt HS (other than PIX) is needed in primary airway cells. A genetic approach would be much more convincing. The authors should also demonstrate whether entry in their primary cell assays is TMPRSS2 vs Cathepsin L dependent (using E64d and camostat, for instance) as mentioned above.

      Each figure should clearly state how many independent experiments and replicates per experiment were performed. What does "3 experiments" mean? Are these three independent experiments or three wells on one day?

      In the well-accepted current model, hACE2 is considered the receptor mediating SARS-CoV-2 cell-surface attachment, entry into cells, and infection, whereas HS is an attachment factor that facilitates SARS-CoV-2 binding to hACE2 at the cell surface. The present work revises this view: HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, with ACE2 acting downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined.

      We made this point clearer throughout the newly revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docks at HS clusters.

      The cited CRISPR-screen literature supports context-dependent host-factor usage. However, the absence of HS biosynthesis genes from a given screen does not prove that HS is irrelevant in that cell type; it only indicates that HS biosynthesis was not detected as a genetic dependency under that assay’s conditions. Such negative results can reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency. In the revised manuscript, we included the following in the Discussion:

      “While some studies using genome-wide CRISPR screening to identify genes involved in SARS-CoV-2 reveal genes for HS biosynthesis, others do not (45-50). The negative result, which might reflect screen sensitivity, incomplete knockout, pathway redundancy, or viral dose/stringency, needs to be verified with specific gene knockout.”

      The ~5-fold reduction is likely due to the inhibitor not completely abolishing HS-virus binding. We revised the Discussion to strengthen the suggestion that targeting the virus cell-surface attachment by interfering HS binding is a therapeutic strategy to prevent and treat COVID-19, as in the following:

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding. However, this strategy has not been the focus for developing methods to prevent and treat COVID-19, likely because HS is considered only a regulator that is not essential for SARS-COV-2 entry. Our finding that HS is the attachment receptor re-emphasizes the importance of perturbing virus-HS binding, the first step of the viral entry, to efficiently block SARS-CoV-2 infection. Further supporting this view, inhibition of HS binding with a clinically used HS-binding agent, pixantrone, inhibits authentic SARS-CoV-2 JN.1 subvariant binding with HS on the cell surface and infection in primary human airway cells (Figs. 6, 7). These results suggest a combinatorial anti-SARS-CoV-2 strategy: early HS blockade to prevent attachment combined with ACE2 targeting to inhibit post-attachment steps”

      We include a sentence in the Discussion that our suggestions are limited to the cells we examined as below.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      Three experiments refer to three independent experiments. We added “independent” accordingly throughout the manuscript.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors define a new paradigm for the attachment and endocytosis of SARS-CoV-2 in which cell surface heparan sulfate (HS) is the primary receptor, with ACE2 having a downstream role within endocytic vesicles. This has implications for the importance of targeting virion-HS interactions as a therapeutic strategy.

      Strengths:

      The authors show that viruses are internalized via dynamin-dependent endocytosis and that endocytic internalization is the major pathway for pseudotyped SARS-CoV-2 genome expression. They show that HS-mediated viral attachment is a critical step preceding viral endocytosis and also subsequent genome expression. Further, they show that hACE2 acts downstream of endocytosis to promote viral infection, and may be co-internalised with virions after HS attachment. Pseudotyped virus and authentic SARS-CoV-2 provide similar results. In addition, the authors demonstrate that remarkable clusters of multiple HS chains exist on the cell surface, visualised by a number of elegant microscopy methods, and that these represent the docking sites for virions. These visualisations are an important general contribution in themselves to understanding the nanoscale interactions of HS at the cell surface.

      The use of a complementary range of methods, virus constructs, and cell models is a strength, and the results clearly support the conclusions.

      Overall, the results convincingly demonstrate a different model to the currently accepted mechanism in which the ACE2 protein is regarded as the cell surface receptor for SARS-CoV-2. Here, the authors provide compelling evidence that cell surface clusters of HS are the primary docking site, with ACE2 interactions occurring later, after endocytosis (whilst still being essential for viral genome expression). This is an exciting and important landmark evidence which supports the view that HS-virion interactions should be viewed as a key site for anti-viral drug targeting, likely in strategies that also target the downstream ACE2-based mechanism of viral entry within endosomes.

      We thank the reviewer for the strong support of the present work.

      Weaknesses:

      This reviewer identified only minor points regarding citing and discussing other studies and typos, which can be corrected.

      We have addressed these points in the revised manuscript. For detail, please see our response to the reviewer’s Recommendations to the authors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Pathway of internalization.

      The authors show clearly that labeled SARS-CoV-2 (pseudovirus or authentic virus) can become internalized in cells lacking hACE2, and this process depends on HS. However, they also show that this pathway is non-productive with regard to infection. Are the entry vesicles mediated by HS alone, HS + hACE2, and hACE2 alone the same? Or does the combination of co-receptor (HS + hACE2) drive SARS-CoV-2 into endocytic vesicles, whereas HS alone promotes macro- or micropinocytosis (lines 361-362). If HS alone directed SARS-CoV-2 into a non-productive entry vesicle, then hACE2 likely would be acting concurrently with HS and not downstream. A more detailed analysis of the different entry vesicles/pathways that occur with HS alone, HS + hACE2, and hACE2 alone is needed.

      During endocytosis, we did not detect a difference in the size distribution of virus-containing vesicles between BHK (HS alone) and BHK<sub>hACE2</sub> cells (HS+hACE2) (Fig. 2D). The similarity in the vesicle size suggests a similar endocytic path with HS alone or with HS + hACE2. In the revised manuscript, we added the following sentence.

      “Third, 3D-STED imaging showed that A490-labeled vesicle’s full-width-at-half-maximum (W<sub>H</sub>) was 363 ± 17 nm (n = 55) in BHK cells, similar to that (333 ± 13 nm, n = 70) in BHK<sub>hACE2</sub> cells (Fig. 2B-D), supporting a similar endocytic path regardless of hACE2 presence or not.”

      (2) TMPRSS2 and plasma membrane fusion.

      Although the authors allude to membrane fusion as an alternate mechanism of entry, their mechanistic experiments do not address the roles of HS and hACE2 in this process, possibly because their BHK and other cells do not co-express significant levels of TMPRSS2. While many Omicron variants preferentially enter cells via endocytosis (relative to antecedent strains in the pandemic) because of spike mutations that reduce cleavage by TMPRSS2 (PMID: 35104837, 36625591, 35145066), plasma membrane fusion can still occur. The authors should add experiments with co-expression of TMPRSS2/hACE2 [with or without HS] and earlier SARS-CoV-2 variants to establish the role of HS in plasma membrane fusion. Also, are there differences in entry pathways if viruses are prepared in cells expressing TMPRSS2?

      We thank the reviewer for this important comment and agree that our mechanistic experiments were not designed to address TMPRSS2-supported plasma membrane fusion. The reviewer’s suggestion for direct testing of HS function in TMPRSS2-supported plasma membrane fusion, including hACE2/TMPRSS2 co-expression and comparison with earlier SARS-CoV-2 variants, will require dedicated experiments and detection of the fusion pathway that we have not yet designed. It is beyond the scope of the present work. In the revised manuscript, we clarify that the observed ACE2-independent uptake and the predominance of endocytic entry refer to the tested cell systems and do not exclude TMPRSS2-dependent plasma membrane fusion in other cell types, as in the following.

      “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Experiments with hACE2 in cells lacking HS.

      Apart from drug treatment (heparinases or pixantrone) studies shown, the current studies do not directly address whether expression of hACE2 on human cells can allow for endocytosis and productive infection in the complete and genetic absence of HS. The only experiments that use genetically deficient cells are the CHO [hamster] cell studies, and these cells lack hACE2 expression. The authors should knock out a key HS biosynthesis gene (e.g., B4GALT7) in more relevant human cells (e.g., A549-hACE2; ideally with sorted subpopulations having different levels of surface hACE2 expression) and assess endocytosis and infection. This is important given studies in the literature by others suggesting that KO of HS expression reduces but does not abrogate SARS-CoV-2 infection.

      We thank the reviewer for these comments. We showed that virus endocytosis is independent of hACE2 (Fig. 1). The reviewer’s question is whether hACE2 alone can allow for endocytosis of viruses. We have shown that in either BHK (without hACE2) or BHK<sub>hACE2</sub> cells (BHK cells expressed with hACE2), heparinase I/II/III mixture (HPRase) nearly abolished cell-surface immunolabelled HS (Fig. 3D), reduced cell-surface virus attachment by ~83-85% (Fig. 3E), reduced viral uptake by ~80% (Fig. 3F, 3G). These results suggest that hACE2 is not essential for viral attachment and endocytosis. We did not test whether hACE2 alone (without HS) plays a minor role for viral attachment and endocytosis, because to our knowledge, HS is present in nearly every cell. Under this physiological condition, it is HS, not hACE2, that plays an essential role in viral cell-surface attachment and endocytosis. In the revised manuscript, we added a sentence admitting that we did not test whether hACE2 alone is sufficient to support viral uptake and productive infection, as in the following.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis.

      (4) hACE2 function in entry.

      In many places, the authors suggest that hACE2-spike functional interaction occurs "downstream" of HS-dependent binding and endocytosis (e.g., lines 25, 33, 48, 210, 309, 312, 318, 333, 346). However, in their model, it is not clear where exactly this interaction occurs. Are the authors suggesting that this spike binds hACE2 on the cell surface, but this has nothing to do with endocytosis, or that the interaction with hACE2 is occurring at a post-entry step? Can they experimentally demonstrate the stage at which hACE2 is functioning? Is it the same or different in cells lacking HS? What about when TMPRSS2 is present?

      We showed that viral attachment and endocytosis are independent of hACE2, whereas entry as determined by viral gene expression, depends on hACE2. We also showed that most virions bind to HS, not hACE on the cell surface. Based on these results, we propose a model that hACE2 functions downstream of virion endocytosis. We cannot rule out the possibility that a small subset of viruses can also bind to hACE2 after their binding with HS at the cell surface.

      We have not been able to design an experiment to visualize hACE2 mediated virion fusion in endosomes, where hACE2 may facilitate virus fusion and delivery of viral genomes to the cytosol. Productive infection requires only a limited number of successful virion–hACE2 engagement events. While many internalized virions can be visualized, the specific virion or vesicle that ultimately gives rise to productive infection cannot be identified from the present imaging data. This makes it difficult to trace the precise stage or compartment in which the functionally relevant spike–hACE2 interaction occurs. In the revised manuscript, we added a paragraph discussing this limitation as below.

      “Our data suggest that, under physiological conditions in which HS is present at the cell surface, hACE2 is not essential for viral cell-surface attachment or endocytosis. We do not know whether hACE2 expression alone, in the absence of HS, can support viral cell-surface attachment and endocytosis. Although our data suggest that hACE2 functions downstream of endocytosis to facilitate viral fusion at the endosome for genome delivery to the cytosol, we do not know whether hACE2 binding with the virus occurs at the cell surface or endosomes. The binding may occur in both places, but not essential for virus attachment and endocytosis.”

      (5) Other comments.

      (a) Figure 1A. "Antibody" is misspelled.

      Corrected. Thank you.

      (b) The imaging experiments with pseudoviruses and authentic viruses lack any information on the multiplicity of infection or the number of virions added per cell. If this is particularly high and non-physiological (e.g., >100), is it possible that such conditions might enable viruses to enter [dominantly] through secondary [non-infectious] pathways?

      To address the reviewer’s concern, we used flow cytometry to measure cell-associated VSV-S signal as we diluted the virus by ~600-fold. We found that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHK<sub>hACE2</sub> cells over a ~600-fold dilution of the virus (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations. In the revised manuscript, we included the following sentence and Fig. S7 (Supplementary Information).

      “Flow cytometry also showed that the V-A647 attachment at the cell surface of BHK cells was similar to that in BHKhACE2 cells over a ~600-fold dilution of the virus concentration (Fig. S7), indicating that the virus cell-surface attachment is independent of hACE2 across a wide range of virus concentrations.”

      (c) Figure 1C and elsewhere. Most of the internalization studies rely on various imaging modalities to demonstrate the pseudovirus or virus on or in the cell. The experiments would be strengthened by inclusion of data from orthogonal binding/internalization assays that measuring virion-associated viral RNA on the surface [4oC binding assay] or inside the cell [after a 37oC temperature shift and exogenous proteinase K and RNAse A treatment]) - such assays can be performed at much lower MOI (e.g., <1, addressed comment #2 above) an also allow more objective quantitation and kinetic analyses of virus internalization (e.g., 0, 5, 15, 30 min at 37oC).

      We demonstrate virion attachment and uptake using multiple approaches, including confocal, STED, and EM analysis, showing virions with the expected morphology at the cell surface and in the cytosol. Furthermore, flow cytometric analysis provides population-level quantitation supporting the same overall conclusion. Thus, while we appreciate and agree that an RNA-based binding/internalization assay would provide additional information, we do not consider it essential to the main conclusion of this work.

      (d) Figure 2. (i) Is there any indication of which vesicles the bath dye is in? Is most of this fluid taken up by micropinocytosis? Are these the same vesicles where the virus that is destined for productive infection (HS/hACE2 engaging) transits? (ii) In all panels, can the authors clearly indicate/label which cells are being used (BHK or BHK-hACE2)? (iii) For the studies with dynasore or dominant-negative dynamin-2-K44A, the readout is at 24 h, a late timepoint, which also could affect virus egress and spread. Can the studies be repeated at much earlier time points (e.g., 15 min to 2 h) to demonstrate that viruses are internalized via dynamin-dependent endocytosis in these cells?

      (i) The bath dye A490 was used as a fluid-phase marker for endocytic uptake, rather than as a marker for a specific vesicle class or intracellular compartment. In principle, any vesicle that takes up extracellular fluid could become labelled by this approach. Since nearly all viruses are in the A490-containing vesicles, productive virus infection must come from some of these vesicles.

      (ii) In the revised Fig. 2 legends, we explicitly indicate which cells are used for each panel.

      (iii) To address the reviewer’s concern, we examined earlier time points for dynasore treatment and found that the virus uptake and genome expression were already markedly reduced at 1 h and 8 h after virus incubation. In the revised manuscript, we described these results as below and in Fig. S5.

      “Fourth, dynasore or dominant-negative dynamin 2-K44A overexpression, which inhibits fission of dynamin-dependent endocytosis [28-30], substantially reduced V-A647 internalized 1-24 h after viral incubation (Figs. 2F-G, S5).

      In addition to inhibiting V-A647 endocytosis, dynasore or dynamin 2-K44A inhibited V-EGFP expression 8-24 h after virus incubation by ~66-77% (Figs. 2F-G, S5), suggesting that endocytosis is the main route for viral genome expression.”

      (e) Line 225. "Envelop" should be "envelope".

      Corrected, thank you.

      (f) Line 235. The authors should clarify that they conclude that the "Omicron variant" of SARS-CoV-2 enters "BHK" cells indistinguishably from VSV-S.

      Thank you for pointing this out. We have rephrased the conclusion as “…omicron variant of SARS-CoV-2 enters BHK cells indistinguishably to VSV-S.”

      (g) Line 278. What happens to virus binding if the authors ectopically express hACE2 in CHO-K1 WT and CHO-pgsA-745 cells?

      We did not perform this experiment (see also our response to major comment 3 above).

      (h) Lines 280-281 and elsewhere (line 635). The authors state "pixantrone (PIX), a drug under clinical trial that binds HS to inhibit HS binding with proteins...." The authors should clarify that the drug is under clinical evaluation for cancer treatment because of its DNA intercalating activity (and not its HS binding activity) and cite any relevant ongoing trials. Also, in line 635, is reference #46 correct?

      As suggested, we modified this sentence as “pixantrone (PIX), a drug under clinical trial for cancer treatment due to its DNA intercalating activity, which can bind HS to inhibit HS binding with proteins”

      (i) Line 281-282. The authors should confirm in a Supplementary Figure that the anti-hACE2 antibody used blocks SARS-CoV-2-JN.1 binding to ACE2.

      In Figure 7D, we showed that PIX and anti-hACE2 antibody block SARS-CoV-2-JN.1 infection, suggesting that anti-hACE2 blocks SARS-CoV-2-JN.1 binding with hACE2.

      (j) Figure 7B. Can hACE2 co-localization be added to this panel?

      We did not perform this experiment. We addressed the role of ACE2 in these airway cells in subsequent panels of Fig. 7.

      (k) Figure 7C. The quantitative data show a 50% reduction in binding signal with pixantrone, whereas the microscopy images appear to show a much greater effect. Can more representative images be shown so that the data better corresponds?

      A ~50% effect is not as visually obvious as the current Fig. 7C. Therefore, we chose not to change the images. However, the statistics in Fig. 7C (right) clearly indicate an average effect of about 50%, as the reviewer pointed out.

      (l) In the Discussion, it is not necessary to use Figure callouts (as done in the Results). Please remove, with the exception of reference to the model.

      We prefer to call out Figures in the Discussion so that we can remind the readers where to find the data. The readers may choose to neglect these callouts. But some readers may read most the discussion part without going through the results carefully. In this case, the figure callouts may help these readers.

      (m) Please delete all references to "new" or "novel" models. It is unnecessary.

      As the reviewer suggested, we deleted “new” and “novel” throughout the revised manuscript.

      (n) Figure legends. Please make sure each panel indicates the # of independent experiments performed. This is included for some but not all panels. Also, a few panels use an unpaired t-test where an ANOVA with multiple comparisons is required (e.g., Figure 1G and S1).

      We agree that, for the three-group sub-comparisons shown within Fig. 1G and Fig. S1, the relevant analyses should account for multiple comparisons. In the revised manuscript, we therefore analyzed these predefined three-group subsets using ordinary one-way ANOVA followed by Dunnett’s multiple-comparisons test, with BHK or Vero used as the reference group as appropriate. The two-group comparisons were analyzed using unpaired two-tailed t-tests.

      Reviewer #2 (Recommendations for the authors):

      (1) It is well established that ACE2 is the receptor for SARS-CoV-2. The authors should not downplay this by saying it is "widely assumed", "typically thought", etc. The specific molecular details at various stages of entry (i.e, the role of HS) remain a bit unclear, but it is disingenuous to imply ACE2 is not the bona fide receptor by any conventional definition.

      The present work does not challenge the well-established view that ACE2 is the receptor for SARS-CoV-2 entry/infection, but suggests that HS is the SARS-CoV-2 attachment receptor mediating virus docking at the cell surface, whereas ACE2 acts downstream of virus endocytosis to enable SARS-CoV-2 infection in the cell types examined. We made this point clearer throughout the revised manuscript. We define the attachment receptor as the docking site where the virus binds to the cell surface. We directly showed with several super-resolution imaging techniques that the virus docked at the HS clusters.

      As the reviewer suggested, we removed “assumed” and “typical” and clarify that our findings do not challenge this concept. For example, we modified the abstract

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely assumed as the receptor.”

      as

      “Virus entry is thought to involve binding a unique receptor for cell attachment and cytosolic entry. For SARS-CoV-2 underlying the COVID-19 pandemic, angiotensin-converting enzyme 2 (ACE2) is widely considered the receptor for cell-surface attachment and subsequent cell entry.”

      (2) When the authors state pseudovirus internalization is independent of ACE2, they should clarify that this is the case in cells not expressing TMPRSS2. Most physiologically relevant cell types express TMPRSS2, which will facilitate entry at the plasma membrane.

      As the reviewer suggested, we included the following sentence in the Discussion section: “For other cells not examined in the present work, if TMPRSS2 is highly expressed, we could not rule out the possibility that the fusion pathway could also be dominant.”

      (3) Line 130: "Endocytic internalization is the main viral infection pathway" and Line 180-181 is not precise and should be rephrased to include the cell types described in the figure. This may be true in BHK-ACE2 cells, but the evidence in this section does not show that this is universally or broadly true.

      We agree and have revised these sentences to limit the conclusions to the experimental context directly supported by our data. Specifically, our results support endocytic uptake as the major route leading to pseudovirus genome expression in the pseudovirus assays and cell types examined here, rather than as a universal entry mechanism for SARS-CoV-2 across cell types. We have therefore modified the subsection title and the relevant sentence in the Results to explicitly refer to the tested cells/assays.

      Across the revised manuscript, we have accordingly revised the text to distinguish initial virion docking/attachment from productive entry, to acknowledge ACE2 as the established receptor for productive infection, and to limit our mechanistic conclusions to the cellular systems directly tested here.

      (4) All bar plots should show individual dots (i.e., Figure 1G) to better reveal the variance of each dataset.

      While we respect the reviewer’s suggestion, this is not required in the journal style. We prefer plotting bar graphs without individual data points, which often makes it difficult to see the mean values.

      (5) Line 57: This is not accurate. HIV uses a receptor and a co-receptor, for instance.

      We thank the reviewer for noting this inaccuracy. We agree that viral entry frequently involves coordinated engagement of multiple host factors rather than a single receptor, for example, HIV requires both a primary receptor and a co-receptor. We have revised the statement in the Introduction (Line 57–58) to reflect that entry can involve receptors together with co-receptors and/or attachment factors, which collectively facilitate membrane fusion or endocytic uptake.

      In the Introduction (Line 57), we replaced the sentence with “Viral entry is often initiated by engagement of host receptors and associated co-factors that together facilitate subsequent viral membrane penetration.”

      (6) Line 60: "most" --> "many"

      As suggested, we have changed “most” to “many”.

      (7) Remove "clinically relevant" in reference JN.1, as JN.1 is not circulating currently. A more appropriate term could be "full-length" or "authentic", or "wild-type".

      As suggested, we changed it to “authentic”.

      Reviewer #3 (Recommendations for the authors):

      (1) The authors omit to mention the work of Zhang et al, 2023 Nature Comms. "Host heparan sulfate promotes ACE2 super-cluster assembly and enhances SARS-CoV-2-associated syncytium formation". These authors also use PIXN and MTN compounds and define different mechanisms based on ACE2 clustering for virus entry. The authors should mention this work in the Discussion and try to reconcile the different findings.

      As suggested, we include the following discussion in the revised manuscript.

      “Consistent with this possibility, HS may promote spike-dependent ACE2 super-cluster assembly at the cell surface and enhance SARS-CoV-2–associated syncytium formation, suggesting that HS may organize ACE2 nanoscale architecture in a cell–cell fusion context [43].”

      (2) The authors should strengthen their case for the validity of HS-virion interactions as a therapeutic target by mentioning studies showing effectiveness of interference with HS-Covid interactions by heparin and other investigational drugs eg. first study to demonstrate heparin inhibition of SARS CoV2 attachment, Mycroft-West et al, Thromb Haemostatis, 2020; recent report of successful clinical trials of nebulized heparin, The Lancet, Sept 2025; and the superior efficacy of HS mimetic Pixatimod/PG545 compared to heparin (Guimond et al 2022 ACS Chemical Sciences).

      We thank the reviewer for this suggestion and add the following paragraph with citations the reviewer mentioned in the Discussion section.

      “Interfering with HS binding has been suggested as a therapeutic strategy to prevent and treat many viral infections that depend on HS for entry, including COVID-19 [1, 2, 9, 12]. Supporting this strategy, disrupting Spike–HS interactions, including inhibition by heparin and related glycans, reduces SARS-CoV-2 attachment/entry [51]. Clinical evaluation of inhaled/nebulized unfractionated heparin has reported improved clinical outcomes without major bleeding signals, supporting the feasibility of targeting airway-surface HS interactions [52]. HS mimetics, such as pixatimod (PG545), inhibit SARS-CoV-2 infection and exhibit greater potency than heparin in assays measuring inhibition of Spike/ACE2 engagement and viral infectivity [53]. These reports support the translational potential of therapeutically interfering with virion–HS binding.”

      (3) Figure 1a: incorrect label for antibody.

      Corrected, thank you.

      (4) Some misspellings noted in the manuscript, e.g., MINFLLUX, so please recheck the manuscript for typos.

      We have rechecked the manuscript and corrected the typos.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      There are two main criticisms:

      (1) It is not clear how much the factors uncovered here are true beyond B6 mice. B6 mice, compared to humans, are known to be very Th1-skewed, and Tbet is a strong inhibitor of Th17-specific T cells. Many people make IL-17-producing T cells in response to Mtb infection.

      We appreciate the point that not all findings in mice are directly translatable to humans. The B6 mouse is widely used as a model organism for tuberculosis due to its tractability and the wealth of genetic tools available for this strain. While it is true that many individuals do produce Th17 cells after infection with Mtb, humans are still very Th1-dominant, and not all infected individuals produce Th17 cells. We can speculate that the mechanisms outlined in this paper may contribute to the reasons that Th17 responses are not more robust in humans, a finding that may be useful in guiding vaccine design in the future.

      (2) Very few novel insights are mechanistically revealed about how Th17 induction is restricted by Mtb. Tbet induction is known to restrict Th17 development, and this is a T-cell intrinsic mechanism. In contrast, the IL-23 association revealed seems to be extrinsic to T cells and to act on T cells. How, if at all, are these factors related to each other in restricting Th17 induction? Also, the conclusion that it is not a result of attenuation is not completely convincing.

      While it is established that Th1 differentiation can inhibit Th17 differentiation, we believe that rigorously demonstrating this genetically in the context of Mtb infection remains important. Moreover, it cannot be assumed that IL-17 elicited by dampening the Th1 response can lead to enhanced control of infection. We view addressing this as a significant contribution. Furthermore, we also show that the ESX-1 and PDIM virulence factors are functionally linked by suppression of IL-17 responses. The effect is unlikely to be simply due to attenuation of the strains as an equally attenuated control strain does not elicit Th17 cells. We believe that these insights are both novel and important for understanding immune responses to Mtb.

      Other points:

      (1) The authors show that mice infected with a deficiency in ESX-1 have more IL-17-producing CD4 T cells in response to stimulation with an ESAT-6 peptide pool (Figure 3B). Because ESAT-6 is encoded by ESX-1, why do mice infected with this Mtb mutant have any ESAT-6-specific T cells? Is it an incomplete knockdown?

      The ESX-1 knock-out M. tuberculosis Erdman strain is a ΔEccC1 mutant. This strain can produce Esat-6 but cannot secrete Esat-6 out of the bacterial cell. Thus Esat-6 protein is present and able to be processed for MHC-II presentation. We also use Ag85b peptide pool stimulation and report similar effects as Esat-6 peptide pool stimulation.

      (2) The manuscript states, "Under the conditions where Th17s are highly induced, mice infected with either ΔESX-1 or PDIM lacking Mtb, the Il17a-/- mice had ~3-5 fold higher CFU than WT mice (Figures 3F-G). These results indicate that the induction of Th17s is not dependent on the attenuation of Mtb in general, but instead Mtb utilizes ESX-1 and PDIM to suppress the induction of a Th17 response that enhances protection against Mtb infection." I don't think the last sentence is necessarily true. I can imagine a scenario in which the induction of the Th17s is, in fact, due to the attenuation, and the Th17 induction still contributes to protection.

      We tested another attenuated M. tuberculosis strain with no known relationship with ESX-1 or PDIM, ΔMmpL4. This attenuated mutant fails to induce IL-17A–producing CD4 T cells to the same extent as observed in mice infected with ESX-1-deficient or PDIM-deficient strains, which is strong evidence that simple attenuation of virulence does not result in higher numbers of Th17 cells being elicited.

      (3) ESX-1, PDIM, and mmpl4 mutants all have similarly reduced CFUs in the lung, but what about the LN? The bacterial burden in the LN may be more important for regulating T-bet, IL-23, and Th17 differentiation, since the LN is where T cell priming occurs, than the CFU in the lung. Perhaps ESX-1 and PDIM mutants have reduced CFU in the LN, but mmpl4 does not. This difference in LN burdens may be the primary driver of Th17 priming, as high avidity interactions are thought to be an important driver of T-bet induction.

      We acknowledge that this is a formal possibility, however we maintain that the phenotype is specific to ESX and PDIM mutants, rather than MmpL4 mutants. Even if this phenotype arises from a tissue-specific attenuation of ESX/PDIM mutants, it remains a specific phenotype of these mutants, and not all attenuated mutants, albeit less directly. More importantly, the observation that these mutants induce higher levels of the Th17-polarizing cytokine IL-23 from infected cells ex vivo suggests that this is not an indirect phenomenon.

      (4) Do LN cDC1 and high levels of IL-12 p35 in mice infected with the mmpl4 mutant? Likewise, LN cDC2's express low levels of IL-12 p19 (akin to those infected with WT Mtb)? If these observations for ESX-1 and PDIM mutants are mechanistically linked to the increased numbers of Th17 cells, then you would expect mice infected with mmpl4 mutants to be more like those infected with WT Mtb than those infected with ESX-1 and PDIM mutants.

      Because ΔMmpL4 and complemented strains resulted in T cell profiles that were not different from the wild-type, we did not measure mediastinal lymph node dendritic cell expression of IL-12 p35 and IL-23 p19 in infections with these mutants.

      (5) ESX-1 and PDIM are very different virulence factors - a protein secretory pathway and cell wall lipid, respectively? Mechanistically, how would mutants in these pathways give very similar outcomes regarding Th17 cells unless it was simply as an aspect of their attenuation? Perhaps, mmpl4 mutants simply differ in some aspects of their attenuation, such as bacterial burdens in LNs, or their interaction with cDCs?

      We are not the first to link phenotypes of ESX-1 and PDIM. Both systems have both been shown to be important for M. tuberculosis permeabilization of the host cell phagosome after phagocytosis, and for suppression of type I IFN responses, among other responses. Thus, these seemingly different virulence factors clearly work together to support specific virulence traits during infection. The exact mechanism of how ESX-1 and PDIM interact is not completely understood and is an area for future investigation.

      Reviewer #2 (Public review):

      The following conclusions and interpretations should be revisited, rephrased, and re-evaluated:

      (1) The manuscript neglects to analyze T cell responses in the dLN, which is the critical site where these responses are initiated (only DC cytokine production is measured in the dLN). The differences in the lungs could reflect trafficking of T cells to the lungs, local lung T cell responses, or durability of the T cell responses in the lungs. The authors state in the last results section that "These results indicate that the ESX-1 and PDIM virulence factors impact naïve T cell differentiation at the draining mediastinal lymph node..." but T cell responses are never measured in the dLN.

      Due to the limited size of the mediastinal lymph node at 3 weeks post infection, we were unable to obtain enough cells for both myeloid cell analysis and T cell analysis, as we perform staining for these panels separately due to the decrease in viability of myeloid cells observed during T cell restimulation. In addition, because T cells in the lung are the population of cells most critical for mediating the outcome of infection, we believe analyzing the T cell response in the lymph nodes though interesting, is not crucial for this study. We have edited the manuscript to be clearer, as suggested by the reviewer.

      (2) Figure 2: The authors state that "Importantly, IFN-γ deficient mice did not exhibit elevated levels of IL-17A producing CD4 T cells demonstrating that IFN-γ production is not the mechanism by which Th1 T cells limit a Th17 response during Mtb infection", but the difference is significantly different and even more obvious in Panel B. In fact, if the Panel D y-axis was on a log scale, the Ifng-/- would likely look more like Tbet-/- than WT. Based on this data, it seems like IFNg is having an effect and should not be completely discounted. Does the deletion of Ifng affect the number of Tbet+ T cells?

      We agree that the IFN-γ<sup>-/-</sup> have only 5x more IL-17 producing CD4 T cells than WT mice while Tbet<sup>-/-</sup>mice exhibit a 25-fold increase compared to WT. We have added this information to the text, and now point out that IFN-γ production is not the sole mechanism by which Th1 T cells limit a Th17 response during Mtb infection.

      In addition, the deletion of Tbet results in an increased number of IFNg+IL-17+ double positive T cells (Figure 2B), in addition to a sizable IFNg single positive T cell population maintained in the Tbet-/- mice (10x the negative control of Ifng-/-). Is this why Tbet deletion is not as severe as Ifng deletion, because T cells are still making IFNg?

      It is possible that the residual IFN-γ produced by T-bet-deficient animals contributes to their relatively modest susceptibility to infection. However, our data show that deletion of IL-17 in this background renders T-bet–deficient mice nearly as susceptible as IFN-γ deficient mice, arguing that the remaining IFN-γ is not a major protective factor.

      Along these lines, the statement in the text that, "Tbet-/-Il17a-/- mice completely lacked both IFN-γ producing...." T cells is not supported by the data in Figure 2C. Tbet-/-Il17a-/- mice look to have more gamma-producing T cells than Tbet-/- mice (which is already 10x the negative control of Ifng-/- in panel 2B if one includes the gamma single positive and IFNg/IL-17 double positive).

      We have amended the language in the text to be more consistent with the data.

      (3) In the Results sections describing Figures 3, 4, and 5, the authors equate IL-17 production by T cells with TH17 responses and IFNg expression with TH1, but Tbet and RORgt expression in the T cells should be measured to make conclusions about TH1 and TH17. Or the authors can rephrase their findings to specifically state the observations as IFNg or IL-17 expressing CD4+ T cells.

      We believe that calling a CD4 T cell in the lung that is producing IFN-γ (and not IL-17) a Th1 cell is appropriate. Potentially confounding cells include those which also produce IL17, which we have ruled out, or T<sub>FH</sub> cells that may be common in lymph nodes but are not common in lungs at this time point and under these conditions.

      (4) Conceptually, do the authors think that ESX1/PDIM promotes TH1 responses and this blocks TH17 or are ESX1/PDIM blocking TH17 responses directly, allowing for increased TH1 responses? It would be helpful to clarify the model in this regard, describe how the data supports one model or the other, and then make sure the language is consistent throughout. Can these effects on T cell responses be tested and recapitulated in vitro using infected APC and T cell co-cultures?

      While it is possible that PDIM and ESAT-6 suppress Th17 through promotion of Th1 differentiation, we do not have data to support this model currently. However, we have added a comment making this point to the discussion.

      Reviewer #3 (Public review):

      Weaknesses:

      (1) The authors should acknowledge and reference key findings from the literature that have identified suppression of Th17 differentiation as an Mtb virulence mechanism, e.g., the role of the Hip1 protease and CD40 signaling (Madan-Lala JI 2014, Sia Plos Path 2017, Enriquez iScience 2022) and Khader JI 2005, showing the requirement of IL-23 for Th17 responses in vivo in a TB mouse model.

      We thank the reviewer for pointing these references out and have added them to the discussion section of the manuscript.

      (2) Addressing several questions related to the Tbet KO mouse experiments would strengthen the study. Do the Tbet KO mice have elevated IL-4/5/13 (which has been previously reported in non-TB studies) in addition to IL-17? The lack of Th17 cells in the IFNg KO compared to the Tbet KO may be due to a difference in timing, since only 3-week data are shown; earlier and later time points would provide better interpretation. The authors do not present any data on neutrophil infiltration in WT vs Tbet KO vs IFNg KO mice. Since IL-17 is known to be important for recruiting neutrophils to the lung, data on neutrophils are important for clarifying the mechanism for the CFU outcomes.

      We agree that it is surprising that, in the context of TB, Th17 responses are protective whereas excessive neutrophil recruitment is detrimental to the host. In IFN-γ–deficient mice, neutrophils are recruited and contribute to the increased susceptibility of this strain (PMID: 21967766). In separate work from our lab, we have shown that the phenotype of neutrophils recruited to the lungs during Mtb infection influences disease outcome (PMID: 40937719). It is possible that differences in the host environment and the timing of the response shape the effects of neutrophils on the host; these and the other questions raised by the reviewer will be the subject of future studies.

      (3) While IL-23 is important for sustaining IL-17 production, IL-6, TGF-b and/or IL-1β are necessary for Th17 polarization. What were the levels of these cytokines in DCs in the lung? (Figure 5). Additionally, Tbet-deficient DCs exhibit impaired activation of antigen-specific Th1 cells and have reduced IL-12 production. Given the data showing higher IL-17 levels in Tbet KO mice, the authors should provide information on the DC phenotype (IL-23, IL-6, etc.) in the Tbet KO experiments.

      While these are interesting points, investigating mechanisms of Tbet-dependent suppression of IL-17 is beyond the scope of this study.

      (4) The mechanism by which ESX-1/PDIM function to impact Th17 differentiation is not clear. While data showing a role for ESX-1 and PDIMs in inhibiting Th17 responses is interesting, there is no insight into the potential mechanism of action. Figure 3 showing reduction in IFNg+ CD4 T cells after infection with eccC1 and fadD28 mutants suggests that this outcome is due to a lower bacterial load relative to WT Mtb at the 3-week time point. Since IFNg is known to suppress IL-17, the higher levels of Th17 cells could be due to the reduction in IFNg due to the attenuated growth of the mutants. Additionally, what was the level of Type I IFNs elicited by these mutants?

      We included the MmpL4 knockout Mtb Erdman strain as a control to ensure that attenuation of mutants is not the cause of the increase in IL-17. We also showed that eliminating type I IFN signaling by deleting its receptor has minimal impact on Th17 differentiation, even in the context of a host that produces excess type I IFN. Therefore we do not believe that type I IFN elicited by these mutants is explanatory for the phenotype.

      (5) Since macrophages have been implicated in the reduced cytokines seen in the ESX-1 mutant, IL-23 and other cytokine data on lung macrophages would complement the DC data.

      Because dendritic cells are primarily responsible for priming CD4 T cell responses, we believe that this result in macrophages would not substantially alter our conclusions. That said, it was demonstrated previously that macrophages infected with ESX-1 mutants produce less IL-12p40, a subunit of IL-23.

      (6) Figure 5. There are many fewer DCs overall in the eccC1 and fadD28 mutant groups, which could account for the increased % IL-23p19 in DCs (5D). What were the levels of IL-23 in DC1s?

      The amount of IL-23 p19+ in type I conventional dendritic cells (cDC1s) was near zero as shown in supplementary figure 6A. cDC1s are known to not express IL-23 p19 in mice.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) What do the authors mean by the alternative secretion part of "ESX1 Type VII alternative secretion system" that they refer to?

      Bacterial alternative secretion systems facilitate the export of proteins from the bacterial cell independent of the canonical Sec-dependent secretion system required for export of most secreted bacterial proteins across the inner membrane. However to avoid confusion, we have removed the word alternative.

      (2) Not sure naïve fits in this sentence at the end of the introduction: "Furthermore, we observe a strong Th17 response during infection with ΔESX-1 or PDIM lacking Mtb in naïve mice....".

      We have removed the word Naïve.

      (3) Figure legend for 1A-C says analysis performed at 21 dpi, but the figure shows the time course.

      We have corrected this error.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 1 should show the non-stimulated flow plot.

      We have added the unstimulated samples.

      (2) The % IL-17 in the flow plots is not consistent across Figures 1, 2, and 3. Not sure why the scales for the Y-axis for IL-17 differ so much between Figures 1 and 2/3. IS there a technical issue with compensation?

      We did not experience any difficulties with compensations. These experiments were done over several years of work. For every experiment, new single-color controls were used and gating was done with the FMO gating strategy. Minor variation such as we see here is not surprising.

      (3) Discuss Yeh et al J Neuroimmunol 2014- show that IFNγ inhibits Th17 differentiation and function via Tbet-dependent and Tbet-independent mechanisms.

      We have added this reference to the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Strengths:

      The strength of this study is its use of a simple behavioral parameter, TOWA, and also a simple design of behavior, WAFO. The importance of the behavioral assay is reproducibility and comparability. In fact, the author demonstrated a summary of comparisons where different treatments result in scalable behavioral changes in WAFO and TOWA.

      We appreciate this assessment and fully agree that the simplicity of the assay and the demonstration of its scalability and reproducibility are a strength.

      Weaknesses:

      The weakness of the study is the lack of further experiments to support their assumption related to TOWA. The authors suggested that TOWA can be interpreted as a behavioral proxy for exogenously induced arousal. However, it could be interpreted as higher activity, although the authors argued that the circadian clock increasing locomotor activity around ZT0 and ZT12 does not affect TOWA, and therefore TOWA is not related to the locomotor activity per se. As the author cited, flies lose locomotor activity in the circular arena of 6.6 cm in diameter, whereas they continuously move during a 1-h recording in the authors' arena of 1 cm in diameter.

      I would agree that the arena of 1 cm in diameter, but not 6.6 cm in diameter, serves as an exogenous stimulus inducing arousal, and TOWA is manifested by arousal. However, TOWA would also be affected by other behavioral parameters, including the activity, motivation for exploration, or perception of the space. Therefore, it could be reasonable to re-examine some of the flies tested in this study in the circular arena of 6.6 cm in diameter. If arousal is biased by the components presented in Figure 6 and TOWA can assess mainly exogenously induced arousal, the treatment altering TOWA in the arena of 1 cm in diameter would not affect their behavior in the arena of 6.6 cm in diameter. My concern is that Figure 6 may demonstrate too simplistic a diagram to interpret the results. I would suggest adding the experiments using the arena of 6.6 cm diameter or softening the argument.

      We are grateful that you prompted us to investigate the relation between TOWA and arousal and different arena diameters in more depth. Based on your comments, we compared naïve and stressed behaviour between arena diameters of 1, 2.2 and 5.8 cm. The sizes were chosen as to optimally comply with the camera field of view in our setup. Naïve flies showed stable locomotor activity throughout the 60 min of recordings in the different arenas (new Figure 1 – figure supplement 1 A-B’’). Moreover, no significant difference in TOWA over the first 10 min was found between the different arena sizes (new Figure 1 – figure supplement 1 C). This suggests to us that arenas with a diameter smaller than 6.0 cm (and not only with a 1 cm diameter) induce some form of activity that resembles stimulated activity as defined by Meehan and Wilson (1987). A mechanical shock (shake) resulted in significantly increased WAFO in all arena sizes (new Figure 1 – figure supplement 1 D-D’’). We did not test the effects in larger arenas > 5.8 cm, as we found that flies do not longer show persistent and quantifiable wall following. We started to see that also in few flies in the 5.8 cm arena – these few flies were excluded from our analysis presented in new Figure 1 – figure supplement 1. Unlike the naïve response, the stress-induced response in WAFO and TOWA appears to be transient (new Figure 1 – figure supplement 2), which is in line with the definition of emotions as a transient state.

      Reviewer #2 (Public review):

      Summary:

      Strengths:

      The main strength of the paper is the rigorous use of several stressful or aversive treatments and their subsequent removal to show that WAFO is a robust proxy for stresslike emotional primitives across multiple stimuli. The pharmacological, molecular, and neuronal activity manipulations, although more limited in scope, lend further credence to the authors' central claim.

      We are glad about this assessment and share your opinion.

      Weaknesses:

      The conceptual advance of this research is unclear, as previous work (Mohammad et al., 2016, Curr Biol.) carried out similar treatments and manipulations and reached largely similar conclusions.

      Thank you very much for bringing this up. We rewrote respective parts of the introduction (second last paragraph) and discussion to more clearly outline the advances over the previous work by Mohammad et al. 2016. While our study builds upon Mohammad et al. 2016, the conceptual advance and novelty is that we constitute and treat TOWA as a second and independent dimension equal to WAFO in the OFT. Mohammad et al. had measured locomotor activity (reported as average speed in their paper (total distance walked/time of recording), but primarily to test the dependency of WAFO on locomotor activity. They found that WAFO metrics were poorly correlated with average walking speed, showing a significant degree of independence of both measures – a finding that our results confirm. However, unlike us, they did not consider average speed/TOWA further for their analysis, possibly because they focused on anxiety-like behaviour while our study looked broader on emotion-like behaviour in general. We further used a round (not square) arena to exclude “cornering” in order to reduce the complexity of the assay, which may explain differences of observed speed/TOWA between our studies.

      Moreover, while WAFO is a good proxy for 'stress', I am not convinced that TOWA necessarily represents an emotional state in all cases. Indeed, as the authors themselves acknowledge, changes in total walking may be associated with other factors, such as starvation-induced hyperactivity, physical exhaustion after sleep deprivation, increased sex drive after mating, alcohol sedation, etc.

      Your comment raises a question in comparative research on emotions which is very difficult if not impossible to conclusively answer. At first sight, the most conservative stance seems to be to completely disregard the idea of emotions and affective experiences in animals. This, however, would mean that we cannot use animal models to study the basics of emotions (= emotion primitives) and would ignore that by all likelihood emotions are a product of evolution and hence should exist at least in more basic forms in animals. Obviously, we have no means to ask flies or any other animal whether they connect “hunger” or “mating” to a feeling or an emotional state (which must not be conscious) but can only observe the behaviour. We here adopt the often-cited “Pankseppian” view (based on the book of Jaak Panksepp: “Affective Neuroscience”) and firmly believe that – in order to fully understand how the brain drives behaviour- we also need to take affective states into account that bias behaviour towards adaptive responses.

      In short, we are unfortunately unable to give a clear and definite answer to your comment whether TOWA represents an emotional state in all cases. Perhaps you are right. We believe, however, that a “Pankseppian” view is adequate, and we may ask what evidence exists that shows that starvation-induced hyperactivity or post-mating is not associated with an affective emotion-like state in the fly or any other animal.

      Another unclear point is the interpretation of some unexpected results, such as the finding that both serotonin transporter overexpression and its knockdown give the same phenotype.

      Thank you very much for this comment, which we also received by reviewer #1. As suggested by the other reviewer, a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT may be a differential effect on the different serotonin receptor subtypes expressed in the brain. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – both higher and lower than optimum levels result in behavioural impairments in adults (see e.g., Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Finally, there are some issues with the use of the OFT in rodent research (e.g., inconsistent effects of anxiolytic drugs; see Rosso et al., 2022, Neurosci Biobehav Rev., for a meta-analysis). These should be explained to place the Drosophila findings in their appropriate context.

      Thank you very much for bringing this systematic review to our attention which assessed the usefulness of various behavioural tests including the OFT to study the effect of anxiolytic drugs in rodents. Overall, the review casts “serious doubt on both construct and predictive validity” of behavioural tests for anxiolytics. While diazepam (the only drug used in our study) was the drug with the most consistent effects across the analysed behavioural assays, only 59% of the OFTs revealed significant effects. We were already aware of earlier findings in the same direction (Prut and Belzung 2003 10.1016/s0014-2999(03)01272-x), but as we only used one drug did not include a discussion in the manuscript. Unfortunately, the number of studies employing the OFT in flies is very small and does not yet allow for a similar comparison. We now changed the respective sentences in the discussion:

      “In rodents, diazepam mostly but not consistently leads to an anxiolytic response in the OFT behaviour which questions the usefulness of the OFT for testing anxiolytic drugs (see (Prut and Belzung 2003; Rosso et al. 2022)).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Overexpression of SerT suppressed increased WAFO after electric shocks. However, knockdown of SerT did similar. It is worth reporting this finding, but possible mechanisms could be mentioned in the manuscript. For example, given that flies carry five serotonin receptor genes, upregulation of overall serotonin level may affect the specific serotonin receptor, but downregulation of it may affect other receptors, leading to unexpected outcomes. Exploring any other possibilities could be better to add to guide future research.

      Thank you very much for this comment and the suggestion of a reasonable mechanism that may underly the similarity of effect after knockdown or overexpression of SerT. The possibility that the concentration-dependent effect of a biogenic amine follows a U-shape is further reasonable and has been demonstrated for dopamine. For example, the relationship between cognitive performance or working memory and dopamine levels in primates follows an inverted U-shape (see e.g. Cools and D’Esposito 2011 10.1016/j.biopsych.2011.03.028, Desimone 1995 10.1038/376549a0). Also in Drosophila, both reduced and increased dopamine levels lead to increased male-to-male courtship behaviour (Liu et al. 2008 10.1523/JNEUROSCI.5290-07.2008, Liu et al. 2009 10.1371/journal.pone.0004574). While we are unaware of similar examples for serotonin, we note that serotonin levels must be kept at optimum level during development – higher or lower levels result in behavioural impairments in adults (see e.g. Shah et al. 2018 10.3389/fnbeh.2018.00114). We have now extended the discussion accordingly.

      Reviewer #2 (Recommendations for the authors):

      (1) The advance over Mohammad et al., 2016, Curr Biol. must be clearly and emphatically articulated in the Introduction and/or Discussion. It is otherwise impossible to appreciate what is conceptually novel about this work.

      We have now rewritten parts of the two last paragraphs in the introduction to make the advances clearer. As outlined above, we conceptually advanced the analysis of OFT behaviour by integrating TOWA as a second and independent dimension in our analysis. During the revision process, we have spent great effort to better characterise the nature of the locomotor activity encountered in the OFT (Figure 1 – supplementary figures 1 and 2). Also, this is an advancement over previous studies, including Mohammad et al. 2016. We further included new results on the general effect of neuropeptides (silver mutants, impaired in neuropeptide processing).

      (2) The most important metric for a stress-like emotional primitive is WAFO. Therefore, figures should be revised in a way that highlights WAFO differences. Figure 4 is a good example of this. In contrast, in Figures 1-3, the WAFO box-and-whisker plots are very small and obscured under the raw WAFO-TOWA plots, which are difficult to see (especially given the light blue background of all plots) and redundant. I strongly recommend just showing the WAFO box-and-whisker plots for the sake of visibility, clarity, and brevity.

      Although we understand your reasoning, we would like to stick with the old figures as we consider TOWA as important to characterise the OFT response as WAFO (see our comments above regarding the conceptual advances of our study over Mohammad et al. 2016). It is true that Figures 1-3 are small, but at least in our print-out well legible. Further, it appears that in the current version of the eLife system the resolution is downsampled. In addition, we anticipate that in the version of record figures can be enlarged online as in other eLife articles.

      (3) Given the criticisms against the OFT in rodent research, one of which is inconsistent effects of anxiolytic drugs (Rosso et al., 2022, Neurosci Biobehav Rev.), it may be useful to expand pharmacological treatments beyond diazepam.

      As our focus is not on the testing of anxiolytic drugs and since it was already very difficult to be granted access to diazepam (we are not at a medical institution), we refrained from testing further drugs. Moreover, as rightfully mentioned by you, the OFT may not be the best test for the efficacy of anxiolytic drugs. On the other hand, diazepam was the most consistent anxiolytic in the OFT in rodents (see Rosso et al. 2022).

      (4) The authors should show results of the effects of at least some stressors/punishments on WAFO/TOWA of female flies to understand if observed effects are sex-specific.

      Thank you very much for bringing this topic to our attention. To test whether the effects are sex-specific, we now performed several new experiments. First, we compared the naïve OFT response of mated and unmated males and females (see new Figure 6). This revealed that without prior stress treatment, the WAFO response is independent of sex and mating status. In contrast, the naïve TOWA response turned up to be sex- and mating state-specific (see new Figure 6). To test whether the mated females show a different stress-induced OFT response to males, we applied mechanical stress (shake) that we had also used to assess the effects of arena diameter (Figure 1 – supplementary Fig. 1 D-D’’). After a first round of shaking, females showed increased TOWA, but WAFO was unaffected. A second round of shaking, however, led to a significant increase in WAFO and TOWA. This suggests that the OFT response is qualitatively similar between the sexes and mating status, yet the threshold for elicited responses differs between males and females. We now added a respective paragraph to the main text in the results section plus a new figure (Fig. 6).

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an interesting study that addresses whether mitochondrial DNA (mitoDNA) variants impact telomere length (TL), which may be relevant to potential maternal inheritance of TL in offspring. The study addresses this question using a cybrid model approach in which mitochondria from donor platelets from 7 individuals that vary in TL and differ in mitoDNA variants are introduced into 143B cells that lack mitochondria. MitoDNA variants that exhibited reduced complex I activity showed telomere shortening in cybrids and increased telomere dysfunction. Interestingly, these phenotypes could be reduced with NAC antioxidant and NAD+ supplementation, suggesting that ROS and oxidative DNA damage at telomeres contributed to the telomere shortening. They further showed that cybrids with lower levels of ROS correlated with longer TL in the lymphocytes of the mitochondrial donors.

      Strengths:

      This study provides compelling evidence that mtDNA variants influence TL through a mechanism involving mitochondrial-derived ROS, potentially causing telomeric oxidative damage. The data are robust, and the manuscript is well written. However, the study could be strengthened by addressing the following questions and minor weaknesses below.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Introduction. Line 81, the relationship between TL and the risk of lymphoid and myeloid leukemia is not straightforward. POT1 variants associated with long TL increase the risk for lymphoid and myeloproliferative neoplasms (see PMID: 41564438 for example).

      We appreciate the reviewer raising the important link between pathogenic POT1 variants and lymphoid malignancies driven by elongated telomeres. We would like to clarify, however, that our introduction focused not on the pathological telomere attrition characteristic of telomere biology disorders, but rather on the natural, non-pathological variations observed in individuals with baseline telomere lengths on the shorter end of the spectrum.

      Nevertheless, we will include this observation regarding patients with long telomeres in the introduction to underscore the complex relationship between telomere length and tumorigenesis.

      The two following publications by the Armanios lab will be added:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      (2) Figure 1. Since sex also influences TL, it would be good to know the sex of the selected individuals or explain why this is not necessary.

      Because this study investigated the potential maternal inheritance of TL via the mitochondrial genome, our cohort consisted predominantly of female donors (6/7). Donor #5, the husband of Donor #7, was the only male included. We will update Figure 1D to include the sex of each donor and clarify this rationale in the text.

      Notably, our analysis revealed minimal influence of sex on TL, which cannot account for the observed differences between the extreme groups.

      (3) Please include a description of the 143B cells that were used for cybrid formation in the Results section when introducing the cybrids.

      We will do so.

      (4) Lines 155-156. The authors note that cybrids from donors 1 and 2 show "pronounced" telomere damage. This result indicates an increase in 53BP1-positive telomeres, which could be indicative of telomere dysfunction or damage. Quantification of the increased chromosome end fusions for cybrids 1 and 2 would strengthen the result.

      We thank the reviewer for this suggestion. Following their advice, we used our metaphase spread FISH analyses to quantify chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells. However, because Cybrid 2 metaphase spreads were of insufficient quality for adequate chromosome analysis, we will restrict our telomere fusion comments exclusively to Cybrid 1 and include the quantifications in our revised manuscript.

      Author response image 1.

      The number of fusions and total chromosomes analyzed are indicated on each bar.

      Do the increased fusions correlate with an increase in telomere signal-free ends? These should be apparent in the telomere FISH images of metaphase chromosomes.

      While this is a strong argument, the telomeres within these cybrid models are critically short. Consequently, the FISH signal intensity falls below the threshold required for reliable quantification of telomere-free ends.

      (5) Lines 168-169. What is the evidence that the "in vitro metabolic shift" causes acute oxidative stress?

      The reviewer is correct that we have not formally demonstrated this in our experimental system. Instead, our assumption was based on the fact that the initial phase of Rho0 cell repopulation involves a temporary ROS burst, partly driven by incompletely assembled ETC supercomplexes that are known to elevate ROS levels (Maranzana et al, 2013). This is supported by our experimental observation that the NAC antioxidant, combined with NR, potently inhibits telomere shortening in cybrids with low CI activity. We propose to add this explanation in the revised manuscript.

      (6) Why did the elevated ROS in cybrid #3 (Figure 4C) not translate to shorter telomeres in the cybrid (Figure 2A)? Perhaps there is a difference between factors that determine TL in the cybrid vs the donor's lymphocytes?

      We cannot fully explain the discrepancy, but we indeed suspect in vivo oxidative stress differs significantly from cell cultures (21% O<sub>2</sub>). Additionally, early cybrid replenishment involves an unknown telomere elongation step that may offset mitochondrial ROS-induced shortening.

      To further investigate this question, we measured telomere length across varying population doublings (PDs) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include (as Supplementary figure) and discuss these data in the revised manuscript.

      Author response image 2.

      In Figure 4B, it appears that the statistical comparisons for mitochondrial superoxide are all relative to Cyb3. If so, why are the comparisons not with the parental 143B rho0 cell line? Please clarify.

      We excluded Rho0 cells from our superoxide measurements because these cells were grown in a different culture medium. Instead, we focused on comparing mitochondrial ROS across cybrids to accurately correlate these values with the telomere length of the corresponding donors’ lymphocytes. Consequently, Rho0 cell measurements would not have contributed to this correlation analysis.

      We propose to discuss this in the revised manuscript.

      (7) Given the heterogeneity in TL and mtDNA variants in the human population, the conclusions could be further strengthened by increasing the number of donors and cybrids analyzed. However, there are admittedly practical factors. Overall, these findings are compelling and provide a solid foundation for expanding this analysis in the future. This is more of a comment than a weakness.

      We thank the reviewer for this constructive feedback and entirely agree that including more donors would have added depth to our findings. While we acknowledge this limitation, pursuing this further is currently impossible without acquiring new ethical approvals and establishing fresh collaborations with clinicians.

      Reviewer #2 (Public review):

      Summary:

      The authors aim to determine whether mitochondrial genotype influences telomere length. By generating cybrids harboring different mitochondrial backgrounds, the authors seek to establish a mechanistic link between mitochondrial status and telomere biology.

      Strengths:

      A major strength of the study is the use of cybrid technology, which provides a great approach to investigate the role of mitochondrial DNA independently of the nuclear genome. The authors also employ multiple complementary assays to assess telomere-related phenotypes associated with mitochondrial dysfunction. Together, these experiments generate an interesting dataset that will be of value to researchers interested in the intersection between mitochondrial biology, genome stability, aging, and development. These results also build on previous work supporting roles for ROS/mitochondria in driving telomere shortening.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      The data support the conclusion that mitochondrial background is associated with differences in telomere length and telomere-related phenotypes. However, some of the mechanistic interpretations would benefit from additional evidence. In particular, the manuscript discusses mitochondrial influences on telomere shortening, yet telomere length in some experiments is assessed at a single time point. Consequently, the current data do not directly address the rate of telomere attrition. Differences observed between cybrid lines could potentially arise from events occurring during cybrid formation, clonal selection, or subsequent cell expansion. Longitudinal analyses across multiple passages, ideally beginning immediately after cybrid generation and controlling for population doublings, would help establish whether mitochondrial function directly affects telomere shortening dynamics. Some experimental results would also benefit from additional quantification, clarification, and some biological replicates are missing.

      We limited our TL measurements to the earliest viable time point after cybrid formation to avoid the confounding effects of cellular adaptation in culture. For instance, the early telomere shortening observed in Cybrid 1 and Cybrid 2 was later alleviated—likely due to the upregulation of the NAD+ salvage pathway genes NAMPT and NAPRT1.

      To accurately capture the effects specific to cybrid formation, we isolated 4 independent clones per donor, all of which showed highly consistent TL values, as shown in Figure 2A and S3D.

      While we acknowledge the reviewer's point about multi-passage longitudinal analyses, we did, in fact, measure telomere length across varying population doublings (PDs) in Cybrid 3 (high mito ROS levels), 6 and 7 (low mito ROS levels) and observed the following shortening after 66-67 PDs:

      - Cybrid 3: about 1.6 kb reduction

      - Cybrid 6: about 1.2 kb reduction

      - Cybrid 7: about 0.8 kb reduction

      These results suggest that the rate of telomere shortening in culture may be higher in Cybrid 3 cells, possibly due to increased ROS levels.

      We propose to include a Supplementary figure and discuss these data in the revised manuscript.

      We further propose to carefully edit the manuscript so as to clarify the text and, whenever possible, add quantifications. Among others, as suggested by Reviewer #1, we will add the quantification of chromosome end fusions in the parental 143B Rho0, Cybrid 1 and Cybrid 6 cells in our revised manuscript (See Author response image 1).

      Overall, this study provides interesting evidence linking mitochondrial background to telomere biology. The cybrid models represent a useful resource for the field, and the work raises important questions regarding mitochondria-telomere communication.

      Reviewer #3 (Public review):

      Strengths:

      Mahieu and colleagues address an interesting and underexplored question: whether non-pathogenic variation in the mitochondrial genome contributes to the inter-individual variability of human telomere length (TL). Using a Belgian Flow-FISH reference cohort (n=491) to identify donors at TL extremes, they generate transmitochondrial cybrids from platelets of seven donors of distinct mtDNA subhaplogroups and characterize the resulting cells with a broad and well-executed toolkit (TRF, TeSLA, ddTRAP, EPR-based mitoROS, Seahorse with permeabilized-cell ETC dissection, LC-MS metabolomics, telomeric PAR-FISH). The most compelling finding is that cybrids derived from donors with low complex I (CI) activity undergo rapid telomere shortening during the glycolysis-to-OXPHOS transition of cybrid formation, and that this is largely prevented by co-treatment with NAC and the NAD⁺ precursor nicotinamide riboside, supporting a model in which CI sustains the NAD⁺ pool required for PARP1-mediated repair of oxidative damage at telomeres. The authors further report an inverse correlation between donor lymphocyte TL and mitoROS in the corresponding cybrids, and provide preliminary evidence that the K1a-defining ATP6 A177T variant (m.G9055>A) may be enriched in long-telomere individuals.

      We thank the reviewer for this very positive evaluation of our work.

      Weaknesses:

      (1) Statistical support and donor sampling for the central in vivo correlation (Figure 4C).

      The inverse correlation between donor lymphocyte TL and cybrid mitoROS (R<sup>2</sup>=0.794, p=0.007) is the principal in vivo claim of the paper, but it is built on seven donors deliberately selected from the extremes of the Flow-FISH distribution. Sampling at the tails of the outcome variable can substantially inflate apparent correlation strength and significance. I would encourage the authors to (i) explicitly state this sampling structure where the correlation is introduced, (ii) report a leave-one-out sensitivity analysis to confirm the relationship is not driven by one or two donors (Cyb3 and Cyb6 appear to anchor the line), and (iii) where feasible, extend the analysis to additional donors with intermediate TL to test whether the relationship holds across the full distribution. Even a modest expansion (e.g., 4 to 5 additional donors at P25 to P75) would substantially strengthen this central claim.

      We thank the reviewer for this insightful comment. As suggested, we will explicitly describe the sampling structure upon introducing the correlation analysis. Furthermore, we have conducted the requested leave-one-out analysis:

      - removing Cyb3: R<sup>2</sup>=0.730; p=0.0302

      - removing Cyb6: R<sup>2</sup>=0.782; p=0.0194

      - removing Cyb1: R<sup>2</sup>=0.740; p=0.0278

      While we agree that additional donors would enhance the study, further experiments are however currently impossible without new ethical clearances and additional clinical partnerships.

      (2) Reconciling the cybrid CI / TL relationship (Fig 3B) with the absence of a CI / TL relationship in donor lymphocytes (Figure 4A).

      Figure 3B shows a strong correlation between CI activity and TL in cybrids (R<sup>2</sup>=0.87), while Figure 4A shows no correlation between donor CI activity (measured in the same cybrids) and donor lymphocyte TL. The authors acknowledge this, but the manuscript subsequently builds toward a CI-centric model of in vivo TL regulation, which seems to outrun the data. The most internally consistent interpretation is that the cybrid CI phenotype reports a sensitized in vitro response to the acute oxidative stress of the metabolic shift, rather than a steady-state determinant of leukocyte TL. I would suggest reframing the abstract, significance statement, and Discussion to make this distinction clearer. The in vitro CI / NAD⁺ / PARP1 axis is a strong finding on its own, while the in vivo role of CI activity (as opposed to ROS more broadly) is not yet established here. Donor #1's profile (very long lymphocyte TL, low CI activity, severe shortening in cybrids, no telomere inheritance in offspring) is informative in this regard and could be discussed more directly as a case that helps delineate where the cybrid model does and does not recapitulate in vivo biology.

      We acknowledge that our study does not establish the in vivo role of CI activity in TL regulation. Our abstract specifically highlights an in vitro phenomenon: “Under the specific conditions of cybrid formation, which involve a metabolic shift from glycolysis to oxidative phosphorylation, mtDNA variants associated with reduced CI activity induced rapid telomere shortening, …”.

      We are nevertheless happy to revise the text to clearly separate our in vitro results from in vivo biology as requested.

      (3) The K1a / ATP6 A177T inheritance claim.

      The proposal that K1a (and specifically ATP6 A177T) contributes to maternal inheritance of long telomeres is intriguing but currently rests on three pedigrees (one of which, donor #1, does not support the hypothesis) and a chi-square test that does not reach significance (p=0.153, Figure 4F). The supporting evidence is also limited by the fact that platelet-mediated mitochondrial transfer delivers donor mitochondrial proteins, lipids, and residual mtRNA in addition to mtDNA, making it difficult to attribute the cybrid phenotype of donor #6 specifically to the ATP6 A177T variant. I would recommend either: (a) extending the genotyping screen to additional unrelated donors and, if feasible, confirming the effect of ATP6 A177T through an isogenic approach (e.g., mtDNA base editing in a clean background), or (b) softening the relevant statements to "suggestive trend warranting larger studies," and presenting the K1a observation as hypothesis-generating rather than supportive. The Ashkenazi-centenarian connection raised in the Discussion is an excellent direction for follow-up and could be framed accordingly.

      We agree that the evidence for the AT6 A177T inheritance claim remains inconclusive. To clarify, we do not argue that this mitochondrial variant is solely responsible for longer telomeres; indeed, the mtDNA genome of donor #1 suggests otherwise. Furthermore, the phenotypic impact of such mtDNA variants likely depends on nuclear variants in other telomere-related genes (e.g., hTERT or hTR), meaning AT6 A177T may not consistently result in elongated telomeres. Unfortunately, our ethical protocol precludes screening additional unrelated donors. We will revise the text to soften our statements accordingly.

      New references:

      DeBoy EA et al. Familial clonal hematopoiesis in a long telomere syndrome. 2023. N Engl J Med 388, 2422-2433.

      Davidson-Swinton HR et al. Lymphoid malignancy and clonality in the POT1-mediated long telomere syndrome. 2026. Blood 147, 2226-2237.

      Maranzana E et al. Mitochondrial respiratory supercomplex association limits production of reactive oxygen species from complex I. 2013. Antioxid Redox Signal 19, 1469-1480.

    1. Author response:

      We thank the reviewers for their assessment of our work and their comments. We are grateful for their evaluation of our findings as fundamental and convincingly supported, and for their appreciation of the relative scope of this manuscript and of future work. The most direct requests for new experimental data are from reviewer #2, who asks for direct assessment of the effects of CHOP deletion on expression of GADD34 and on protein synthesis. We agree that these are important experiments to conduct for the revision.

      The reviewers requested more clarity on the experimental logic of the paper and on the place of our findings in the broader context of ER stress signaling, which we will be happy to provide in a revised manuscript. These revisions will include a more explicit consideration of how the regulation of metabolic genes by CHOP contributes to its effects in the liver independently of its role in regulating eIF2a dephosphorylation.

      In particular, there were concerns about the logic of the time points chosen that we feel are important to also address here. For analysis of ChopHKO animals, all experiments were carried out 8 hours after ER stress challenge. This is because, as we show in Fig. 1B and also in our previous paper on CHOP (1), this is the time point at which CHOP expression is at its maximum. Thereafter, hepatocytes become heterogeneous with respect to whether they do or do not express CHOP. This is an interesting finding because it suggests that CHOP is part of a cellular switch, and potentially even an effector of that switch—a point currently raised in the Discussion but worth further highlighting in a revision. At the practical level, it means that discerning the contribution of CHOP to ER stress signaling and adaptation at subsequent time points will require sophisticated single cell analyses that can discriminate cells that express CHOP from cells that do not, which are an important future direction.

      In contrast, for Atf6aHKO animals, all experiments were carried out 48 hours after ER stress challenge. As we have previously shown (2), at short time points after a stress challenge, such as 8 hours, there is very little difference in ER stress signaling between wild-type animals and those lacking ATF6a. The reason for this lack of distinction is that the major targets of ATF6a are ER chaperones and the like. Because adaptation to ER stress in the early phases of the response depends more on non-transcriptional mechanisms such as inhibition of protein synthesis and IRE1-dependent mRNA decay (RIDD), the failure to fully upregulate ATF6a targets is initially of little consequence. It is only at later time points when wild-type animals restore ER homeostasis and largely silence ER stress signaling. In contrast, at these same later points, animals lacking ATF6a show evidence of persistent ER stress, most notably in the form of persistent Xbp1 mRNA splicing and profound suppression of metabolic genes. Although the 8 hour time point for experiments in ChopHKO animals differs from the 48 hour time point for Atf6aHKO animals, the two lines of experimentation are united by the persistence of ongoing ER stress and of ISR signaling despite diminished eIF2a phosphorylation at the points when the presence of CHOP or the absence of ATF6a are of the greatest impact. A revised manuscript will present this logic more clearly.

      References

      (1)  Liu K, et al., EMBO Reports 25, 228 (2024)

      (2)  Rutkowski DT, et al., Dev. Cell 15, 829 (2008)

    1. Author response:

      We would like to express our deepest gratitude to the Editors and Reviewers for their highly rigorous and constructive evaluation of our manuscript. We are greatly encouraged by the recognition of our study’s ambition, the unique value of the in vivo intrathecal contrast MRI dataset, and the conceptual novelty of linking macroscopic glymphatic physiology with neural activity and regional proteopathy.

      We fully agree with the thoughtful limitations and methodological concerns raised in the eLife Assessment and the Public Reviews. In our upcoming revised manuscript, we are implementing a comprehensive set of revisions to address these points. Specifically, our planned revisions focus on the following key areas:

      - Tempering Causal Interpretations: We agree that our cross-sectional design precludes definitive causal inferences. We are systematically revising the manuscript to soften causal language (e.g., replacing "drives" with "is spatially associated with"). We will explicitly frame our findings as macroscopic spatial associations and discuss the potential influence of joint physiological confounders.

      - Tightening Terminology and Imaging Physics: We are refining our terminology to more accurately reflect our MRI measurements. We will replace assertive terms like "direct glymphatic flow" with precise descriptors such as "imaging proxies for tracer enhancement and retention." Furthermore, we are expanding the Limitations section to explicitly acknowledge the confounding effects of Partial Volume Averaging (PVE), systemic tracer redistribution, and renal clearance kinetics.

      - Conducting Supplementary Imaging & Robustness Analyses: To address concerns regarding cohort heterogeneity and the sample size of the rs-fMRI subgroup (n=15), we are performing a series of rigorous supplementary analyses. This includes conducting sensitivity analyses (e.g., excluding the motor neuron disease subgroup) and applying leave-one-out cross-validation to rigorously assess the subject-level stability and robustness of the spatial coupling between neural activity and tracer clearance.

      - Clarifying the Conceptual Model and "Mismatch" Index: To improve readability, we are moving the anatomical definitions of the cortical gradients directly into the Results section. Additionally, we are introducing schematic diagram to intuitively explain the mathematical formulation and biological interpretation of the "activity-clearance mismatch" index.

      - Re-framing External Dataset Analyses: We are carefully re-framing the interpretations of the Allen Human Brain Atlas (AHBA) transcriptomic data and the external PiB-PET amyloid dataset, emphasizing that these reflect spatial correspondences of intrinsic regional vulnerability across groups, rather than individual-level direct interactions.

      We believe these revisions will significantly enhance the scientific rigor, clarity, and precision of our study.

    1. Author response:

      The following is the authors’ response to the original reviews

      Public Reviews:

      Reviewer #1 (Public Review):

      This work provides a new dataset of 71,688 images of different ape species across a variety of environmental and behavioral conditions, along with pose annotations per image. The authors demonstrate the value of their dataset by training pose estimation networks (HRNet-W48) on both their own dataset and other primate datasets (OpenMonkeyPose for monkeys, COCO for humans), ultimately showing that the model trained on their dataset had the best performance (performance measured by PCK and AUC). In addition to their ablation studies where they train pose estimation models with either specific species removed or a certain percentage of the images removed, they provide solid evidence that their large, specialized dataset is uniquely positioned to aid in the task of pose estimation for ape species.

      The diversity and size of the dataset make it particularly useful, as it covers a wide range of ape species and poses, making it particularly suitable for training off-the-shelf pose estimation networks or for contributing to the training of a large foundational pose estimation model. In conjunction with new tools focused on extracting behavioral dynamics from pose, this dataset can be especially useful in understanding the basis of ape behaviors using pose.

      We thank the reviewer for the kind comments.

      Since the dataset provided is the first large, public dataset of its kind exclusively for ape species, more details should be provided on how the data were annotated, as well as summaries of the dataset statistics. In addition, the authors should provide the full list of hyperparameters for each model that was used for evaluation (e.g., mmpose config files, textual descriptions of augmentation/optimization parameters).

      We have added more details on the annotation process and have included the list of instructions sent to the annotators. We have also included mmpose configs with the code provided. The following files include the relevant details:

      File including the list of instructions sent to the annotators:

      OpenMonkeyWild Photograph Rubric.pdf

      Mmpose configs:

      i) TopDownOAPDataset.py

      ii) animal_oap_dataset.py

      iii) init.py

      iv) hrnet_w48_oap_256x192_full.py

      Anaconda environment files:

      i) OpenApePose.yml

      ii) requirements.txt

      Overall this work is a terrific contribution to the field and is likely to have a significant impact on both computer vision and animal behavior.

      Strengths:

      Open source dataset with excellent annotations on the format, as well as example code provided for working with it.

      Properties of the dataset are mostly well described.

      Comparison to pose estimation models trained on humans vs monkeys, finding that models trained on human data generalized better to apes than the ones trained on monkeys, in accordance with phylogenetic similarity. This provides evidence for an important consideration in the field: how well can we expect pose estimation models to generalize to new species when using data from closely or distantly related ones?

      Sample efficiency experiments reflect an important property of pose estimation systems, which indicates how much data would be necessary to generate similar datasets in other species, as well as how much data may be required for fine-tuning these types of models (also characterized via ablation experiments where some species are left out).

      The sample efficiency experiments also reveal important insights about scaling properties of different model architectures, finding that HRNet saturates in performance improvements as a function of dataset size sooner than other architectures like CPMs (even though HRNets still perform better overall).

      We thank the reviewer for the kind comments.

      Weaknesses:

      More details on training hyperparameters used (preferably full config if trained via mmpose).

      We have now included mmpose configs and anaconda environment files that allow researchers to use the dataset with specific versions of mmpose and other packages we trained our models with. The list of files is provided above.

      Should include dataset datasheet, as described in Gebru et al 2021 (arXiv:1803.09010).

      We have included a datasheet for our dataset in the appendix lines 621-764.

      Should include crowdsourced annotation datasheet, as described in Diaz et al 2022 (arXiv:2206.08931). Alternatively, the specific instructions that were provided to Hive/annotators would be highly relevant to convey what annotation protocols were employed here.

      We have included the list of instructions sent to the Hive annotators in the supplementary materials. File: OpenMonkeyWild Photograph Rubric.pdf

      Should include model cards, as described in Mitchell et al (arXiv:1810.03993).

      We have included a model card for the included model in the results section line 359. See Author response image 1:

      Author response image 1.

      It would be useful to include more information on the source of the data as they are collected from many different sites and from many different individuals, some of which may introduce structural biases such as lighting conditions due to geography and time of year.

      We agree that the source could introduce structural biases. This is why we included images from so many different sources and captured images at different times from the same source—in hopes that a large variety of background and lighting conditions are represented. However, doing so limits our ability to document each source background and lighting condition separately.

      Is there a reason not to use OKS? This incorporates several factors such as landmark visibility, scale, and landmark type-specific annotation variability as in Ronchi & Perona 2017 (arXiv:1707.05388). The latter (variability) could use the human pose values (for landmarks types that are shared), the least variable keypoint class in humans (eyes) as a conservative estimate of accuracy, or leverage a unique aspect of this work (crowdsourced annotations) which affords the ability to estimate these values empirically.

      The focus of this work is on overall keypoint localization accuracy and hence we wanted a metric that is easy to interpret and implement, in this case we made use of PCK (Percentage of Correct Keypoints). PCK is a simple and widely used metric that measures the percentage of correctly localized keypoints within a certain distance threshold from their corresponding groundtruth keypoints.

      A reporting of the scales present in the dataset would be useful (e.g., histogram of unnormalized bounding boxes) and would align well with existing pose dataset papers such as MS-COCO (arXiv:1405.0312) which reports the distribution of instance sizes and instance density per image.

      We have now included a histogram of unnormalized bounding boxes in the manuscript, see Author response image 2:

      Author response image 2.

      Reviewer #2 (Public Review):

      The authors present the OpenApePose database constituting a collection of over 70000 ape images which will be important for many applications within primatology and the behavioural sciences. The authors have also rigorously tested the utility of this database in comparison to available Pose image databases for monkeys and humans to clearly demonstrate its solid potential.

      We thank the reviewer for the kind comments.

      However, the variation in the database with regards to individuals, background, source/setting is not clearly articulated and would be beneficial information for those wishing to make use of this resource in the future. At present, there is also a lack of clarity as to how this image database can be extrapolated to aid video data analyses which would be highly beneficial as well.

      I have two major concerns with regard to the manuscript as it currently stands which I think if addressed would aid the clarity and utility of this database for readers.

      (1) Human annotators are mentioned as doing the 16 landmarks manually for all images but there is no assessment of inter-observer reliability or the such. I think something to this end is currently missing, along with how many annotators there were. This will be essential for others to know who may want to use this database in the future.

      We thank the reviewer for pointing this out. Inter-observer reliability is important for ensuring the quality of the annotations. We first used Amazon MTurk to crowd source annotations and found that the inter-observer reliability and the annotation quality was poor. This was the reason for choosing a commercial service such as Hive AI. As the crowd sourcing and quality control are managed by Hive through their internal procedures, we do not have access to data that can allow us to assess inter-observer reliability. However, the annotation quality was assessed by first author ND through manual inspections of the annotations visualized on all of the images the database. Additionally, our ablation experiments with high out of sample performances further vaildate the quality of the annotations.

      Relevant to this comment, in your description of the database, a table or such could be included, providing the number of images from each source/setting per species and/or number of individuals. Something to give a brief overview of the variation beyond species. (subspecies would also be of benefit for example).

      Our goal was to obtain as many images as possible from the most commonly studied ape species. In order to ensure a large enough database, we focused only on the species and combined images from as many sources as possible to reach our goal of ~10,000 images per species. With the wide range of people involved in obtaining the images, we could not ensure that all the photographers had the necessary expertise to differentiate individuals and subspecies of the subjects they were photographing. We could only ensure that the right species was being photographed. Hence, we cannot include more detailed information.

      (2) You mention around line 195 that you used a specific function for splitting up the dataset into training, validation, and test but there is no information given as to whether this was simply random or if an attempt to balance across species, individuals, background/source was made. I would actually think that a balanced approach would be more appropriate/useful here so whether or not this was done, and the reasoning behind that must be justified.

      This is especially relevant given that in one test you report balancing across species (for the sample size subsampling procedure).

      We created the training set to reflect the species composition of the whole dataset, but used test sets balanced by species. This was done to give a sense of the performance of a model that could be trained with the entire dataset, that does not have the species fully balanced. We believe that researchers interested in training models using this dataset for behavior tracking applications would use the entire dataset to fully leverage the variation in the dataset. However, for those interested in training models with balanced species, we provide an annotation file with all the images included, which would allow researchers to create their own training and test sets that meet their specific needs. We have added this justification in the manuscript to guide the other users with different needs. Lines 530-534: “We did not balance our training set for the species as we wanted to utilize the full variation in the dataset and assess models trained with the proportion of species as reflected in the dataset. We provide annotations including the entire dataset to allow others to make create their own training/validation/test sets that suit their needs.”

      And another perhaps major concern that I think should also be addressed somewhere is the fact that this is an image database tested on images while the abstract and manuscript mention the importance of pose estimation for video datasets, yet the current manuscript does not provide any clear test of video datasets nor engage with the practicalities associated with using this image-based database for applications to video datasets. Somewhere this needs to be added to clarify its practical utility.

      We thank the reviewer for this important suggestion. Since we can separate a video into its constituent frames, one can indeed use the provided model or other models trained using this dataset for inference on the frames, thus allowing video tracking applications. We now include a short video clip of a chimpanzee with inferences from the provided model visualized in the supplementary materials.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Please provide a more thorough description of the annotation procedure (i.e., the instructions given to crowd workers)! See public review for reference on dataset annotation reporting cards.

      We have included the list of instructions for Hive annotators in the supplementary materials.

      An estimate of the crowd worker accuracy and variability would be super valuable!

      While we agree that this is useful, we do not have access to Hive internal data on crowd worker IDs that could allow us to estimate these metrics. Furthermore, we assessed each image manually to ensure good annotation quality.

      In the methods section it is reported that images were discarded because they were either too blurry, small, or highly occluded. Further quantification could be provided. How many images were discarded per species?

      It’s not really clear to us why this is interesting or important. We used a large number of photographers and annotators, some of whom gave a high ratio of great images; some of whom gave a poor ratio. But it’s not clear what those ratios tell us.

      Placing the numerical values at the end of the bars would make the graphs more readable in Figures 4 and 5.

      We thank the reviewer for this suggestion. While we agree that this can help, we do not have space to include the number in a font size that would be readable. Smaller font sizes that are likely to fit may not be readable for all readers. We have included the numerical values in the main text in the results section for those interested and hope that the figures provide a qualitative sense of the results to the readers.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This rigorous and creative study uses an elegant combination of metabolomics, transcriptomics, and budding yeast molecular genetics to discover that (i) activating AMPK to maintain mitochondrial respiration fueled by cytosolic Acetyl CoA and (ii) increasing fatty acid synthesis independent of respiration drive independent pathways that increase the fitness of replicatively-aged budding yeast cells, albeit without increasing their lifespan. This work will be of interest to scientists in the field of aging and metabolism. Some clarifications in the text would address the following concerns, which would increase the impact of the study:

      (1) What does activation of AMPK (via PGDP-Sak1 expression) do to the replicative lifespan? How many bud scars, in general, do the subpopulations that are older - yet have less Tom70 (increased mitochondrial fitness) - have, after the 48 hrs timepoint that they are examining? How many divisions occurred in this 48hr time period - i.e. is it long enough to have all cells reach the end of their replicative lifespan? This information is important to rule out that a subset of the mutant cells just divided faster and hence had more divisions within 48 hrs (growing faster and living longer are different things). Having identical growth curves doesn't indicate per se that they all divide at the same rate, as there may be a subpopulation that divides faster and a subpopulation that doesn't grow so well.

      Increasing AMPK activity increases replicative lifespan [PMID: 25869125], but given our finding that AMPK activation splits the population, such replicative lifespan assays are hard to interpret. Bud scar counts have a similar issue. Hence we restricted the lifespan and bud scar analyses to wt and A2A which are more homogenous (Figures S2 B and E). A2A cells at 48 h have ~25% more bud scars than wt cells. Yes, by 48 h most of the cells have lost viability (Figure 2E). The reviewer is correct that you can't properly compare the lifespan curves if the cells divide at different rates, hence our follow-up test of wt at 48 h vs A2A at 40 h viability after we had confirmed that these time points captured cells at equivalent replicative ages (Figure 2D, E). This shows that viability of A2A is slightly lower than wt at matched age, indicating a slightly shorter lifespan. 

      (2) A2A cells do not have an extended replicative lifespan (RLS) but show an increase in the "low senescence" population (Figure 2). If the cells are not becoming senescent, why don't they have longer RLS? Not having a longer lifespan seems inconsistent with the statement that "bud scar counting confirmed that A2A cells reach a higher age than wild type", which comes back to how many times the cells can divide in the 48hr timepoint studied and their rate of cell division? Also, the lifespan curve shown is plotted against time, not cell division number, which does not take into account different division times of cells within the population (described above). It would be much more useful to show standard lifespan curves showing cell division numbers per lifespan per cell.

      Our observation that cells can reach the end of life without senescing is consistent with other studies that have studied the life course of individual cells by microscopy [PMID: 31291577, 32675375]. These studies always highlight some proportion of the cells that reach the end of life with no or minimal senescence, though this fraction varies with the experimental system. The question of why cells lose viability without senescing is a complete unknown in the field, and reflects a wider lack of consensus as to why yeast lose viability with replicative age.

      In liquid culture we can only assess viability over time, not cell division number, which we agree is not optimal and we are wary about making strong statements on lifespan for exactly the reasons the reviewer notes. Unfortunately, it is clear from the comparison of liquid and solid media lifespans performed by the Gottschling lab [PMID: 19652178] that culture system has a huge effect on lifespan, with cells in classical plate-based microdissection assays living far longer than the same strains do in liquid. This means that lifespans determined by microdissection-based assays are of questionable relevance to ageing studies performed in liquid culture. Senescence cannot be assayed on plates, while microfluidic systems lack the throughput necessary and preclude key techniques like RNA-seq, so liquid culture assays were the only option for this work. We agree that this leaves an unsatisfactory approximation for lifespan measurements, but we consider it critical that everything is measured in the same system. We therefore restricted our conclusion on lifespan to simply say that lifespan of A2A cells is not extended which our data in Figures 2D, E, S2B does support (see also answer to Q1), and therefore with the majority of A2A cells showing low senescence marks and high fitness at 48 h we can conclude that lifespan and fitness loss must be separable.

      We have added a note of these limitations of lifespan measurements in the materials and methods section of the manuscript.

      (3) Increased "fitness" of the old cells is implied from the increased size of the colonies that the old cells can make. However, this is a measure of the fitness of the daughters per se, not the old mother cells. Are the old mothers just passing on healthier mitochondria and more lipids to the daughters, such that they can divide more times? If the aged cells have an "increased fitness", why don't they divide more times themselves (i.e. live longer?).

      Yes, colony growth speed is defined by daughter cell replication, but as long as the daughters and subsequent generations divide at the same rate irrespective of whether they come from a young or old mothers then the size of the colony after 24 hours varies based on the time it took the initial mother to produce a daughter. This is what the assay really measures. We note that aged wildtype mothers often do not divide at all in the first 24 hours after being put on an agar plate (hence the tiny reported colony size), even though they do eventually produce a daughter which then forms a colony, whereas A2A cells tend to produce the first daughter rapidly whether young or old. It is known that daughters of aged wildtype mothers also divide slower, as to some extent do grand-daughters (PMID: 2644196), which will also contribute to differences in colony size, and this may well result from a lipid and/or mitochondrial contribution, but the primary driver of colony size in 24 hours is the time the mother took to initially divide. We have added this detail to the materials and methods section of the manuscript.

      As noted above, the mechanistic basis of lifespan is unknown, but although senescence can shorten lifespan, our work and that of others shows that lifespan is still limited in the absence of senescence.

      (4) The statement is made that "these experiments define two classes of aging cells with distinct metabolic needs, coherent with the model of two aging trajectories previously proposed (referencing Nan Hao's work)". However, the big difference here is that in Nan Hao's work, their two aging trajectories influenced the length of lifespan, but that does not appear to be the case here. That distinction should be made clear. Perhaps the authors could also speculate as to why the A2A yeast stops dividing after presumably the same number of cell divisions, even though they have an activated AMPK and activated fatty acid synthesis pathway.

      Yes, this is a good point and we have added this distinction to the Discussion:

      “Here we have characterised two classes of ageing cells seemingly differentiated by high and low availability of cytosolic Acetyl-CoA, consistent with a previous demonstration that ageing follows two trajectories in yeast though it should be noted that in this previous report, the two trajectories also differed in replicative lifespan (6).”

      We would love to speculate on why the A2A cells don't have an extended lifespan, but at this point we don't have a strong hypothesis. We have come up with many theories for this, but none that we haven’t managed to disprove experimentally. One thing worth considering is that many cells which lose replicative viability in liquid culture and probably in plate assays remain intact – for example, DNA and RNA integrity is not compromised over 24- 48 h – so those cells are probably not dead per se. But we also detect apoptosis-sized DNA fragments, which must come from dead cells, so there is clearly not a single mechanism defining the end of replicative lifespan.

      (5) I am a bit confused by the use of the word "senescence" by this lab here and in their previous growth on galactose studies. If yeast don't senesce, which is usually defined as an irreversible arrest of the cell cycle where cells stop dividing, shouldn't the yeast that do not senesce still be dividing and hence have a longer lifespan? Should a different term be used rather than senescence? Such as "fitness late in life". The authors giving their definition of senescence may help reduce this apparent contradiction.

      We completely agree, this is confusing and noted this distinction in the Introduction. Use of the term senescence to mean a loss of fitness late in life in yeast stems from the classical definition of senescence as applied to whole organisms. However, the term senescence as applied to cells has a more specific meaning in terms of the cell cycle as the reviewer notes. As an individual S. cerevisiae is both a cell and an organism, the terminology clashes. However, the marker we largely employ (Tom70-GFP) which in our hands is a very good proxy for fitness was originally defined as marking the senescence entry point (SEP), so overall we feel we can't avoid the term.

      Reviewer #2 (Public review):

      Summary:

      In this study, the authors investigate how cytosolic acetyl-CoA metabolism influences replicative aging in budding yeast. They propose that acetyl-CoA regulates aging through three major pathways: (1) mitochondrial transport to support mitochondrial function, (2) fatty acid synthesis, and (3) global protein acetylation. The data show that AMPK activation promotes mitochondrial import of acetyl-CoA and partially mitigates mitochondrial decline in a subset of aging cells.

      Furthermore, the engineered A2A strain, which enhances mitochondrial acetyl-CoA utilization while relieving inhibition of fatty acid synthesis, increases the proportion of cells exhibiting a "low senescence" phenotype.

      Overall, this is a thoughtful and potentially impactful study that advances our understanding of metab to olic control of aging. Addressing the points below, particularly by refining interpretations and, where feasible, incorporating additional analyses, will further strengthen the manuscript and its conclusions.

      Strengths:

      The study has several notable strengths. It addresses an important question by shifting the focus from lifespan to preservation of late-life fitness, which is highly relevant to aging biology. The work integrates metabolic, genetic, and functional analyses to link cytosolic acetyl-CoA flux with distinct aging outcomes, and the engineering of the A2A strain provides a clear and elegant demonstration of how coordinated pathway modulation can improve cellular fitness.

      Weaknesses:

      (1) While the manuscript focuses on mitochondrial transport and fatty acid synthesis, cytosolic acetyl-CoA is also a key regulator of histone acetylation and chromatin silencing. It would strengthen the study to consider whether acetyl-CoA depletion contributes to improved fitness through enhanced rDNA silencing. Given the well-established role of rDNA instability in yeast aging, additional experiments examining rDNA silencing and stability would be valuable. For example, monitoring rDNA copy number changes (not necessarily ERCs) under AMPK activation, oleic acid supplementation, and in the A2A strain, similar to approaches used in the authors' prior work, would help clarify whether chromatin regulation contributes to the observed phenotypes.

      We have added data addressing these points to the manuscript and Supplemental Figures 2, 3 and 4, though the outcomes are complex. Histone acetylation changes chromatin accessibility and could therefore alter global gene expression; in accord with this, RNA-seq shows that P<sub>GPD</sub>-SAK1 reduces known age-linked gene expression dysregulation. However, A2A does not further reduce the effect, meaning either that another driver exists in addition to cytosolic acetyl-CoA, or that age-linked gene expression dysregulation is unrelated to cytosolic acetyl-CoA. Oleic acid has little effect on age-linked gene expression dysregulation despite rescuing fitness. With regard to rDNA silencing, transcription of the rDNA intergenic spacer non-coding RNAs promotes ERC formation; we have added data showing that ERC accumulation is not reduced in A2A but slightly higher coherent with the higher replicative age of A2A at 48 h, which suggests silencing is not better in A2A. By RNA-seq, these intergenic spacer transcripts are massively upregulated with age, but this will be a consequence of the increased genomic copy number on ERCs; the upregulation is less in A2A than other conditions, but this arises because the log phase spacer transcript levels are higher and so does not reflect better rDNA silencing. We have previously assayed for heritable changes in rDNA copy number arising during ageing and found (to our surprise) absolutely nothing, so we don't expect any changes under these conditions. The upregulation of transcripts from Sir2-repressed telomeric and MAT loci with age is decreased in P<sub>GPD</sub>-SAK1 and A2A, but the effect size is not different from any other low-expressed genes so we do not think there is a particular effect at loci subject to chromatin silencing (see our previous study Zylstra et al PMID 37643194 for evidence that Sir2-mediated gene silencing is not affected by age). We have added our conclusions from these experiments to the Discussion.

      (2) The current data do not fully distinguish whether AMPK activation and oleic acid supplementation act on distinct subpopulations of aging cells. An alternative explanation is that oleic acid supplementation enhances mitochondrial function and acts additively with AMPK activation, thereby increasing the fraction of cells in the "low senescence" state. Since this distinction is not central to the main conclusions, I suggest softening the language around subpopulation specificity. Emphasizing instead that the A2A strain coordinately modulates multiple branches of acetyl-CoA metabolism to improve late-life fitness would maintain the strength of the central message without over interpretation.

      We respectfully disagree with the reviewer on this point. We show that P<sub>GPD</sub>-SAK1 rescues senescence in ~half the population by a Cat2/Mls1 dependent mechanism (Figure 1F). We then show that in A2A, which rescues most cells, deletion of CAT2/MLS1 restores senescence in ~half the cells (Figure 3F/G). This cannot be explained by an additive mechanism as this would either result in all cells being partially rescued in the P<sub>GPD</sub>-SAK1 and in the A2A cat2Δ mls1Δ mutants, which is definitely not the case either by Tom70-GFP or fitness. Instead the population splits into high/low senescence and fit/unfit cells in the different assays.

      On the specific point of whether lipid synthesis additively increases mitochondrial function, we have added oxygen consumption rate data showing that A2A cells respire more than P<sub>GPD</sub>-SAK1 at 48h but only by a relatively small amount (Figure S3D), so there is indeed an additive improvement in mitochondrial function, but too little to explain the difference in population fitness in our opinion.

      We realise that the reviewer is asking more specifically about oleic acid, but again in the flow data, Figure 4C, what changes with oleic acid or P<sub>GPD</sub>-SAK1 is the proportion of cells in the low Tom70 / high WGA sector. Under an additive effect model, oleic acid or P<sub>GPD</sub>-SAK1 individually would partially reduce Tom70 and partially increase WGA, but the population in the low Tom70 / high WGA sector has the same average Tom70/WGA values in oleic acid, P<sub>GPD</sub>-SAK1 or P<sub>GPD</sub>-SAK1+oleic acid. It is the proportion of cells in this population that changes. Furthermore, under an additive model, wildtype cells aged with oleic acid would not have highest fitness than P<sub>GPD</sub>-SAK1 or A2A (Figure 4D) as these individual cells would lack the mitochondrial upregulation from P<sub>GPD</sub>-SAK1.

      (3) The manuscript proposes that lipid starvation and excess acetyl-CoA are major drivers of senescence in distinct subpopulations of wild-type aging cells. This conclusion is not yet fully supported by the presented data. Direct measurements of age-dependent divergence in acetyl-CoA and fatty acid levels at the single-cell level would be needed to substantiate this model. Based on the current evidence, a more conservative interpretation would be that aging cells exhibit differential sensitivity to perturbations in acetyl-CoA and lipid metabolism. Accordingly, I recommend revising the statement in the Abstract ("We further implicate lipid starvation and excess acetyl coenzyme A availability as major drivers of senescence...") and the corresponding discussion text to better align with the data.

      We agree and have adjusted the abstract to make it clearer that the lipid starvation / excess acetyl-coA interpretation is a model.

      “Our findings support a model in which lipid starvation and excess acetyl-coenzyme A availability are major drivers of senescence in replicatively aged wild-type yeast.”

      Reviewer #3 (Public review):

      Summary:

      These findings suggest that PGPD-SAK1 yeast show a subpopulation with lowered TOM70-GFP expression in high bud scar staining aged cells. Deletion of CAT2 or MLS1 reduces this effect. A PGPD-SAK1 acc1S1157A double mutant (called "A2A" here) shows an even larger effect of lowered tom70 expression in high bud scar staining aged cells. Utilization of various additional mutants involved in acetyl-CoA transport, carnitine shuttle, respiration, etc., leads the authors to conclude that these shifts in TOM70-GFP in aged cells are linked to the AMPK-fatty acid metabolic regulatory system.

      Strengths:

      These extensive and clearly described experiments reveal interesting changes in TOM70-GFP intensity in subsets of aged yeast in several mutants eventually identified as linked to the AMPK-fatty acid metabolic regulatory system.

      Weaknesses:

      (1) 3 biological replicates for mRNASeq is low.

      Thank you for pointing this out. We performed another replicate after posting the initial preprint to confirm the finding but didn’t update the figure in the eLife-reviewed version. We have added this to the scatter plots and analysis in Figure 1, there are minor changes but the set of genes we followed up are still highly significant. For ageing experiments, we sequence to n=3 as a first pass which is sufficient to detect widespread age-linked gene expression effects, and add more replicates if required to solidify findings for specific sets of genes. Hence, the additional RNAseq experiments we have added to the manuscript to Address Reviewer 2’s comments on widespread gene expression effects are also n=3-4.

      (2) While "Traditional conceptions of ageing implicate a progressive accumulation of damage leading to systemic degradation in performance until death, with evolutionary pressures acting to maximise early life fitness and fecundity at the expense of ageing health." is tangential perhaps to the data and conclusions of the study, both claims of this sentence are at best controversial, and the manuscript is no weaker for their omission.

      We would prefer not to remove this sentence, which we see as important to a major message of the manuscript: that ageing does not have to involve a loss of fitness before death. Outside the ageing biology field, ageing is often described as the progressive wearing out of components leading to decline and death (‘like an old car’ is a common analogy); in the ageing field this is certainly controversial, but outside the field it remains the normal understanding. This is what we mean by traditional conceptions, and it is important to consider the contradiction between this widely held viewpoint and our findings (and of course those of many others in the ageing field).

      The second part of the sentence about evolutionary pressures alludes to antagonistic pleiotropy, which we have now made explicit. Antagonistic pleiotropies as a driving mechanism for ageing, while not universally accepted, are as far as we can tell the most widely accepted type of theory in the ageing field. Our interpretation that yeast are bet-hedging as a population growth strategy and this drives ageing in the long term is a classic antagonistic pleiotropy and we need to raise this concept in the introduction.

      (3) The statement that "Here, we determine the basis of senescence and fitness loss in replicatively ageing yeast" is a bit strong as a summary of the present careful work presented here. If the authors had created yeast mutants that retained fitness indefinitely, this would be a more appropriate strength of claim to summarize the work.

      We agree and have moderated this sentence:

      “Here, we show that senescence and fitness loss in replicatively ageing yeast can be almost completely avoided without extension of lifespan by rewiring the conserved AMPK-fatty acid metabolic regulatory system.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The labelling of Figure 3G horizontal axis needs to be realigned with the data.

      Fixed – thank you.

      Reviewer #2 (Recommendations for the authors):

      (1) In Figure 3G, the x-axis labels appear misaligned and should be corrected for clarity.

      Fixed – thank you.

      (2) Figures S3B and S3C appear to be mislabeled and should be revised.

      Fixed – thank you.

      (3) On page 6 (3rd paragraph), the statement that the beneficial impact arises from acetyl-CoA removal "rather than a benefit of respiration" may be overstated. The data support a role for acetyl-CoA removal but do not fully exclude a contribution from respiration. A more balanced phrasing would improve accuracy.

      We have revised this sentence and also added data:

      “Working in sip2Δ to avoid an increase in AMPK activity due to reduced Acetyl-CoA availability, we observed that ald6Δ increased the low senescence population through decreasing Tom70-GFP (S3C), and therefore the beneficial impact of PGPD-SAK1 on this pathway arises primarily through Acetyl-CoA removal. It is possible that respiration is adding to this benefit, and we detect a significant increase in Oxygen Consumption Rate in aged PGPD-SAK1 cells, but the further increase in A2A is smaller and we consider that this cannot fully explain the effect of acc1S1157A.”

      Reviewer #3 (Recommendations for the authors):

      This manuscript is clearly written, and the data are clearly presented. While 3 biological replicates is inadvisably low for mRNASeq, the subsequent experiments motivated by the genes identified there nevertheless stand on their own as presented.

      Thank you.

    1. Author response:

      The following is the authors’ response to the original reviews.

      This is a summary of the changes that have been made to the Reviewed Preprint:

      (1) The data from RettBASE which was analysed in the manuscript has been added in the form of four supplementary tables. Supplementary Table 1 contains the download of all MeCP2 mutations contained in RettBASE. Supplementary Tables 2-4 contain subsets of this data which were used in Figure 2B and Figure 3 SF2. Supplementary Table 4 also has the HGVS nomenclature for both e1 and e2 isoforms and the ClinVar Variation ID for each allele. Wording has been changed to clarify that analysis in the manuscript used this data from RettBASE and not the information that was deposited in ClinVar.

      (2) Similarly, Supplementary Tables 5-8 contain the gnomAD data that was used in the preparation of Figure 2, Figure 2 SF1 and Figure 3 SF1. The “high confidence” alleles in Supplementary Table 8 have been annotated with their HGVS names.

      (3) The criteria for selecting “high confidence” RettBASE and gnomAD alleles have been more explicitly stated in both the Results and Materials and Methods sections.

      (4) An additional “high confidence” RTT mutation (c.1152_1195) has been added to Figures 2B and 3B.

      (5) Figure 3 Supplementary Figure 2 has been added to show reading frame data for all frameshifting deletions in the C-terminal deletion-prone region (CT-DPR), showing that +2 frameshifts predominate in this larger data set, not just in the “high confidence” set. This has necessitated changing the previous Fig. 3 SF2 to Fig. 3 SF3.

      (6) A summary of the genetic alterations described in the manuscript, and their outcomes, has been added as Figure 7.

      (7) A simple flow chart which assists in the classification of human CTDs as “likely benign” or “likely pathogenic” has been added as Figure 8. This will aid future assessment of novel mutations in this region.

      (8) Additions have been made to the Materials and Methods section to comply with reporting guidelines.

      (9) Minor changes have been made to the text to correct typographical errors and to clarify meaning.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors scrutinized differences in C-terminal region variant profiles between Rett syndrome patients and healthy individuals and pinpointed that subtle genetic alternation can cause benign or pathogenic output, which harbors important implications in Rett syndrome diagnosis and proposes a therapeutic strategy. This work will be beneficial to clinicians and basic scientists who work on Rett syndrome, and carries the potential to be applied to other Mendelian rare diseases.

      Strengths:

      Well-designed genetic and molecular experiments translate genetic differences into functional and clinical changes. This is a unique study resolving subtle changes in sequences that give rise to dramatic phenotypic consequences.

      Weaknesses:

      There are many base-editing and protein-expression changes throughout the manuscript, and they cause confusion. It would be helpful to readers if authors could provide a simple summary diagram at the end of the paper.

      We have added a summary diagram, as suggested (Figure 7). We have also provided a flowchart which shows how to classify human CTDs as “likely benign” or “likely pathogenic” based on their location.

      Reviewer #2 (Public review):

      Summary:

      This study by Guy and Bird and colleagues is a natural follow-up to their 2018 Human Molecular Genetics paper, further clarifying the molecular basis of C-terminal deletions (CTDs) in MECP2 and how they contribute to Rett syndrome. The authors combine human genetic data with well-designed experiments in embryonic stem cells, differentiated neurons, and knock-in mice to explain why some CTD mutations are disease-causing while others are harmless. They show that pathogenic mutations create a specific amino acid motif at the C-terminus, where +2 frameshifts produce a PPX ending that greatly reduces MeCP2 protein levels (likely due to translational stalling) whereas +1 frameshifts generating SPRTX endings are well tolerated.

      Strengths:

      This is a comprehensive and rigorous study that convincingly pinpoints the molecular mechanism behind CTD pathogenicity, with strong agreement between the cell-based and animal data. The authors also provide a proof of principle that modifying the PPX termination codon can restore MeCP2-CTD protein levels and rescue symptoms in mice. In addition, they demonstrate that adenine base editing can correct this defect in cultured cells and increase MeCP2-CTD protein levels. Overall, this is a well-executed study that provides important mechanistic and translational insight into a clinically important class of MECP2 mutations.

      Weaknesses:

      The adenine base editing to change the termination codon is shown to be feasible in generated cell lines, but has yet to be shown in vivo in animal models.

      This work is the obvious next step and is in progress. However, with the rise in pre- and neonatal genetic testing we felt it was important to disseminate our findings as soon as possible. The family pedigree in Figure 3C is a clear illustration of this need

      Reviewer #3 (Public review):

      Summary:

      Guy et al. explored the variation in the pathogenicity of carboxy-terminal frameshift deletions in the X-linked MECP2 gene. Loss-of-function variants in MECP2 are associated with Rett syndrome, a severe neurodevelopmental disorder. Although 100's of pathogenic MECP2 variants have been found in people with Rett syndrome, 8 recurrent point mutations are found in ~65% of disease cases, and frameshift insertions/deletions (indels) variants resulting in production of carboxy-terminal truncated (CTT) MeCP2 protein account for ~10% of cases. Many of these occur in a "deletion prone region" (DPR) between c.1110-1210, with common recurrent deletions c.1157-1197del (CTD1) and c.1164_1207del (CTD2). While two major protein functional domains have been defined in MeCP2, the methyl-binding domain (MBD) and the NCoR interacting domain (NID), the functional role of the carboxy-terminal domain (CTD, beyond the NID, predicted to have a disordered protein structure) has not been identified, and previous work by this group and others demonstrated that a Mecp2 "minigene" lacking the CTD retains MeCP2 function suggesting that the CTD is dispensable. This raises an important question: If the CTD is dispensable, what is the pathogenic basis of the various CTT frameshift variants? Prior work from this group demonstrated that genetically engineered mice expressing the CTD1 variant had decreased expression of Mecp2 RNA and MeCP2 protein and decreased survival, but those expressing the CTD2 variant had normal Mecp2 RNA and protein and survival. However, they noted that differences between the mouse and human coding sequences resulted in different terminal sequences between the two common CTD, with CTD1 ending in -PPX in both mouse and human, but CTD2 ending in -PPC in human but -SPX in mouse, and in the previous paper they demonstrated in humanized mouse ES cells (edited to have the same -PPX termination) containing the CTD2 deletion resulted in decreased Mecp2 RNA and protein levels. This previous work provides the underlying hypotheses that they sought to explore, which is that the pathological basis of disease causing CTD relates to the formation of truncated proteins that end with a specific amino acid sequence (-PPX), which leads to decreased mRNA and protein levels, whereas tolerated, non-pathogenic CTD do not lead to production of truncated proteins ending in this sequence and retain normal mRNA/protein expression.

      In this manuscript, they evaluate missense variants, in-frame deletions, and frame shift deletions within the DPR from the aggregated Genome Aggregated Database (gnomAD) and find that the "apparently" normal individuals within gnomAD had numerous tolerated missense variants and in-frame deletions within this region, as well as frameshift deletions (in hemizygous males) in the defined region. All of the gnomAD deletions within this region resulted in terminal amino acid sequences -SPRTX (due to +1 frameshift), whereas nearly all deletion variants in this region from people with Rett syndrome (from the Clinvar copy of the former RettBase database) had a terminal -PPX sequence, due to a +2 frameshift. They hypothesized that terminal proline codons causing ribosomal stalling and "nonsense mediated decay like" degradation of mRNA (with subsequent decreased protein expression) was the basis of the specific pathogenicity of the +2 frameshift variants, and that utilizing adenine base editors (ABE) to convert the termination codon to a tryptophan could correct this issue. They demonstrate this by engineering the change into mouse embryonic stem cell lines and mouse lines containing the CTD1 deletion and show that this change normalized Mecp2 mRNA and protein levels and mouse phenotypes. Finally, they performed an initial proof-of-concept in an inducible HEK cell line and showed the ability of targeted ABE to edit the correct adenine and cause production of the expected larger truncated Mecp2 protein from CTD1 constructs.

      The findings of this manuscript provide a level of support for their hypothesis about the pathogenicity versus non-pathogenicity of some MECP2 CTT intragenic deletions and provide preliminary evidence for a novel therapeutic approach for Rett syndrome; however, limitations in their analysis do not fully support the broader conclusions presented.

      Strengths:

      (1) Utilization of publicly available databases containing aggregated genetic sequencing data from adult cohorts (gnomAD) and people with Rett syndrome (Clinvar copy of RettBase) to compare differences in the composition of the resulting terminal amino acid sequences resulting from deletions presumed to be pathogenic (n+2) versus presumed to be tolerated (n+1).

      (2) Evaluation of a unique human pedigree containing an n+1 deletion in this region that was reported as pathogenic, with demonstration of inheritance of this from the unaffected father and presence within other unaffected family members.

      (3) Development of a novel engineered mouse model of a previously assumed n+1 pathogenic variant to demonstrate lack of detrimental effect, supporting that this is likely a benign variant and not causative of Rett syndrome.

      (4) Creation and evaluation of novel cell lines and mouse models to test the hypothesis that the pathogenicity of the n+2 deletion variants could be altered by a single base change in the frameshifted stop codon.

      (5) Initial proof-of-concept experiments demonstrating the potential of ABE to correct the pathogenicity of these n+2 deletion variants.

      Weaknesses:

      (1) While the use of the large aggregated gnomAD genetic data benefits from the overall size of the data, the presence of genetic variants within this collection does not inherently mean that they are "neutral" or benign. While gnomAD does not include children, it does include aggregated data from a variety of projects targeting neuropsychiatric (and other conditions), so there is information in gnomAD from people with various medical/neuropsychiatric conditions. The authors do make some acknowledgement of this and argue that the presence of intragenic deletion variants in their region of interest in hemizygous males indicates that it is highly likely that these are tolerated, non-pathogenic variants. Broadly, it is likely true that gnomAD MECP2 variants found in hemizygous males are unlikely to cause Rett syndrome in heterozygous females, it does not necessarily mean that these variants have no potential to cause other, milder, neuropsychiatric disorders. As a clear example, within gnomAD, there is a hemizygous male with the rs28934908 C>T variant that results in p.A140V (p.A152V in e1 transcript numbering convention). This pathogenic variant has been found in a number of pedigrees with an X-linked intellectual disability pattern, in which males have a clear neurodevelopmental disorder and heterozygous females have mild intellectual disability (see PMIDs 12325019, 24328834 as representative examples of a large number of publications describing this). Thus, while their claim that hemizygous deletion variants in gnomAD are unlikely to cause Rett syndrome, that cannot make the definitive statement that they are not pathogenic and completely benign, especially when only found in a very small number of individuals in gnomAD.

      We have included the possibility that mutations found in gnomAD may give rise to less severe neurological conditions in the discussion.

      (2) The authors focus exclusively on deletions within the "DPR", they define as between c.1110-1210 and say that these deletions account for 10% of Rett syndrome cases. However, the published studies that are the basis for this 10% estimate include all genetic variants (frameshift deletions, insertions, complex insertion/deletions, nonsense variants) resulting in truncations beyond the NID. For example, Bebbington 2010 (PMID: 19914908), which includes frameshift indels as early as c.905 and beyond c.1210. Further specific examples from RettBase are described below, but the important point is that their evaluation of only frameshift variants within c.1110-1210 is not truly representative of the totality of genetic variants that collectively are considered CTT and account for 10% of Rett cases.

      The vast majority of C-terminal truncating mutations do occur within the “CT-DPR”, likely due to its C-rich nucleotide sequence and the presence of microhomologies within the region. Looking at frameshifting deletions in RettBASE that start after the NID, a large proportion of these end within the CT-DPR and result in a -PPX ending. We decided to restrict our analysis to the c.1110-1210 region to avoid including the rarer examples that may have a different reason for their pathogenicity. We do not assert that all C-terminal truncations are pathogenic due to this mechanism, but current evidence suggests that most are.

      (3) The authors say that they evaluated the putative pathogenic variants contained within RettBase (which is no longer available, but the data were transferred to Clinvar) for all cases with Classic Rett syndrome and de novo deletion variants within their defined DPR domain. Looking at the data from the Clinvar copy of RettBase, there are a number (n=143) of c-terminal truncating variants (either frameshift or nonsense) present beyond the NID, but the authors only discuss 14 deletion frameshift variants in this manuscript. A number of these variants have molecular features that do not fall into the pathogenic classification proposed by the authors and are not addressed in the manuscript and do not support the generalization of the conclusions presented in this manuscript, especially the conclusion that the determination of pathogenicity of all c-terminal truncating variants can be determined according to their proposed n+2 rule, or that all of the 10% of people with Rett syndrome and c-terminal truncating variants could be treated by using a base editor to correct the -PPX termination codon.

      It is important to state here that we did not use the data in ClinVar for our analysis, but the original information that was held in RettBASE. We have clarified this in the manuscript and have now included supplementary tables containing the data we downloaded from both RettBASE and gnomAD. Table 1 contains all the RettBASE entries with MECP2 mutations, while Tables 2-4 contain subsets of this data pertaining to CTDs. We have extended our “high confidence” set of RTT alleles to contain one more that was previously overlooked (c.1152_1195) to bring the total number of alleles to 15. Taken together these alleles account for 158 individual entries in RettBASE. We have now included an analysis of all frameshifting mutations in the CT-DPR (Supplementary Table 3, Fig. 3 SF2) which covers 69 different mutations and 260 individual entries. Of these, +2 frameshifts make up the large majority, in contrast to the gnomAD data shown in Supplementary Table 8 and Fig. 3 SF1.

      (4) The HEK-based system utilized is convenient for doing the initial experiments testing ABE; however, it represents an artificial system expressing cDNA without splicing. Canonical NMD is dependent on splicing, and while non-canonical "NMD-like" processes are less well understood, a concern is whether the artificial system used can adequately predict efficacy in a native setting that includes introns and splicing.

      We disagree with this opinion. We show that the loss of protein and mRNA seen with knock in mouse and human alleles is recapitulated when using a cDNA-based transgene in the HEK system, demonstrating that the mechanism of loss does not involve factors bound at splice junctions etc. We also demonstrate the effect of the A to G change at the stop codon is the same whether we do this by base editing our cDNA transgene in T-REx cells or by making the CTD1 X>W knock-in mouse. Both result in increased levels of a slightly extended but still truncated protein.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The phrase in the title, "an alternative therapeutic approach" is only insinuated in the manuscript, making it rather inappropriate to be in the title.

      In this study we use adenine base editing to modify RTT-causing CTD mutations in T-REx cells which is clearly a precursor to developing a therapy, utilising the new findings in this study. We therefore feel that the use of “an alternative therapeutic approach” can be justified.

      Reviewer #2 (Recommendations for the authors):

      I have a few minor comments for the authors to consider:

      (1) Please double-check Figure 2 Supplementary Figure 1, as the allele count for E394K does not appear to be in the thousands; rather, E397K seems to be the variant shown in the graph.

      Yes, this was an error and has been corrected to E397K in the text. Thank you for spotting it.

      (2) On page 12, the phrase “common DNA sequence features shared by all CTDs that give rise to RTT” might be better described as “amino acid sequence features.”

      This has been altered in the text as suggested.

      (3) On page 3, the sentence "analysis of patient mutations and experimental data from mouse models support a role in transcriptional repression" cites Gabel et al. 2015 and Kinde et al. 2016, which focus on null alleles but not patient mutations. It would be appropriate to also cite Johnson et al. 2017, which analyzed MeCP2 T158M and R106W patient mutations.

      This has now been cited.

      (4) In the same sentence, Bajikar et al. 2025 are described as studying the "acute loss of MeCP2," but gene expression was analyzed after a week or longer period of time, not minutes to hours as in degron-mediated degradation systems.

      The word “acute” has been removed from the text.

      (5) On page 7, the statement that "E394K is common... who were later found to have additional pathogenic MECP2 mutations" should include a supporting reference.

      We have now included the references Moncla et al (2002) and Wan et al (1999) to address this.

      (6) Similarly, on page 10, the sentence "these findings question the validity of two cases where individuals presented with classical Rett..." is missing a reference to the case report mentioned.

      The references Bienvenu et al (2000) and Philippe et al (2006) have been added.

      (7) While the manuscript is well written and full of detail, if space is an issue, the authors might consider tightening sections that reiterate findings from their 2018 HMG paper.

      A section discussing the CTD2 allele from the 2018 HMG paper has been removed from the results section.

      Reviewer #3 (Recommendations for the authors):

      (1) Overall, the manuscript is rather dense and potentially challenging to follow easily, especially for a non-expert reader.

      Minor edits have been made to the text which will hopefully make it easier to follow. We have also added two new figures (Figures 7 and 8) to summarize the different alleles and edits which appear in the paper, and to show how to determine whether a CTD in the region is likely to be benign or pathogenic.

      (2) The introduction of data presented in Figure 1 within the manuscript introduction seems inappropriate and should be moved to the results section.

      We would say that this is unconventional rather than inappropriate, and is referred to in the introduction, so we would prefer to leave it as it is.

      (3) Providing specific, common nomenclature for genetic repository variants (rs numbers, gnomAD IDs, etc) somewhere would be beneficial. This is an issue because of the complexity of numbering (either coding or protein) for MECP2 due to the different transcript-based numbering systems.

      This nomenclature is now included in the supplementary tables of data from RettBASE and gnomAD.

      (4) As described in the public comments, there are a number of MECP2 genetic variants listed in the Clinvar copy of RettBase, resulting in c-terminal truncations that are not mentioned or discussed within the manuscript. Without the level of detail present in the original form of RettBase (number of events, de novo, etc) in the currently available Clinvar iteration, it is unclear why a number of variants, even within the limited DPR region, were not mentioned. A supplementary file including the more complete information from the RettBase version, with a complete listing of all c-terminal truncating variants, and an explicit rationale for the exclusion of variants would be helpful.

      We have now included supplementary tables with our download of all MECP2 mutations which were held in RettBASE. We have further added tables with the subset of mutations that we have analysed and have more explicitly stated our criteria for defining the “high confidence” sets of mutations. We did not download the data relating to “evidence of pathogenicity” (ie de novo?, absent from parents etc) from RettBASE, but annotated our list of CTDs with this information while RettBASE was still available. This was used in Supplementary Table 3.

      (5) A discussion of the limitations, notably that the fact that the focus exclusively on deletion variants within a restricted region (c.1110-1210) does not truly represent all genetic variants that cumulatively account for 10% of Rett cases, is needed. Furthermore, as pointed out, not all frameshift variants, even those that are n+2, result in the -PPX termination that is presented as the pathogenic basis of c-terminal truncations and amenable to correction by ABE. This should be noted in the discussion, as well as consideration of the late nonsense variants that cause c-terminal truncations (some of which would be very similar to the deletion variants discussed but without the proposed primary pathogenic driver, -PPX).

      We do not claim to explain the pathogenic mechanism of all C-terminal frameshift mutations found in cases of Rett syndrome. There will certainly be some that do not fit our explanation. However, we believe we have shown evidence that a large proportion of CTDs in RTT will be amenable to the therapy we propose.

      (6) Regarding point 3 in the public review, specifically:

      (a) n=7 nonsense variants (S360X, K363X, E397X, R453X, E455X) that do not carry the destabilizing -PPX sequence.

      (b) n=136 frameshift indel variants beyond the NID.

      (i) n=11 that have indels that extend past the native stop codon, n=4 of which start within the DPR domain (c.1110-1210) but would have a different terminal sequence than their proposed pathogenic -PPX sequence.

      (ii) n=125 frameshift indels with terminal breakpoint before the native stop codon

      (c) n=89 that have start or stop points within c.1110-1210

      (d) n=72 not mentioned within the manuscript.

      (e) n=27 are n+1, with 26/27 having what the authors term as the "tolerated" -SPRTX ending, but 1/27 having a frameshift beyond this region (c.1133_1361)

      (f) n=45 are n+2, with 32/45 ending in -PPX (supporting authors conclusion), but 10/45 will use the frameshift stop codon preceding the -PPX and have a different terminal sequence, and 3/45 result in a frameshift termination beyond the -PPX sequence.

      (g) n=36 have breakpoints either before c.1110 or after c.1210

      (h) n=16 start before c.1110, with 11/16 ending before c.1110. 4 of these 11 are n+2, but would use the earlier frameshift stop codon and not have -PPX terminal sequence. For 5/16, the indel extends past c.1210, with the n+2 leading to frameshift termination codons beyond the -PPX sequence.

      (i) n=20 indels start beyond c.1210, with 8/20 being n+1 and 12/20 being n+2, with neither leading to the -SPRTX or --PPX termination sequences characterized in this manuscript.

      As mentioned previously, we have used data taken from RettBASE, not from ClinVar. Both RettBASE and ClinVar will contain MECP2 mutations found in cases of RTT which are not the causative mutation. Databases of this kind contain sequencing errors and mutations that have been mistakenly assigned as causative. It is therefore imperative that the publicly available information is screened to only include mutations that meet stringent criteria. This is why we chose to start by looking at high confidence sets of mutations, with our conclusions supported by analysis of all such mutations in our region of study.

      As mentioned in response to point 5 above, we do not claim to explain every mutation in the region, but believe this study reveals an important disease mechanism for a large proportion of CTDs, leading to a potential therapy. It also contains significant information for predicting the likely prognosis of individuals with CTDs, who may remain healthy but are currently informed that their mutation is likely pathogenic. At present it seems that this is often based solely on the presence of a frameshifting mutation with similarities to bona fide RTT CTDs, without strong evidence of pathogenicity.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript by Noirot-Gros et. al. presents a herculean effort to map the protein-protein interactome of the c-di-GMP signaling network in Pseudomonas fluorescens (Pf). C-di-GMP, the key driver of biofilm formation in bacteria, is controlled by a highly complex network of synthesis, degradation and effector proteins. Pf is no exception as it encodes dozens of such proteins. The authors use a Yeast Two-Hybrid approach genome-wide screen with 10 diguanylate cyclase (DGC) enzymes as bait to assess protein-protein interactions in this network. The results identify over one hundred such interactions with several different hubs, including c-di-GMP signaling, other signaling systems, membrane proteins, etc. The authors then explore the original bait proteins as well as identify interactors on biofilm formation-related phenotypes and swarming using a high-throughput CRISPRi expression knockdown approach. The amount of data generated is quite impressive. Much of the manuscript uses statistical-based network analysis to group different proteins based on their interactions or impact on phenotypes, which is a high-level analysis that can catalyze further study into this system. The authors chose three specific proteins to assess their impact on cell morphology, DNA repair, and protein localization. Overall, in my view, this is perhaps the best analysis of a c-di-GMP protein-protein interactome, and it provides a multitude of hypotheses to be tested. However, therein lies the weakness of the manuscript in that very few of these hypotheses are actually tested. But such is not the goal of this network analysis type of approach. Overall, I think the work will be highly impactful to those in the c-di-GMP field, and it provides a template for others attempting such analyses of protein-protein interactions.

      Strengths:

      The manuscript is impressive in the sheer scale of the protein-protein interactions identified, network analysis, and phenotypic analysis of specific proteins in the network. It is an impressive amount of work that could be very useful to the field. It is also statistically rigorous in its analysis of significant interactions or network nodes.

      Weaknesses:

      The weakness of the manuscript is that, with three exceptions, very few of the hypotheses are actually tested. For example, BifA is shown to be a network hub protein that interacts with many other diguanylate cyclases, and this is hypothesized to be through GGDEF heterodimerization. I appreciate that experimentally testing such a hypothesis is probably another entire manuscript, but some early forays into such ideas could be undertaken using AlphaFold structural modeling of protein-protein interactions compared with GGDEFs that don't form heterodimers. Also, an inherent weakness is that such detailed analyses of a c-di-GMP signaling network, in which each diguanylate cyclase and phosphodiesterase may respond to a unique cue, is that the network identified and the conclusions made are highly specific to the experimental conditions in which the work was done. Therefore, it is unclear how broadly these conclusions (i.e. BifA is the central regulator of c-di-GMP signaling) apply to other conditions. But it is impossible to get around such a limitation, and this work can lead to testing the robustness of the identified network in other environments.

      We would like to thank the reviewer sincerely for their positive comments on our manuscript and for their constructive feedback. We recognize the limitations arising from the lack of extensive knowledge regarding the environmental cues that trigger the regulation of all CDG activities in P. fluorescens. We hypothesize that DipA acts as a central local hub that positively or negatively regulates the activity of its interacting CDG partners throughout the cell life cycle, lifestyle transitions and environmental signals. Testing this hypothesis would indeed require extensive biochemical and omics approaches. However, strengthening the significance of DipA complexes in silico using AlphaFold is a very appealing proposition and we are currently considering including this analysis in the revised version of the manuscript.

      Reviewer #2 (Public review):

      Summary:

      In this manuscript, Noirot-Gros and coworkers investigated the network of c-di-GMP associated protein complexes in Pseudomonas fluorescens. They did so by using a genome-wide yeast two-hybrid screen, and that was further probed by phenotypic screening that focused on biofilm and motility phenotypes. From this network map, they discovered that the phosphodiesterase DipA interacts with the GGDEF domains of many c-di-GMP-binding proteins.

      Strengths:

      (1) Broadness of screen led to identification of new interactions: The genome-wide yeast two-hybrid screening approach permitted broad investigation of c-di-GMP-associated protein-protein interactions. These interactions included some previously validated interactions as well as newly discovered interactions.

      (2) Complementary experimental validation: The proposed network was experimentally validated, including by using a CRISPRi-based approach in which the expression of genes encoding proteins identified in the network was systematically suppressed, and then the impact on the biofilm and motility phenotypes was assessed.

      Weaknesses:

      The findings would have been strengthened by further biochemical analysis, but this is likely beyond the scope of the paper.

      We would like to express our gratitude to the reviewer for their positive evaluation assessment, and for taking into account the limitations of the study's scope.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, Noirot-Gross et al take an open-ended approach to elucidate the c-diGMP-associated protein complexes in Pseudomonas fluorescens. Starting with 10 cyclic d-GMP putative proteins, they use a combination of genome-wide two-hybrid system followed by CRISPRi-mediated exploration of phenotypes to describe the cyclic di-GMP-associated regulation of biofilm formation, and how it relates to other functions. Overall, this work presents an excellent example of how genome annotations can be further confirmed with the use of integrated functional genomic approaches. Some areas of improvement can be applied to this manuscript to enhance readability and provide a clearer distinction between confirmatory results and new findings, which are provided below:

      Strengths:

      (1) The authors have explored their findings extensively and provide a comprehensive view of the topic.

      (2) The combination of genome-wide explorations of protein-protein interactions with the more focused phenotypic exploration of the interactions found provides a solid framework for the work presented.

      Weaknesses:

      (1) Overall goal of the work:

      While articles that describe open-ended approaches can be comprehensive and descriptive in nature, the authors should have a main overall goal, which can guide the reader through the main and most compelling findings at the end. As written, the overall goal is not clear. The network perspective is interesting, and the focus on biofilm formation appears in the title. Why P. fluorescens? How is cyclic di-GMP-mediated regulation of biofilm formation in P. fluorescens different from P. aeruginosa? Why would it be studied? (Positive or negative regulation of biofilm formation?)

      We would like to express our appreciation to the reviewer for their thorough evaluation of our manuscript and for the constructive feedback they provided. The overall goal of this study will be further refined, and outlined in the introduction in the revised version of the manuscript.

      (2) Abstract:

      The abstract is very well written and guides the reader to the DipA as a hub protein in the network. From further reading, the article could clarify whether this finding is confirmatory or novel (does DipA play a similar role in P. aeruginosa?) It would be appropriate to mention the role of DipA in other Pseudomonas species from the beginning, and not only in the discussion session.

      (3) Introduction:

      The introduction is nicely written. An area of improvement could be giving more attention to protein interactions as relevant to c-di-GMP. The authors could consider an independent paragraph starting with line 84-85 "Protein-protein interactions involving DGCs, PDEs, and target effectors are crucial in establishing localized signalling through the generation of local pools of c-di-GMP", expanding on this particular aspect with an example of localized signal, after explaining that localization could help decipher specific function within the network of DGCs and PDEs. Then go into connecting biofilms with c-di-GMP and protein-protein interactions, using the example of GcbC and LapD.

      We propose highlighting the example to the local signalling cascade formed by the tripartite system YdaM, YciR and MlrA. This will be addressed in the revised version of the manuscript.

      (4) The rationale of choosing 10 PDEs could be clarified. The nice diagrams shown in the supplementary table could be used as part of Figure 1, so the reader understands why these proteins were used, and what is known about them (for example, add them as Figure 1a).

      We propose to include a specific section in the supplementary file to explain the whole rationale behind choosing these CDGs. These proteins were selected based on their involvement in different steps of biofilm formation in Pseudomonas, as well as their role in the ability of P. fluorescens strains to colonize plant roots.

      (5) Figures 1b and 2 convey the same information as in Figure 1a. They could be removed without affecting the understanding of the article.

      Figure 2 will be transferred in Supplementary as part of the Figure S1

      (6) CRISPRi and Figure 3. Figure 3 shows the methodology of CRICPR phenotypic screening. A diagram showing the CRISPRi system in P. fluorescens could help the non-expert reader. While the choice of 23 proteins related to the emerging hub DipA is clear, the choice of the other 33 genes could be better explained. Are these proteins already related to biofilm formation? Where are they part of the network detected? How about the other 14 SBW25 genes? The authors could clarify the rationale of the choices. Figure 4 could be combined with Figure 3 or moved to the supplementary material.

      A better description of the rationale behind the choice of tested interacting protein partners will be provided. We also agree to combine Figure 4 with Figure 3.

      (7) Figures 5, 6 and 7 represent solid network analysis of the findings. Still, they could be improved in clarity on the main findings. The authors conclude at the end of section 3.2.3 that there are networks that exert a "positive role" and a "negative role". The authors could show that in the figures, explaining what those roles are: more biofilm structural coding genes? positive or negative regulation of biofilm formation?)

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors claim that bacteria are guided by diffusiophoresis. They perform experiments of bacterial motility in microfluidic channels with salt gradients. The data show that P. putida bacteria swim towards higher sodium chloride concentrations, but there is no evidence that this is due to diffusiophoresis.

      Weaknesses:

      It is well known that bacteria perform chemotaxis in salt gradients (see e.g., PNAS 86, pp. 8358-8362, 1989). The underlying mechanism based on chemoreceptors is widely accepted, but the authors do not mention this possibility. I recommend a control experiment where the chemotaxis genes are knocked out. Even if this mechanism can be ruled out, the current data show no evidence for a mechanism based on diffusiophoresis.

      We thank the reviewer for raising this important comment. We agree that receptor-mediated salt taxis is a well-established mechanism in some bacteria, including the classic study by Qi and Adler (PNAS, 1989). We note, however, that the original manuscript did discuss this possibility and cited Qi and Adler in line 117: “We also note that we did not observe any significant difference in the tumble rates between the control and NaCl gradient cases (Figure 3h; Figure S2, SM), suggesting that NaCl gradients do not interfere with chemoreceptors (Qi and Adler, 1989).” In Figure S2, the run-time distributions show no significant difference between the no-gradient condition and the NaCl-gradient condition, with fitted tumble rates of λ = 0.41 s<sup>−1</sup> and λ = 0.38 s<sup>−1</sup>, respectively. We also do not observe a directional bias in run duration, namely longer runs up the salt gradient and shorter runs down the gradient, which would be expected for canonical chemoreceptor-mediated taxis.

      The basis for assigning the observed migration to diffusiophoresis is that the NaCl gradient produces a directional drift of the bacterial body without a measurable change in the run-and-tumble statistics. This behavior is consistent with our previous work [1], where non-motile bacteria were shown to undergo diffusiophoretic migration toward higher salt concentration. Because that migration occurred in non-motile cells and across different bacterial types and morphologies, it supports the interpretation that native bacterial surface charge can drive a non-specific diffusiophoretic response in salt gradients.

      That said, we agree with the reviewer that a genetic control would provide a stronger test against receptor-mediated chemotaxis. We will therefore perform additional experiments using a ∆cheA strain. Because CheA is required for canonical chemotactic signal transduction, observing the same directional migration in the ∆cheA mutant would directly test whether the NaCl gradient response persists in the absence of receptor-mediated chemotaxis. We will include these new data and revise the manuscript to more explicitly distinguish diffusiophoretic drift from chemoreceptormediated salt taxis.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate how salt gradients influence the transport of Pseudomonas putida in confined microfluidic environments. They report that salt gradients enhance directional migration, increase run persistence, and promote transport toward contaminant-rich regions. To explain these observations, the authors propose a physical steering mechanism in which differential diffusiophoretic mobilities of the cell body and flagellar bundle generate an aligning torque that reorients cells along the salt gradient.

      Strengths:

      The study addresses an interesting question at the interface of microbiology, complex fluids, and active matter. Their experiments suggest that salt gradients influence bacterial transport behavior and lead to more persistent, directional motion. Once confirmed, the proposed mechanism would broaden our understanding of how environmental gradients can shape microbial migration through physical interactions in addition to more traditional sensing-based pathways.

      Weaknesses:

      The main limitation of the current study is that the proposed steering mechanism is not directly demonstrated. The evidence for the diffusiophoretic torque is largely inferred from trajectory statistics and theoretical modeling. While the observed transport behavior is convincing, the causal link between the observed migration patterns and the proposed reorientation mechanism remains less well established. In particular, the manuscript focuses primarily on cell trajectories and transport properties, whereas the proposed mechanism fundamentally involves changes in cell orientation. Additional evidence connecting orientation dynamics to the proposed torque mechanism would strengthen the conclusions.

      We thank the reviewer for the constructive comments. We agree that the proposed steering mechanism should be supported by evidence that directly connects the salt gradient response to bacterial orientation dynamics, not only to trajectory-level transport statistics.

      We would like to clarify that the original manuscript already includes an orientation-based analysis in Figure 4f,g in the main text. The corresponding methodology and results are described in lines 173–183 and in the Supporting Information. Specifically, we quantified cell steering by measuring the change in body angle, ∆θ, along individual run trajectories as a function of arc length, s, using the orientation correlation ⟨cos(∆θ)⟩<sub>s</sub>. In the absence of salt gradients, the orientation correlation decays slowly with arc length, indicating persistent swimming along the initial

      Author response image 1.

      Instantaneous angular velocity as a function of heading angle relative to the salt gradient orientation. (a) Experimental and (b) simulated mean angular velocity of cells as a function of heading angle θ (measured relative to the gradient direction; θ = 0° points toward the gel/high-salt side) in the absence (blue) and presence (red) of a NaCl gradient. Positive and negative values indicate counterclockwise and clockwise rotation, respectively, with arrows showing rotation direction. Under the gradient, both experiments and simulations show a signed, angle-dependent rotation rate that is largest near θ = ±90° and approaches zero near θ = 0° and 180°, consistent with a restoring torque that steers cell heading toward the gradient direction. Simulations reproduce this behavior with comparable magnitude to the experimental measurements, and in the absence of a gradient, angular velocity remains relatively small in the no-gradient case with no consistent directional bias across heading angles.

      run direction. Under a salt gradient, the correlation decays more rapidly, indicating stronger directional reorientation during runs. Because the tumble statistics do not change significantly between the no-gradient and salt-gradient conditions, this enhanced orientational decorrelation is not attributed to increased tumbling or rotational noise. Instead, it is consistent with continuous deterministic steering during runs, as expected from a diffusiophoretic torque acting on the cell body–flagellar bundle system.

      To further address the reviewer’s concern, we performed an additional orientation-dynamics analysis using the same dataset shown in Figure 4. Following the approach used by Stehnach et al. [2], we calculated the instantaneous angular velocity during individual runs as a function of the cell heading angle relative to the salt gradient direction. The cell orientation was obtained from the run trajectories, and tumble events were excluded because they produce large transient angular velocity spikes that are not representative of continuous steering during runs.

      The new analysis is shown in Author response image 1. Under the no-gradient condition, the angular velocity remains small and nearly independent of heading angle. In contrast, under the salt gradient condition, the angular velocity becomes strongly heading-dependent. The angular velocity is largest when cells swim nearly perpendicular to the salt gradient, where a steering torque is expected to be maximal. Moreover, the sign of the angular velocity indicates rotation toward alignment with the gradient direction. This behavior is consistent with the proposed diffusiophoretic torque mechanism and provides a direct link between the observed transport behavior and salt-gradient-induced reorientation dynamics.

      We plan to add this angular velocity analysis to the revised manuscript and revise the relevant text to make the connection between trajectory statistics, orientation dynamics, and the proposed torque mechanism clearer.

      A related concern is whether alternative physical mechanisms associated with the imposed salt gradients have been fully excluded. For example, weak flow-mediated effects or other hydrodynamic influences could potentially contribute to the observed transport behavior. The manuscript would benefit from a more thorough discussion of such possibilities and a clearer justification for why the proposed diffusiophoretic mechanism should be regarded as the dominant explanation.

      We thank the reviewer for raising this important point. We agree that alternative physical mechanisms associated with the imposed salt gradient should be considered explicitly. In the revised manuscript, we will add a more detailed discussion explaining why flow-mediated or other hydrodynamic mechanisms are unlikely to account for the observed steering behavior.

      First, the characteristic diffusiophoretic drift velocity in our experiments is approximately u<sub>d</sub> ≈ 1 µm/s, corresponding to only about 1–5% of the typical swimming speed of P. putida. Thus, the proposed mechanism does not require externally driven advection of the cells. Instead, the salt gradient produces a weak but persistent differential diffusiophoretic slip on the cell body and flagellar bundle, which can generate a reorienting torque during active swimming.

      We also considered whether diffusio-osmotic flow along the channel walls could generate sufficient shear to induce rheotaxis. In a dead-end channel, the diffusio-osmotic velocity profile can be estimated as [3]

      which gives a wall shear rate (at z = h) of

      This value is below the shear rate threshold reported by Marcos et al. [2], where rheotactic drift becomes negligible for S < 0.1 s<sup>−1</sup>. Therefore, the shear generated by diffusio-osmotic wall flow in our experiments is too weak to explain the observed directional reorientation. This effect would be even smaller for P. putida, whose thin flagellar bundle is expected to experience weaker shear-induced alignment than organisms with larger flagellar structures.

      We further considered viscotaxis as a possible mechanism. However, the viscosity difference between 1 mM and 100 mM NaCl solutions is marginal. This contrast is far smaller than the approximately 4–5-fold viscosity difference reported to induce strong viscophobic turning in bacteria [4]. Thus, salt-gradient-induced viscosity variations are insufficient to account for the measured steering response.

      Taken together, these estimates indicate that hydrodynamic shear, rheotaxis, and viscotaxis are too weak under our experimental conditions to explain the observed migration and orientation dynamics. In contrast, the proposed nonuniform diffusiophoretic mechanism naturally accounts for the key observations: directional migration up the salt gradient, enhanced curvature during runs, heading-dependent angular velocity, and the absence of significant changes in tumble statistics. We will include this analysis in the revised manuscript to clarify why diffusiophoretic steering is the dominant mechanism under the present conditions.

      The manuscript would also benefit from a clearer positioning within the broader literature on physically induced microbial transport and swimmer reorientation. Previous studies have demonstrated directed migration arising from rheotaxis (Marcos et al., 2012, PNAS) and viscosity-gradient-induced steering (Stehnach et al., 2021, Nature Physics). While the mechanism proposed here appears distinct, a more explicit discussion of how the present work relates to these earlier studies would help readers better understand the specific conceptual advance being made.

      We thank the reviewer for this helpful suggestion. We agree that the manuscript should more clearly position the proposed mechanism within the broader literature on physically induced microbial transport and swimmer reorientation.

      In the revised manuscript, we have added a discussion comparing our results with prior studies on rheotaxis and viscosity-gradient-induced steering. Specifically, we now discuss the work of Marcos et al. [4], which showed that shear flow can generate a torque on the helical flagellum of Bacillus subtilis, producing rheotactic alignment independent of chemical sensing. We also discuss the work of Stehnach et al. [2], which showed that viscosity gradients can steer Chlamydomonas reinhardtii through asymmetric viscous drag on its two flagella, producing viscophobic turning down the viscosity gradient.

      The mechanism proposed in the present work is distinct from both of these cases. Unlike rheotaxis, it does not require externally imposed shear flow. Unlike viscophobic turning, it does not rely on a substantial viscosity contrast. Instead, we propose that a salt concentration gradient generates differential diffusiophoretic motion of the cell body and flagellar bundle, producing a torque that continuously reorients swimming cells along the gradient. This mechanism therefore identifies salt gradients as a distinct physical cue capable of steering bacteria through surface-mediated transport rather than through flow, viscosity contrast, or canonical chemoreceptor signaling.

      We have added the following text to the revised manuscript: ”Our findings also relate to other physical mechanisms of microbial reorientation. Bacterial rheotaxis, arising from a torque generated by shear flow acting on the helical flagellum, steers Bacillus subtilis independent of any chemical gradient [4]. Viscosity gradients similarly drive a viscophobic turning in the alga Chlamydomonas reinhardtii, where uneven viscous drag on its two flagella produces a torque that reorients cells down the gradient [2], a behavior confirmed by measuring angular velocity as a function of heading angle and revealing a sinusoidal form, ω(θ) = −ω <sub>visc</sub> sin(θ), which resembles what we report here for diffusiophoresis in (Figure S4). This identifies salt gradients, independent of flow or viscosity, as a distinct physical route by which swimming cells can be steered and guided, broadening the set of known nonchemoreceptor mechanisms for directed microbial transport.” References

      (1) V. S. Doan, P. Saingam, T. Yan, S. Shin, A trace amount of surfactants enables diffusiophoretic swimming of bacteria, ACS Nano 14 (10) (2020) 14219–14227.

      (2) M. R. Stehnach, N. Waisbord, D. M. Walkama, J. S. Guasto, Viscophobic turning dictates microalgae transport in viscosity gradients, Nature Physics 17 (8) (2021) 926–930.

      (3) S. Shin, E. Um, B. Sabass, J. T. Ault, M. Rahimi, P. B. Warren, H. A. Stone, Size-dependent control of colloid transport via solute gradients in dead-end channels, Proc. Natl. Acad. Sci. 113 (2) (2016) 257–261.

      (4) Marcos, H. C. Fu, T. R. Powers, R. Stocker, Bacterial rheotaxis, Proceedings of the National Academy of Sciences 109 (13) (2012) 4780–4785.

    1. Author response:

      General Statements:

      We appreciate the reviewers for the critical review of the manuscript and the valuable comments. We have carefully considered the reviewer’s comments and have revised our manuscript accordingly.

      Point-by-point description of the revisions:

      Reviewer #1 (Evidence, reproducibility and clarity):

      Major comments

      (1) This study leaves out lipid metabolism as a major energy metabolism pathway relevant to AD. The authors themselves cite the significance of acylcarnitines and CPT1A in AD (pg. 3, lines 32-33, pg. 4, lines 1-2). Lipid metabolism and homeostasis is known to be disrupted in AD1. Fatty acid oxidation is a known energy source in the prefrontal cortex2 and will also generate acetyl coA, which this study reveals is a significant decreased metabolite in AD. Furthermore, sphingomyelin emerges as one of the major decreased DEMs as well. Thus, lipid metabolism should be highlighted in Figure 3 and discussed throughout the manuscript; otherwise its omission should be clearly stated and justified.

      We appreciate the reviewer’s insightful comment regarding a critical role of lipid metabolism in AD. We recognize that lipid metabolism is a metabolic pathway deeply involved in AD pathology (Baloni et al., 2022, 2020; Varma et al., 2021). Accordingly, we have revised the Limitations section to more strongly emphasize its role as a vital energy source (pg. 13, lines 15-17). Regarding the visualization of lipid metabolism, we extracted lipid-related pathway from the trans-omic network but found that the regulatory relationships among DEPs and DEMs were excessively complex and interconnected. Thus, interpreting this regulatory network seemed to be more challenging compared to the other energy production pathways presented in our manuscript. Therefore, we have concluded that the pathway analysis in our trans-omic network may not be suitable for deeply elucidating the lipid dysregulation in AD. We have added a statement acknowledging this as a limitation of our current methodology in the revised manuscript (pg. 13, lines 13-22).

      (2) The covariates used for differential analysis should be discussed and justified. Notably, age is used as a covariate for transcriptomic analysis but not proteomic and metabolomic analysis, with no justification. Additionally, given the known importance of lipid metabolism in AD and the putative role of APOE in lipid homeostasis3, APOE genetic status should be considered as a covariate, or its omission should be justified.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, such as age at death and RIN, is that these data were not available for each sample. Thus, we referred to the original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, and years of education. Regarding the proteomic dataset, in the original article (Johnson et al., 2020), age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (3) The authors make a conclusion statement that suggests intervention: "Collectively, our data suggests that preserving or improving the ability to produce ATP and early intervention in the process of nitrogen metabolism are candidates for the prevention and treatment of dementia" (pg. 12, lines 12-14). This claim is not well-supported by the evidence provided in the study. There are a few limitations: (a) This was an observational, not interventional study; (b) The study did not establish whether the metabolic disruptions are causes or effects in AD; and (c) ATP or other bioenergetic indicators were not directly measured. Therefore, any statements about potential interventions should be removed or qualified as highly speculative.

      We agree with the reviewer that the statement regarding potential interventions was not sufficiently supported by our analyses. Accordingly, we have removed the sentence regarding prevention and treatment from the revised manuscript (e.g., we have deleted final paragraph of the previous manuscript).

      (4) In conjunction with the last point, the main conclusion of the study is that energy production is down in AD. The data presented in Figure 3 are consistent with this conclusion, but it is far from definitive due to limitations stated above in comments 3a and 3b. The authors should offer additional support for this conclusion: experimental follow-up, flux modeling, analysis of alternative datasets with ATP measurement, causal inference.

      We sincerely thank the reviewer for this valuable and constructive suggestion. Regarding flux modeling, we agree that metabolic flux analysis could provide important mechanistic insight. Indeed, previous studies have applied flux modeling in the context of lipid metabolism in Alzheimer’s disease (Baloni et al., 2022). We also attempted to perform flux modeling focusing on energy metabolism. However, we found it difficult to obtain biologically meaningful and robust results and therefore decided not to include these analyses in the current manuscript.

      With respect to ATP measurements, we fully agree that direct evidence of altered ATP levels would further strengthen our conclusion. However, to the best of our knowledge, there are currently no publicly available large-scale datasets that directly measure ATP levels in human postmortem brain tissues. This limitation makes it challenging to incorporate validation in the present study.

      Regarding experimental follow-up, we agree that functional validation is essential to confirm the mechanistic implications of our findings. We are actively considering follow-up experimental studies. However, we consider the present work to be a multi-omic integrative analysis aimed at identifying key molecular alterations and generating biologically important hypotheses. We have revised the Limitation section to more clearly position this manuscript as an observational systems-level analysis (pg. 13, lines 20-22).

      (5) The validation analysis did not sufficiently show the generalizability of this study's results. The authors demonstrated a correlation of 0.53 to the MSBB transcriptomics data and 0.60 to the AMP-AD DiverseCohorts proteomics data. Beyond these correlation coefficients, no meaningful comparison between the datasets is offered. How concordant are the differentially expressed features (or pathways) between the datasets? How robust would the trans-omic network be if incorporating the alternate datasets? Is the main conclusion (energy metabolism is down in AD) supported by the validation datasets? We think this analysis should be expanded and described in the main text.

      Although the results for external metabolomics datasets are reported in Fig S2C, correlation coefficients with the external data are not reported. The authors state, "Note that each study used different definitions for AD and CT groups, had variations in measurement methods and brain regions analyzed." We appreciate these limitations. However, the external data should be re-analyzed using the same definitions of AD and CT, if possible. The limitations and results (which DEMs are shared between datasets) should be discussed in the main text.

      We thank the reviewer for this important comment regarding the generalizability of our findings. In the revised manuscript, we have expanded the validation analyses and summarized the results in Figure S2. First, at the transcriptomic level, Figure S2B and S2C show the overlap between up- and downregulated genes in AD identified in our ROSMAP-derived analyses and those reported in a previously published large-scale meta-analysis of 2,114 postmortem samples across seven brain regions (Wan et al., 2020). A substantial proportion of DEGs were shared, supporting cross-cohort and cross-region robustness to some extent. At the proteomic level, Figure S2E shows a comparison between the ROSMAP and the AMP-AD DiverseCohorts datasets. We highlighted the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3 and calculated a separate correlation coefficient for this subset (Pearson coefficient = 0.86, p-value = 1.5e-7), further supporting our main conclusion. In addition, to assess the concordance between the two datasets in a threshold-independent manner, we additionally performed Rank-Rank Hypergeometric Overlap (RRHO) analysis (Figure S2E). RRHO analysis (Cahill et al., 2018; Plaisier et al., 2010) enables the comparison of ranked protein lists without relying on arbitrary differential expression cutoffs and has been used for cross-dataset comparison in several previous studies (Fröhlich et al., 2024; Maitra et al., 2023). The RRHO heatmaps demonstrated significant enrichment in the concordant quadrants, confirming systematic agreement between datasets beyond simple correlation coefficients. For metabolomics, Figure S2G shows RRHO analyses comparing the ROSMAP metabolomic data with other datasets measured by the same UPLC-MS/MS platform (Batra et al., 2024; Novotny et al., 2023), demonstrating significant concordance in ranked metabolite changes in AD.

      (6) The glycolysis analysis and discussion needs more development. Glycolysis and gluconeogenesis share many of the same enzymes, but they are not the same pathway and should not be discussed as such. To make a claim about the overall influence of enzyme and metabolite levels on glycolysis, the authors should focus on the energetically committing steps of glycolysis (hexokinase, phosphofructokinase, pyruvate kinase) in Figure 3A, and include the full/current version of the figure in the supplement. Gluconeogenesis-specific enzymes (pyruvate carboxylase, PEPCK) are not mentioned at all - are they among the DEPs/DEGs?

      We appreciate the reviewer’s comment regarding the distinction between glycolysis and gluconeogenesis pathway. Among the gluconeogenesis-specific enzyme proteins, G6PC1, FBP1, PC, and PCK2 were measured in our dataset, but none of them were identified as DEPs. In addition, gluconeogenesis is a process that occurs primarily in the liver and kidney rather than the brain. Given this biological context and the lack of significant changes in relevant enzymes, we have revised the terminology throughout the manuscript, replacing “glycolysis/gluconeogenesis pathway” with “glycolysis pathway” in the revised version.

      (7) Given that there wasn't good concordance between the DEGs and DEPs, did including the mRNA and transcription factor layers in the network really add anything useful? It seems like the main conclusions of the manuscript were driven by the protein and metabolite layers only. How many of the DE metabolic enzymes were coregulated at the transcript and protein level? It would be useful to include the 5-layer trans-omic network in the supplement to display these results. Given your network, at what level does it appear that energy metabolism is regulated?

      It is true that our primary conclusion regarding the regulation of energy metabolism is driven by the changes in protein and metabolite abundance. However, we consider the low concordance between mRNA and protein expression itself to be an important feature of AD pathology, as also reported in previous studies (Johnson et al., 2022; Tasaki et al., 2022). Although we did not perform a further analysis of this discordance, we believe that including the TF and mRNA layers into the metabolic trans-omic network strengthens a system-wide view of metabolic dysregulation in AD.

      Regarding the mRNA changes corresponding to the DEP enzymes, please refer to Figure S7A.

      (8) Comment further on the results from Figure 2D. What can be learned from identifying metabolites with the greatest degree centrality? What pathways other than energy metabolism are highlighted by the trans-omic network?

      We assume that some energetic indicators, including AMP and acetyl-CoA, and nitrogen metabolism-related metabolites, Glu, 2-oxoglutarate, and urea, can be potential key regulators of dysregulated metabolism in AD.

      (9) (Suggestion) We suggest the authors leverage their trans-omic network in additional ways beyond giving a snapshot of a few energy metabolism pathways. The analysis of top DEMs could go further. What pathways are impacted beyond energy metabolism? Among the metabolic reactions allosterically regulated by top DEMs, what metabolic pathways are enriched?

      We identified the enriched metabolic pathways that were allosterically regulated by DEMs in AD using Fisher’s exact test. Alanine, aspartate, and glutamate metabolism pathways were significantly enriched in 2-oxoglutarate, glutarate, alanine, and glutamate-regulating metabolic reactions. Arginine and proline metabolism pathway was enriched in N-methyl-L-arginine and putrescine-regulating metabolic reactions. Arginine biosynthesis pathway was enriched in arginine-regulating metabolic reactions. Glycerophospholipid metabolism pathway was enriched in CDP-ethanolamine-regulating metabolic reactions. Glycine, serine, and threonine metabolism pathway was enriched in serine-regulating metabolic reactions. Purine metabolism pathway was enriched in AMP-regulating metabolic reactions. Pyrimidine metabolism pathway was enriched in deoxyuridine and thymidine-regulating metabolic reactions. Sphingolipid metabolism pathway was enriched in sphingosine-regulating metabolic reactions. However, this analysis did not yield sufficiently valuable insights into the regulatory relationships among biomolecules in AD. Thus, we did not include these results in the revised manuscript.

      (10) (Suggestion) Figure 3 shows that most differential signal in AD points to lower energy production due to the combination of differentially expressed metabolites and enzymes, but we are not given much context about the strength of these among all the differential signals. We would suggest including volcano plots where the features of interest, i.e. DE enzymes and metabolites, are colored differently (or a similar figure).

      We thank the reviewer for this constructive suggestion. To provide better context regarding the importance of the differential signals, we have added volcano plots for mRNAs, proteins, and metabolites in Figure S4A, B, and C.

      (11) (Suggestion) The PPI network could be better leveraged to understand metabolic changes in AD. If nodes are grouped into subnetworks (e.g. by Louvain / Leiden clustering) and tested for pathway enrichment, could you find functional subnetworks of coordinately up- and down- regulated metabolic enzymes? This could yield some pathways of interest beyond the energy metabolism pathways already highlighted.

      We appreciate the reviewer’s suggestion to utilize the PPI network for subnetwork analysis. However, it is important to note that the proteomic dataset analyzed in this study is derived from the original work of (Johnson et al., 2020). In that paper, the authors already performed a Weighted Gene Co-expression Network Analysis (WGCNA) across several datasets to identify co-expressed modules and functional pathways.

      Given this, we assumed that applying additional clustering methods to the same dataset would be unlikely to yield significant biological insights beyond the established findings.

      Minor comments

      (1). "All genes" and "all metabolites" should not be the background for the proteomic and metabolic pathway enrichment analysis by Metascape and MetaboAnalyst. The background should be limited to the proteins and metabolites that were measured.

      We fully agree with the reviewer that using “all gene” or “all metabolites” as a background is not suitable for enrichment analyses. As suggested, we have revised the enrichment analyses using the measured proteins and metabolites as a background in both Metascape and MetaboAnalyst (Fig. S4D).

      (2) Highlight the metabolic enzymes in Fig S2B. Calculate a separate correlation coefficient for the enzymes extracted in the energy metabolism analysis from Fig 3.

      We appreciate the reviewer’s suggestion to refine the correlation analysis. As requested, we have revised Fig. S2D to explicitly highlight the subset of enzymes involved in the energy metabolism analysis shown in Fig. 3. We calculated a separate correlation coefficient for the subset (Pearson coefficient = 0.86, p-value = 1.5e-7).

      (3) Use a multiple hypothesis adjusted p-value or q-value in Figure S3.

      We agree with the reviewer regarding the necessity of correcting for multiple comparisons. Accordingly, we have revised Fig. S4D using q-values.

      (4) Describe the methods used to calculate the logFC values from the validation dataset.

      We have revised the Methods to include a detailed description of the procedure used to calculate the log2FC values for the validation datasets (pg. 21, lines 13-15).

      (5) It is difficult to read Figure 3. We would recommend really emphasizing to the reader to refer to Fig S7B as a "key" to this figure. The description of the red/blue arrows and nodes in the methods section (pg. 24, lines 21-36, pg 25, lines 1-4) were also helpful, but very lengthy. We recommend putting an abridged version of this description into the Fig S7 figure legend.

      We appreciate the feedback regarding the readability of Fig. 3. As recommended, we have revised the manuscript to explicitly direct readers to Fig. S8B as an essential “key” for interpreting the network visualization (pg. 8, lines 28). Furthermore, we have added an abridged description of the network elements to the legend of Fig. S8B.

      (6) The S7 figure legend should refer to panels A and B, not E and F.

      We apologize for this oversight. We have corrected the legend of Fig. S8.

      (7) (Suggestion) Are any of the differentially expressed metabolites allosteric regulators of the DE transcription factors? This could be interesting to discuss.

      We appreciate the reviewer’s insightful suggestion about the potential allosteric regulation of the DETFs by DEMs. We conducted an extensive literature search to identify any reports related to this perspective. However, to the best of our knowledge, no such direct interactions have been reported to date.

      Reviewer #1 (Significance):

      The study's strength lies in leveraging three omics modalities across large patient cohorts (n ~ 150-240) to identify coherent signals between transcriptomics, proteomics, and metabolomics in postmortem DLPFC tissue. It was encouraging to see that the main result, showing downregulation for TCA, oxidative phosphorylation, and ketone body metabolism, emerged from consistent signals across both proteomics and metabolomics. This result was consistent with previous findings in other models cited by the author4,5 and other studies 6,7 demonstrating deficiency in energy-producing pathways in AD.

      Another strength of the study is the application of thoughtful methodology to connect differentially expressed proteins and metabolites via an intermediate data layer of metabolic reactions. The authors leverage the KEGG and BRENDA databases and apply sound logic to estimate the effects of enzyme level and metabolite level on pathway activity, with metabolites serving as substrate, product, or allosteric regulator for reactions. This trans-omic network methodology was developed in previous studies cited by the author8,9.

      However, as written, this study is limited in its contribution of new knowledge to the AD research field. The main conclusion (energy production is down in AD, due to regulatory disruption of energy metabolism) is not strongly supported (see comments 1, 3, and 4 for elaboration). The evidence could be improved by orthogonal approaches: further experimentation, further integration of external datasets, causal modeling, or flux modeling. Alternatively, even in the absence of new experimental and computational approaches, the story could be made more complete by further leveraging the trans-omic network to provide insights into (a) the regulation of energy metabolism; and (b) the impacts of key disrupted metabolites (see comments 7-9).

      The study is also limited in its demonstrating the power of these methodologies to provide integrative insights. As mentioned above, the integration of enzyme levels and metabolite levels is clearly useful (Figure 3). In contrast, the utility of the mRNA and transcription factor layers was not evident. The study did not appear to improve or expand upon trans-omic network methodology described in the previous works. Finally, the various analyses (analyzing the trans-omic network for nodes with the highest degree centrality, the PPI analysis, and viewing the energy metabolism pathways in the network) provided disparate results that were only tenuously connected in the discussion section.

      Reviewer #2 (Evidence, reproducibility and clarity):

      Summary

      This manuscript integrates public transcriptomic, proteomic, and metabolomic datasets from ROSMAP DLPFC samples to construct a multi-layer metabolic trans-omic network in Alzheimer's disease. By linking transcription factors, enzyme mRNAs, proteins, metabolic reactions, and metabolites, the authors report coordinated downregulation of the TCA cycle, oxidative phosphorylation, and ketone body metabolism, along with mixed regulatory signals in glycolysis/gluconeogenesis. They interpret these patterns as indicative of broad energetic dysfunction and alterations in amino-acid/nitrogen metabolism in AD. While the framework is conceptually appealing, much of the analysis remains descriptive, and several biological interpretations extend beyond what the data can robustly support. The reliance on bulk tissue without accounting for cell-type composition, limited covariate adjustment, and the absence of validation or sensitivity analyses reduce confidence in the mechanistic conclusions. Overall, the study provides a preliminary systems-level overview, but additional rigor is needed before the proposed trans-omic regulatory insights can be considered convincing.

      Major Comments

      (1) Interpretation requires more cautious phrasing, and validation is essential. The manuscript frequently asserts that specific pathways are "inhibited" or that energetic deficits are "compensated," but these conclusions extend beyond what the descriptive, bulk-level data can support. Because no metabolic flux, causality, or direct functional measurements are included, the results should be framed as putative regulatory shifts, not confirmed impairments. Critically, key claims about pathway inhibition would require flux modeling, perturbation analyses, or experimental validation to be convincing. Without such validation, the mechanistic interpretations remain speculative.

      We thank the reviewer for this crucial comment. We fully agree that, given the descriptive and bulk-level nature of our analysis, mechanistic interpretations must be made with caution. In the absence of direct metabolic flux measurements or experimental validation, our findings should be interpreted as putative regulatory shifts rather than confirmed functional impairments. Accordingly, we have revised the manuscript to temper mechanistic claims. We have replaced definitive statements with more speculative phrasing (e.g., “Our analysis revealed a putative coordinated downregulation …” instead of “Our analysis revealed a coordinated downregulation …” in Abstract section; “we demonstrate the systems-level view of the potential dysregulated energy production …” instead of “we demonstrate the systems-level view of the dysregulated energy production …” in pg. 10, lines 25-26).

      (2) Although the authors acknowledge this in the limitations, bulk-level differences may primarily reflect altered proportions of neurons, astrocytes, microglia, and oligodendrocytes rather than true within-cell-type regulation. Incorporating a cell-type deconvolution or performing a sensitivity analysis would substantially improve interpretability. This issue also impacts the trans-omic network: if the molecules included originate from different cell types, the inferred regulatory relationships may not reflect true intracellular processes.

      We appreciate the reviewer’s point that bulk-level differences can reflect altered proportions of different brain cell types, subsequently affecting the inferred trans-omic network analysis. To assess the changes in cell type proportions of the samples that we used in our study, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglias, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two group. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      (3) Differential analysis covariates. For the differential expression analyses, only gender and PMI were included as covariates. Additional variables, such as age at death, RIN, neuropathological measures, and comorbidities, can strongly influence molecular profiles and should be considered to ensure that the observed differences reflect AD-related biology rather than confounding pathological or technical factors.

      We appreciate the reviewer’s comment regarding the included covariates in differential analyses of our study. The reason we did not include other variables, including age at death and RIN, is that these data for each sample were not available. Thus, we referred to original research articles from which proteomic or metabolomic datasets used in our study were derived. Regarding the metabolomic dataset, in the original article (Batra et al., 2023), only two metabolites, 1-methyl-5-imidazoleacetate and N6-carboxymethyllysine, were significantly associated with age. In addition, no metabolites were significantly associated with sex, BMI, or education. Regarding the proteomic dataset, in the original article, age at death, PMI, and sex were included as covariates in the analyses, though these variables were not found to strongly influence the data (Extended Data Fig.2 in (Johnson et al., 2020)).

      (4) Network stability and sample non-overlap. Proteomic, transcriptomic, and metabolomic data come from partially overlapping individuals. The authors should test whether the reconstructed network is robust to: different significance thresholds, restricting analyses to overlapping samples and alternative definitions of AD vs control.

      We appreciate the reviewer’s comment for the trans-omic network stability. In our study, the number of individuals for whom all omic modalities were measured was relatively small (n=25 in CT and n=35 in AD). This limited overlap reduces statistical power and can affect the downstream network construction. We have acknowledged this limitation in the revised manuscript and clarified that the reconstructed networks should be interpreted with caution regarding reproducibility and generalizability (pg. 13, lines 13-23).

      Minor Comments

      (1) Some TF enrichment and regulatory inferences lack explicit mention of multiple-testing correction.

      We apologize for the lack of clarity in our original description. We have corrected for multiple-testing for the TF inference. Thus, we have revised the Methods section to explicitly describe the correction method used and the threshold applied (pg. 23, lines 23-24).

      (2) The limitations section is strong but should explicitly discuss the influence of postmortem interval on metabolite levels.

      We appreciate the reviewer’s comment about the effect of postmortem interval on changes in metabolite levels. Accordingly, we have added the description of this perspective in our revised manuscript (pg. 13, lines 1-5).

      Reviewer #2 (Significance):

      The study extends a trans-omic integration framework, originally applied to metabolic disease, into the context of Alzheimer's pathology. Although the biological findings largely confirm known alterations in mitochondrial and energy metabolism, the network-based approach offers a structured way to view cross-layer regulatory changes. Its main advance is conceptual rather than biological, providing a unified framework rather than uncovering fundamentally new mechanisms. This work will primarily interest researchers in neurodegeneration and systems biology, as well as computational groups developing multi-omics integration methods.

      Reviewer #3 (Evidence, reproducibility and clarity):

      This study leverages existing transcriptomic, metabalomic and proteomic datasets from prefrontal cortex (PFC) to assess metabolic dysregulation in Alzheimer's disease (AD). They found a downregulation of multiple metabolic pathways, including TCA cycle, oxidative phosphorylation, and ketone metabolism, that may explain bioenergetic alterations in AD.

      The study used matching ROSMAP omics datasets from the DLPFC that have allowed more robust data integration. However, the datasets are all generated using bulk tissue, which makes data interpretation difficult. For example, the AD changes they observed may be due to shifts in cell type proportion with disease (e.g. cell death, neuron inflammation). Did the authors account for any potential shifts in cell type proportion in their analysis?

      If the assumption is that the changes in AD are cell intrinsic, which cell types are likely to be impacted? Can the authors integrate any existing single-cell analysis to infer which cell types may be driving the signals they detect, and whether this accounts for some of the antagonistic regulatory effects that were detected?

      We thank the reviewer for their insightful comments. We agree that the use of bulk tissue datasets cannot account for cell-type heterogeneity. As noted in our Limitations section (pg. 12, lines 24-27), we recognize that previous studies have found that the Braak stage is correlated positively with microglia and astrocyte proportions and negatively with oligodendrocyte proportion (Hannon et al., 2024; Shireby et al., 2022). Regarding the integration of single-cell analysis, we have referenced recent snRNA-seq findings (Mathys et al., 2024) in our Limitations section (pg. 12, lines 28-32) to deconvolve our bulk signatures.

      Furthermore, in our revised manuscript, we additionally used public single-cell transcriptomic datasets, which were obtained from DLPFC tissue of 465 subjects in the ROSMAP cohort (Green et al., 2024). For each omic data that we used in our analyses, we matched the same subjects and calculated the following cell type proportions, astrocytes, excitatory neurons, inhibitory neurons, microglia, oligodendrocytes, and OPCs. Then, we statistically compared the cell type proportions between control subjects and patients with AD (Fig. S3). In the transcriptomic data, we confirmed that the proportion of inhibitory neurons in the AD group was smaller than in the CT group, and that the proportion of oligodendrocytes in the AD group was larger than in the CT group. In the proteomic data, we did not observe any statistically significant changes in the cell type proportion between the two groups. In the metabolomic data, we found that the proportion of inhibitory neurons in the AD group was smaller than in the CT group (pg. 6, lines 8-11).

      Reviewer #3 (Significance):

      The manuscript provides multimodal insight into metabolic dysregulation in AD in the PFC. Given that metabolic dysfunction is likely to play a major in disease pathogenesis, this is a study of importance. However, the findings lack granularity at the cell type level, which limits the impact of the study.

      Reference

      (1) Baloni, P., Arnold, M., Buitrago, L., Nho, K., Moreno, H., Huynh, K., Brauner, B., Louie, G., Kueider-Paisley, A., Suhre, K., Saykin, A. J., Ekroos, K., Meikle, P. J., Hood, L., Price, N. D., Alzheimer’s Disease Metabolomics Consortium, Doraiswamy, P. M., Funk, C. C., Hernández, A. I., … Kaddurah-Daouk, R. (2022). Multi-Omic analyses characterize the ceramide/sphingomyelin pathway as a therapeutic target in Alzheimer’s disease. Communications Biology, 5(1), 1074.

      (2) Baloni, P., Funk, C. C., Yan, J., Yurkovich, J. T., Kueider-Paisley, A., Nho, K., Heinken, A., Jia, W., Mahmoudiandehkordi, S., Louie, G., Saykin, A. J., Arnold, M., Kastenmüller, G., Griffiths, W. J., Thiele, I., Alzheimer’s Disease Metabolomics Consortium, Kaddurah-Daouk, R., & Price, N. D. (2020). Metabolic Network Analysis Reveals Altered Bile Acid Synthesis and Metabolism in Alzheimer’s Disease. Cell Reports. Medicine, 1(8), 100138.

      (3) Batra, R., Arnold, M., Wörheide, M. A., Allen, M., Wang, X., Blach, C., Levey, A. I., Seyfried, N. T., Ertekin-Taner, N., Bennett, D. A., Kastenmüller, G., Kaddurah-Daouk, R. F., Krumsiek, J., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2023). The landscape of metabolic brain alterations in Alzheimer’s disease. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(3), 980–998.

      (4) Batra, R., Krumsiek, J., Wang, X., Allen, M., Blach, C., Kastenmüller, G., Arnold, M., Ertekin-Taner, N., Kaddurah-Daouk, R., & Alzheimer’s Disease Metabolomics Consortium (ADMC). (2024). Comparative brain metabolomics reveals shared and distinct metabolic alterations in Alzheimer’s disease and progressive supranuclear palsy. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 20(12), 8294–8307.

      (5) Cahill, K. M., Huo, Z., Tseng, G. C., Logan, R. W., & Seney, M. L. (2018). Improved identification of concordant and discordant gene expression signatures using an updated rank-rank hypergeometric overlap approach. Scientific Reports, 8(1), 9588.

      (6) Fröhlich, A. S., Gerstner, N., Gagliardi, M., Ködel, M., Yusupov, N., Matosin, N., Czamara, D., Sauer, S., Roeh, S., Murek, V., Chatzinakos, C., Daskalakis, N. P., Knauer-Arloth, J., Ziller, M. J., & Binder, E. B. (2024). Single-nucleus transcriptomic profiling of human orbitofrontal cortex reveals convergent effects of aging and psychiatric disease. Nature Neuroscience, 27(10), 2021–2032.

      (7) Green, G. S., Fujita, M., Yang, H.-S., Taga, M., Cain, A., McCabe, C., Comandante-Lou, N., White, C. C., Schmidtner, A. K., Zeng, L., Sigalov, A., Wang, Y., Regev, A., Klein, H.-U., Menon, V., Bennett, D. A., Habib, N., & De Jager, P. L. (2024). Cellular communities reveal trajectories of brain ageing and Alzheimer’s disease. Nature, 633(8030), 634–645.

      (8) Hannon, E., Dempster, E. L., Davies, J. P., Chioza, B., Blake, G. E. T., Burrage, J., Policicchio, S., Franklin, A., Walker, E. M., Bamford, R. A., Schalkwyk, L. C., & Mill, J. (2024). Quantifying the proportion of different cell types in the human cortex using DNA methylation profiles. BMC Biology, 22(1), 17.

      (9) Johnson, E. C. B., Carter, E. K., Dammer, E. B., Duong, D. M., Gerasimov, E. S., Liu, Y., Liu, J., Betarbet, R., Ping, L., Yin, L., Serrano, G. E., Beach, T. G., Peng, J., De Jager, P. L., Haroutunian, V., Zhang, B., Gaiteri, C., Bennett, D. A., Gearing, M., … Seyfried, N. T. (2022). Large-scale deep multi-layer analysis of Alzheimer’s disease brain reveals strong proteomic disease-related changes not observed at the RNA level. Nature Neuroscience, 25(2), 213–225.

      (10) Johnson, E. C. B., Dammer, E. B., Duong, D. M., Ping, L., Zhou, M., Yin, L., Higginbotham, L. A., Guajardo, A., White, B., Troncoso, J. C., Thambisetty, M., Montine, T. J., Lee, E. B., Trojanowski, J. Q., Beach, T. G., Reiman, E. M., Haroutunian, V., Wang, M., Schadt, E., … Seyfried, N. T. (2020). Large-scale proteomic analysis of Alzheimer’s disease brain and cerebrospinal fluid reveals early changes in energy metabolism associated with microglia and astrocyte activation. Nature Medicine, 26(5), 769–780.

      (11) Maitra, M., Mitsuhashi, H., Rahimian, R., Chawla, A., Yang, J., Fiori, L. M., Davoli, M. A., Perlman, K., Aouabed, Z., Mash, D. C., Suderman, M., Mechawar, N., Turecki, G., & Nagy, C. (2023). Cell type specific transcriptomic differences in depression show similar patterns between males and females but implicate distinct cell types and genes. Nature Communications, 14(1), 2912.

      (12) Mathys, H., Boix, C. A., Akay, L. A., Xia, Z., Davila-Velderrain, J., Ng, A. P., Jiang, X., Abdelhady, G., Galani, K., Mantero, J., Band, N., James, B. T., Babu, S., Galiana-Melendez, F., Louderback, K., Prokopenko, D., Tanzi, R. E., Bennett, D. A., Tsai, L.-H., & Kellis, M. (2024). Single-cell multiregion dissection of Alzheimer’s disease. Nature, 632(8026), 858–868.

      (13) Novotny, B. C., Fernandez, M. V., Wang, C., Budde, J. P., Bergmann, K., Eteleeb, A. M., Bradley, J., Webster, C., Ebl, C., Norton, J., Gentsch, J., Dube, U., Wang, F., Morris, J. C., Bateman, R. J., Perrin, R. J., McDade, E., Xiong, C., Chhatwal, J., … Harari, O. (2023). Metabolomic and lipidomic signatures in autosomal dominant and late-onset Alzheimer’s disease brains. Alzheimer’s & Dementia: The Journal of the Alzheimer’s Association, 19(5), 1785–1799.

      (14) Plaisier, S. B., Taschereau, R., Wong, J. A., & Graeber, T. G. (2010). Rank-rank hypergeometric overlap: identification of statistically significant overlap between gene-expression signatures. Nucleic Acids Research, 38(17), e169.

      (15) Shireby, G., Dempster, E. L., Policicchio, S., Smith, R. G., Pishva, E., Chioza, B., Davies, J. P., Burrage, J., Lunnon, K., Seiler Vellame, D., Love, S., Thomas, A., Brookes, K., Morgan, K., Francis, P., Hannon, E., & Mill, J. (2022). DNA methylation signatures of Alzheimer’s disease neuropathology in the cortex are primarily driven by variation in non-neuronal cell-types. Nature Communications, 13(1), 5620.

      (16) Tasaki, S., Xu, J., Avey, D. R., Johnson, L., Petyuk, V. A., Dawe, R. J., Bennett, D. A., Wang, Y., & Gaiteri, C. (2022). Inferring protein expression changes from mRNA in Alzheimer’s dementia using deep neural networks. Nature Communications, 13(1), 655.

      (17) Varma, V. R., Wang, Y., An, Y., Varma, S., Bilgel, M., Doshi, J., Legido-Quigley, C., Delgado, J. C., Oommen, A. M., Roberts, J. A., Wong, D. F., Davatzikos, C., Resnick, S. M., Troncoso, J. C., Pletnikova, O., O’Brien, R., Hak, E., Baak, B. N., Pfeiffer, R., … Thambisetty, M. (2021). Bile acid synthesis, modulation, and dementia: A metabolomic, transcriptomic, and pharmacoepidemiologic study. PLoS Medicine, 18(5), e1003615.

      (18) Wan, Y.-W., Al-Ouran, R., Mangleburg, C. G., Perumal, T. M., Lee, T. V., Allison, K., Swarup, V., Funk, C. C., Gaiteri, C., Allen, M., Wang, M., Neuner, S. M., Kaczorowski, C. C., Philip, V. M., Howell, G. R., Martini-Stoica, H., Zheng, H., Mei, H., Zhong, X., … Logsdon, B. A. (2020). Meta-Analysis of the Alzheimer’s Disease Human Brain Transcriptome and Functional Dissection in Mouse Models. Cell Reports, 32(2), 107908.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Summary

      Large language models (LLMs) have been developed rapidly in recent years and are already contributing to progress across scientific fields. The manuscript tries to address a specific question: whether LLMs can accurately infer signaling networks from gene lists. However, the evaluation is inadequate due to four major weaknesses described below. Despite these limitations, the authors conclude that current general-purpose LLMs lack adequate accuracy, which is already widely recognized. Its key contribution should instead be to provide concrete recommendations for the development of specialized LLMs for this task, which is completely absent. Developing such specific LLMs would be highly valuable, as they could substantially reduce the time required by researchers to analyze signaling networks.

      Strengths

      The manuscript raises a good question: whether current LLMs can accurately generate signaling networks from gene lists.

      Weaknesses:

      (1) The authors evaluate LLM performance using only three signaling networks: "hypertrophy", "fibroblast", and "mechanosignaling". Given the large number of well-established signaling pathways available, this is not a comprehensive assessment. Moreover, the analysis need not be restricted to signaling networks. Other network types, including metabolic and transcriptional regulatory networks, are already accessible in well-known databases such as KEGG, Reactome, BioCyc, WikiPathways, and Pathway Commons. Including these additional networks would substantially strengthen the evaluation.

      We agree with the reviewer that our evaluation of LLM performance is not comprehensive of all signaling networks, and that the benchmarking was previously limited to signaling networks. The purpose of this study is to benchmark LLMs against peer-reviewed computational models that make testable predictions and are highly validated experimentally, of which these three signaling networks are strong examples. KEGG, Reactome, Biocyc, WikiPathways, and Pathway Commons are databases that contain collections of individual reactions or pathways but are not computational models themselves and have not been experimentally validated in that sense.

      While this study focuses primarily on signaling networks, we agree that it would be useful to evaluate how well LLMs perform in generating networks of another type, for which predictive and validated computational models are available. Therefore, in new Figure 3 we now test the ability of LLMs to reconstruct the E. Coli core metabolic network, as well as test its ability to predict growth on metabolic substrates using flux balance analysis. We find that Claude Opus 4.6, GPT 5.2 Pro, and Gemini 3 Pro Preview perform well at reconstructing reactions from E. Coli core metabolism, but these reactions are not sufficient to predict growth on a variety of substrates. Given this expansion of scope, we replaced “signaling networks” in the title with “biochemical networks”.

      (2) In LLM evaluation, the authors use the gene lists that exactly match those in their "ground truth" networks, thereby fixing the set of nodes and evaluating only the predicted edges. However, in practical research, the relevant genes or nodes are not fully known. A more realistic assessment would therefore include gene lists with both genes present in the ground-truth network and additional genes absent from it, to evaluate the ability of the LLM to exclude irrelevant genes.

      We agree with the reviewer that evaluating the capacity of these LLMs to exclude additional genes is interesting. But because biological networks are always incompletely known, there is no “ground truth” of genes absent from a given network. Therefore, for the most rigorous benchmarking against a “ground truth”, we examine the positive predictive ability of LLMs. However, in response to this comment and point 3 below, we further examine additional measures of performance that include “false positives”.

      (3) The authors report only the recall/sensitivity of the LLM, without assessing specificity. In practical applications, if an LLM generates a large number of incorrect interactions that greatly exceed the correct ones, researchers may be misled or may lose confidence in the LLM output. Therefore, a comprehensive evaluation must include both sensitivity and specificity. Furthermore, it would be informative to check whether some of the "false positives" might in fact represent biologically plausible interactions that are absent from the manually curated "ground truth". Manually generated "ground truth" can overlook genuine interactions, and the ability of LLMs to recover such missing edges could be particularly valuable. This may even represent one of the most important potential contributions of LLMs.

      We agree with the reviewer that additional metrics could inform the evaluation of LLM performance. Therefore, as recommended, we calculated sensitivity, specificity, precision, negative predictive value (NPV), accuracy, and F1 score for each of the network models (hypertrophy, fibroblast, and mechanosignaling). These new results are summarized in confusion matrices shown in a new Supplementary Figures 3, 4, and 5. We performed this additional benchmarking across the 10 replicates for each LLM.

      One limitation of this approach is the substantial class imbalance within these confusion matrices. Because we interpret actual negatives as connections that are not found between any nodes of the ground truth models, there will be >10x more true negatives than any of the other classes. This makes specificity, accuracy, and the negative predictive value less informative.

      To illustrate this point, consider the ground-truth hypertrophy network which contains 191 connections between 106 nodes. The total number of possible connections between any two nodes is 106<sup>2</sup> = 11,236. Given that there are 191 actual positives, that leaves 11,045, actual negatives (as illustrated in the null predictor of Supplementary Figure 3A). As the number of node-to-node connections predicted by LLMs is on the order of a few hundred, the number of true negatives is always in the thousands, often outweighing the true positives in the specificity or accuracy calculations.

      The effects of this class imbalance are illustrated with the results of the “null predictors” in Supplementary Figure 4A, which are hypothetical models that fail to predict any connection between nodes (have predicted positive values of 0). These null predictors have sensitivities of 0 and specificities of 1 and high accuracies and NPVs because of the high true negative rates.

      The precision and F1 scores calculated using these confusion matrices are robust to these class imbalances. Indeed, there is substantial heterogeneity in the number of “false positives” connections generated by the LLMs as illustrated by the precision and F1 scores. We are hesitant to unequivocally label these novel, predicted connections as false positives because, as the reviewer points out, these connections could represent true molecular interactions that were undiscovered at the time of the ground truth models’ conception but have since been demonstrated experimentally and published.

      (4) It is widely known that applying differential equation models to highly complex biological networks, such as the three networks in the manuscript, is meaningless, because these systems involve a large number of parameters whose values can drastically alter the results. As Richard Feynman once said: "with four parameters I can fit an elephant, and with five I can make him wiggle his trunk." Thus, the evaluation of LLMs on "logic-based differential equation models" does not make much sense.

      Differential equation models have been the primary mathematical framework for studying complex systems for decades. We refer readers unfamiliar with differential equations to the Nobel prize-winning work of Hodgkin/Huxley (Physiology or Medicine1963, action potential of neurons), Prigogine (Chemistry 1977, non-equilibrium thermodynamics and pattern formation), John Nash (Economics 1994, game theory and Nash Equilibrium), Merton/Scholes (Economics 1997, dynamics of financial derivatives), and Manabe/Hasselmann (Physics 2021, dynamic modeling of atmosphere and oceans).

      We are confused by the quote of a joke by Richard Feynman about fitting equations to data in the shape of an elephant. While this famous joke is amusing, it is both misattributed (it was a recollection by Enrico Fermi in 1953 of a joke once made by John von Neumann) and deliberately hyperbolic (see https://en.wikipedia.org/wiki/Von_Neumann%27s_elephant). Regardless, the relevance of this joke to our study is unclear, because we are not fitting equations to data. As described in the text, in previous studies we validated the predictions of these three logic-based network models with experimental data that was not used to develop the models.

      Reviewer #1 (Recommendations for the authors):

      (1) All figures are in very poor resolution.

      Thank you for identifying this. We have fixed this issue, which was due to SVG embedding. We now embed as higher resolution PNG and provide full resolution files separately.

      (2) The manuscript does not include data availability or code availability.

      As described in the Methods, all code and data is now available via GitHub (https://github.com/saucermanlab/LLM-network-generation).

      Reviewer #2 (Public review):

      (1) Information on the accuracy of directionality of interaction would help understand if there is a bias towards either a positive or negative association.

      To assess if LLMs are biased in predicting either stimulatory or inhibitory connections, we examined the proportion of stimulatory and inhibitory connections for each ground truth model along with the prediction sets from the different LLMs (Author response table 1). These findings suggest that any directional bias is minimal and not conserved across the different ground truth models. These tables were not included in the revised manuscript.

      Author response table 1.

      Proportion of stimulatory and inhibitory connections present in each ground truth model and in the sets of connections predicted by each LLM (Claude, GPT, and Gemini).

      The primary subset of reaction types that the LLM’s tend to miss are often downstream, cell-type specific, and/or gene regulatory connection (e.g. Figure 1B). This observation is consistent with the fact that the ground truth models were constructed using experimental evidence from specific publications involving defined experimental models and cell types whereas the LLM’s presumably draw from the entire corpus of published literature.

      (2) Do all LLMs capture similar information, or are some LLMs better at capturing certain information than others? Further to this, it would be interesting to look into whether amalgamating information across all three LLMs results in a more accurate network.

      The reviewer asks interesting questions that can be qualitatively answered in Figure 1B and in the network visualizations (Supplementary Figures 1 and 2). These diagrams illustrate predicted connections that are shared between the nodes. To include a more quantitative, comprehensive assessment of this overlap, we have included Venn diagrams showing the extent to which LLMs capture shared information (Supplementary Figures 3-5). One such Venn diagram for the hypertrophy model is included in Supplementary Figure 3C. Indeed, it seems that there is substantial overlap in the “false positive” connections predicted by Claude and GPT. It could be interesting to evaluate if these connections are reflective of newly discovered molecular interactions, as referenced in our response to reviewer one point 3.

      While amalgamating information across all three LLMs might generate a more accurate network, these Venn diagrams illustrate that there remain connections within the ground truth models that are not represented in any of the prediction sets from the LLMs. Indeed, the union of all predicted connections for the hypertrophy network made by any of the 10 replicates from the different LLMs would still lack 33 ground truth connections (Supplemental Figure 3C).

      (3) Would it be possible to retrieve a confidence value of the interactions from the LLMs and conduct Precision, recall rate, AUPR and calibration analyses? These metrics would also help with the benchmarking process.

      We agree with the reviewer that including additional evaluative metrics is instructive. Therefore, for each reference network, we have included precision, specificity, recall, accuracy, and F1 score calculations (Supplementary Figures 3-5). Additionally, we conducted calibration analyses to illustrate the LLM’s reliability. While this latter analysis is interesting, we note that the calibration curves were generated using the frequency of predictions among the 10 replicates as a proxy for confidence. This frequency is highly sensitive to the temperature parameter of the LLMs. Different temperature settings could substantially influence the trajectories of these confidence curves and the distributions of the associated histograms. See Supplemental Figure 3B for calibration analyses of the hypertrophy model.

      We considered performing analyses resembling AUPR as suggested by the reviewer, but we did see a defensible way to vary a “threshold” for the classification of a connection as positive or negative. In this study, a connection predicted by the LLM either does or does not exist within the ground truth mode.

      Typos:

      (1) Generated networks to predict THE CLASSIC "fetal gene program" gene expression.

      Thank you for catching this error. We have corrected it.

      (2) A manually curated network has A functional accuracy of

      Thank you for catching this error. We have corrected it.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states, but also enabled the real time monitoring of protein conformational changes precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results and conclusions supported by the experimental data. Authors have determined the conformational dynamics of TmrAB across different ATP concentrations including physiological ones and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies. Authors have also mentioned limitations in the study.

      Comments on revised version.

      Authors have worked on most of the revisions stated in previous feedback and included in the newer version, which has been significantly improved. Other comments have been described to be out of scope from this study.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATPbound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I had three major concerns with the original version, all of which have been addressed by the authors in this revised version.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      I mentioned that the final section of the Results part seems like an afterthought, especially since the heading suggests a broader scope.

      Reply: We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      The changes made to the section do not align with the wording of the reply. Please consider modifying it further.

      We appreciate the positive feedback and this final comment. We have revised the final section of the Results to better reflect the scope indicated by the heading. In addition to clarifying the kinetic analysis, we now explicitly relate our kinetic observations to previously determined thermodynamic measurements, showing that the rapid interconversion of ATP-bound conformations observed during steady-state turnover is consistent with a thermodynamic landscape characterized by a near-zero free-energy difference and entropy–enthalpy compensation. This revision more clearly integrates the kinetic and thermodynamic aspects of the transporter cycle.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript aims to test the idea that visual recognition (of faces) is hierarchically organized in the human ventral occipital-temporal cortex (VOTC). The paper proposes that if VOTC has a hierarchical organization, this should be seen in two independent features of the VOTC signal. First, hierarchy assumes that signals along the hierarchy increase in representational complexity. Second, hierarchy assumes a progressive increase in the onset time of the earliest neural response at each level of the hierarchy. To test these predictions, the authors extract high-frequency broadband signals from iEEG electrodes in a very large sample of patients (N=140). They find that face selectivity in these signals is distributed across the VOTC with increasing posterior-anterior face selectivity, hence providing evidence for the first prediction. However, they also find broadband activity to occur concurrently, therefore challenging the view of a serial hierarchy.

      Strengths:

      (1) The hypothesis (that VOTC is hierarchically organized) and predictions (that hierarchy predicts increases in representational complexity and increases in onset time) were clearly described.

      (2) The number of subjects sampled (140) is extremely large for iEEG studies that typically involve <10 subjects. Also, 444 face selective recording contacts provide a very nice sampling of the areas of interest.

      We would like to thank this reviewer for their positive comments and evaluation of our manuscript.

      Weaknesses:

      (1) A control analysis where areas have known differences in response onset should be performed to increase confidence that the proposed analyses would reveal expected results when a difference in response onset was present across areas. From Figure 3, it can be seen that many electrodes are placed in earlier visual areas (V1-V3) that have previously been shown to have earlier broadband responses to visual images compared to VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018). The same analyses as in Figures 4 and 5 should be used comparing VOTC to early visual areas to confirm that the analyses would detect that V1-V3 have earlier onsets compared to VOTC.

      First, we would like to mention that the analyses performed in our paper are commonly accepted analyses to extract time-domain information.

      Yet, the reviewer is right that, considering our claim, evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies would provide further support for our claims. The solution proposed by the reviewer is interesting but the number of face-selective recording contacts/sites in posterior occipital cortex (colored disks in figure 3) is too small for any meaningful comparison. Moreover, while absolute responses to visual images should indeed emerge earlier in early visual cortex than association cortex, it may not be the case for category-selective responses to visual images (which is what our claim is about).

      To address the reviewer’s concern, we used face-selective responses from regions known to have different response onset latencies: occipital and posterior temporal lobe vs. medial temporal lobe structures (e.g., Mormann et al., 2008, https://pmc.ncbi.nlm.nih.gov/articles/PMC2676868/). Waveforms and onset latencies for these regions (OCC, PTL, MTL) are shown in Author response image 1. Despite the small number of contacts showing significant face-selective activity in the MTL (N=20) and the lower SNR in this region, the onset latency differences between OCC/PTL and MTL are significant using all 4 methods of latency estimation (see methods in the main manuscript), with medium to large effect sizes. This was despite noisy latency estimates for the MTL (in particular, the ‘delta slope’ method could not be used to get meaningful Cohen’s d when comparing MTL to OCC). Latency estimates are also slightly higher for the ‘% of peak’ method than in the manuscript, as we estimated the latency at 25% of the peak (instead of 20% in the manuscript), again to allow meaningful estimations for the MTL.

      Author response image 1.

      In addition to this, we performed a simulation analysis where we statistically compared the measured PTL signals to ATL signals that have been artificially, incrementally, shifted forward in time. Author response image 2 shows (top row) the measured onset latencies differences between the 2 regions (PTL minus ATL) estimated using 4 different approaches as a function of the temporal shift applied to ATL, as well as the associated p-values (bottom row). As shown in Author response image 2, the ‘original’ unshifted data yields no significant difference between the 2 regions. The difference however becomes significant when ATL signals is shifted forward in time 10 to 30 ms, depending on the method used.

      Author response image 2.

      These 2 observations provide evidence that our approach does indeed allow revealing ‘true’ differences in onset latencies, further supporting our claims.

      Last, as also suggested by reviewer 2, we conducted a thorough equivalence testing using ROPE and Bayesian factor to support the lack of differences between regions. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from early visual cortex than OCC). These analyses, now reported in the result section of the revised manuscript (Table 1), largely confirm the hypothesis of concurrent onset latencies across VOTC.

      (2) It is unclear why correlating mean timeseries helps understand how much variance is shared between regions (Figure 4). Any variance between images is lost when averaging time series across all images, and this metric thus overestimates the variance shared between areas. Moreover, the finding that correlating time domain signals across VOTC areas does not differ from correlating signals within an area could be driven by this averaging. For example, if the same analysis was done on electrodes in left and right V1 when half of the images had contrast in the left hemifield and the other half had contrast in the right hemifield, the average signals may correlate extremely well, while this correlation falls apart on a trial-by-trial basis. These analyses therefore need to be evaluated on a trial-by-trial basis.

      This is an interesting comment. We agree that variance between images is lost when averaging time-series across all images. However, to use the reviewer’s analogy, in order to support the claim that left and right V1 would show the same onset times and time-course (i.e., no hemispheric lateralization) for lateralized presentations (= the same kind of claim that we make in our paper), it’s the average response across images that should be compared, not a correlation run on a trial-by-trial basis (which would indeed falls apart because of a lack of response in the ipsilateral V1).

      Moreover, we would like to emphasize that the goal of this analysis in our paper is not to make claims about the variance shared between regions. In fact, this is not a key analysis in our paper, the outcome of which is not strictly necessary for the main argument made. Finally, if we were to perform a (time-consuming) image-by-image analysis in our study, correlations would be weak due to low signal-to-noise ratio (each face image appears only 1.6 times per stimulation sequence on average) and the fact that each face image appears after a different non-face image across presentations.

      (3) Previous studies on visual processing in VOTC have shown that evoked potentials are more predictive of the onset of visual stimuli than broadband activity (e.g. Miller et al., 2016, PLOS CB, https://doi.org/10.1371/journal.pcbi.1004660). Testing the prediction from a hierarchical representation that signals along the VOTC increase in onset time should therefore include an evaluation of evoked potential onsets in addition to broadband signals.

      We have used HFB responses in our study as these signals tend to be easier to characterize in the time domain than evoked potentials, and they are more local given their reduced SNR compared to evoked potentials (Jacques et al., 2022; https://pmc.ncbi.nlm.nih.gov/articles/PMC9457683/). Moreover we have previously shown highly correlated time courses across HFB and low frequency evoked potentials in the same paradigm (Jacques et al., 2022, eLife).

      Yet, to address this reviewer’s concern, we replicated the main analyses on low-frequency event-related potential signal, identifying contacts exhibiting significant face-selective responses in the same manner as in Jacques et al (2022). Namely, we start from bipolar-referenced sequences of recording corresponding to the full visual stimulation sequences (~70 s). For each recorded intracerebral contact, we average sequences in the time-domain, crop the average to contain an integer number of face frequency (1.2 Hz) cycles, run an FFT on the cropped sequences and identify the significant contacts with a Z-score procedure identical to that used for HFB signals. We then notch-filter out the visual response (6 Hz and harmonics) from the full length sequences, extract short epochs from the filtered sequences around the onset of each face image, average across epoch for each recording contact, subtract the mean signal measured in the baseline (-0.166 to 0 s relative to face onset) and take the absolute value (to be able to average across contacts despite differences of morphology and polarity). Significant contacts are then subjected to the same analyses as for the HFB signal.

      Results from these analyses are presented as supplementary material (Figure S9, Table S4) in the revised manuscript (referenced in lines of the main manuscript). While we were not able to obtain reliable latency estimates using the z-score method with the same parameters as for HFB signal, these analyses with ERP signal indicate similar onset latencies for ERPs as for HFB activity and largely replicate observations made with HFB. In particular, onset latencies were in a very similar range (~100 to 140 ms) with similar patterns across regions or along VOTC and between-region signal correlations. There were also a few significant face-selective responses over posterior ventro-medial occipital cortex, likely overlapping ‘early visual cortex’ (V1,V2v,V3v,hV4), probably due to limited low-level contributions in this paradigm (see Or et al., 2019, JOV; https://jov.arvojournals.org/article.aspx?articleid=2734585). Over these regions, onset latency was systematically earlier (up to 40ms) than in slightly more anterior regions, (i.e. anterior to -80 mm) where very little variability in onset latency was found up to the ATL region. We have acknowledged this in the revised manuscript.

      (4) Testing the second prediction, that the onset time of processing increases along the VOTC posterior to anterior path, is difficult using the iEEG broadband signal, because from a signal processing perspective, broadband signals are inherently temporally inaccurate, given that they are filtered. Any filtering in the signal introduces a certain level of temporal smoothing. The manuscript should clearly describe the level of temporal smoothing for the filter settings used.

      The reviewer is right that HFB signals are temporally smoothed, potentially yielding slightly underestimated onset latencies. However, our time-frequency analyses parameters ensured a minimal degree of smoothing. In fact, the original submission already contained a description of the expected temporal smoothing resulting from the wavelet transform. This is what we wrote in the original submission: “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had 20 ms of full width at half maximum (FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin), ensuring that onset timing information is accurate up to 10 ms (i.e half of the FWHM).”

      In the revised manuscript we further elaborate as follows:

      “The number of cycles (i.e., central frequency) of the wavelet was adapted as a function of frequency from 2 cycles at the lowest frequency to 9 cycles at the highest frequency. The temporal smoothing resulting from the wavelet transform was minimal: wavelets had a temporal spread of 20 ms (full width at half maximum - FWHM) across the frequency range (i.e. median of FWHM computed at each frequency bin). A simulation of HFB signals with a constant abrupt onset time and realistic signal-to-noise ratio indicates that the potential underestimation of onset latency due to the wavelet analysis is around 5-10 ms, which is on par with the value of the half width at half maximum (= FWHM/2 = 20/2 ms).”

      Author response image 3 displays simulated HFB signal (using identical wavelet parameters than in our manuscript) in an ideal scenario with a response starting at 150 ms in all trials (N=150 trials), reaching maximum 10 ms later. This provides a theoretical estimate of the slight underestimation of onset latency due to the wavelet transform. It shows onset latency estimates are at most 12 ms underestimation of true onset time.

      Author response image 3.

      That being said, given the physiological noise in the data, the fact that the response onset likely varies slightly from trial to trial, with a variable slope in activity increase, these wavelet parameters (within a certain margin) have likely little influence on the actual latency estimation.

      (5) The onsets of neural activity in VOTC are surprisingly early: around 80-100 ms. This is earlier than what has previously been reported. For example, the cited Quian Quiroga et al. (2023) found single neuron responses to have the earlier onset around 125 ms (their Figure 3). Similarly, the cited Jacques et al., 2016b and Kadipasaoglu et al., 2017 papers also observe broadband onsets in VOTC after 100 ms. Understanding the temporal smoothing in the broadband signal, as well as showing that typical evoked potentials have latencies compared to other work, would increase confidence that latencies are not underestimated due to factors in the analysis pipeline.

      In the revised manuscript, as suggested by reviewer 2, we have modified the data resampling strategy (using hierarchical bootstrap and permutation test that respects the nested structure of the data) to estimate onset latencies, confidence interval and permutation tests. Moreover, since the absolute onset latency estimates depend on the methods used, we now provide estimates using 4 different methods. The overall absolute onset latencies differ slightly across the 4 methods but all median onset latencies vary between 95 ms and 130 ms, which is similar to what was reported in some of the participants in Kadipasaoglu et al., 2017 (https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0188834; note that latencies reported at the level of single sites or individual are usually higher due to lower signal-to-noise ratio). Moreover, these latencies for face-selective responses are actually very similar to those measured for non-selective/absolute responses to visual stimuli (e.g. Jacques et al., 2016: around 90-100 ms for faces in FG; Yoshor et al., 2007: ~100 ms in posterior Fusiform Gyrus; Regev et al. 2018: 90-120ms in posterior to middle FG). Other studies, measuring non-selective responses using more ‘conservative’ onset detection methods, find slightly later onset latencies (e.g. Cao et al. 2025: 139 ms in posterior FG; Martin et al., 2019: ~150 ms in ventral occipital -VO- regions).

      Cao R, Zhang J, Zheng J, Wang Y, Brunner P, Willie JT, Wang S. 2025. A neural computational framework for face processing in the human temporal lobe. Current Biology. DOI: https://doi.org/10.1016/j.cub.2025.02.063

      Martin AB, Yang X, Saalmann YB, Wang L, Shestyuk A, Lin JJ, Parvizi J, Knight RT, Kastner S. 2019. Temporal dynamics and response modulation across the human visual system in a spatial attention task: An ECoG study. Journal of Neuroscience 39:333–352. DOI: https://doi.org/10.1523/JNEUROSCI.1889-18.2018, PMID: 30459219

      Regev TI, Winawer J, Gerber EM, Knight RT, Deouell LY. 2018. Human posterior parietal cortex responds to visual stimuli as early as peristriate occipital cortex. European Journal of Neuroscience 48:3567–3582. DOI: https://doi.org/10.1111/ejn.14164, PMID: 30240547

      Yoshor D, Bosking WH, Ghose GM, Maunsell JHR. 2007. Receptive fields in human visual cortex mapped with surface electrodes. Cerebral Cortex 17:2293–2302. DOI: https://doi.org/10.1093/cercor/bhl138, PMID: 17172632

      As an important note, in the revised manuscript, we have removed data from 3 recording contacts from 1 participant that were located in the upper bank of the Calcarine Sulcus, which is actually outside of our VOTC region of interest. The 3 contacts being located in dorsal V1 or V2 were showing very early responses and were biasing our latency estimates for the OCC region.

      (6) Understanding the extent to which neural processing in the VOTC is hierarchical is essential for building models of vision that capture processing in the human brain, and the data provides novel insight into these processes.

      For additional context, a schematic figure of the hierarchical view and a more parallel system described in the paragraph on models of visual recognition (lines 553) would help the reader interpret and understand the implications of the paper.

      Our observations in the current study clearly indicates concurrent face-selective processing in the VOTC, which is incompatible with a serial hierarchical model. While we discuss how such concurrent activity could be implemented in the cortex (e.g., via direct input from ‘early visual cortex’ to different VOTC face-selective clusters), our data do not allow to provide more evidence in that respect to what already exists in the literature. Moreover, we are not providing data regarding connectivity (feedforward or re-entrant) either between face-selective regions or between these regions and ‘early visual cortex’.

      Author response image 4 shows a very simplified versions of standard hierarchical/serial versus concurrent/parallel models.

      Author response image 4.

      Reviewer #2 (Public review):

      Summary:

      This very ambitious project addresses one of the core questions in visual processing related to the underlying anatomical and functional architecture. Using a large sample of rare and high-quality EEG recordings in humans, the authors assess whether face-selectivity is organised along a posterior-anterior gradient, with selectivity and timing increasing from posterior to anterior regions. The evidence suggests that it is the case for selectivity, but the data are more mixed about the temporal organisation, which the authors use to conclude that the classic temporal hierarchy described in textbooks might be questioned, at least when it comes to face processing.

      Strengths:

      A huge amount of work went into collecting this highly valuable dataset of rare intracranial EEG recordings in humans. The data alone are valuable, assuming they are shared in an easily accessible and documented format. Currently, the OSF repository linked in the article is empty, so no assessment of the data can be made. The topic is important, and a key question in the field is addressed. The EEG methodology is strong, relying on a well-established and high SNR SSVEP method. The method is particularly well-suited to clinical populations, leading to interpretable data in a few minutes of recordings. The authors have attempted to quantify the data in many different ways and provided various estimates of selectivity and timing, with matching measures of uncertainty. Non-parametric confidence intervals and comparisons are provided. Collectively, the various analyses and rich illustrations provide superficially convincing evidence in favour of the conclusions.

      We thank the reviewer for their positive comments on our manuscript.

      Weaknesses:

      (1) The work was not pre-registered, and there is no sample size justification, whether for participants or trials/sequences. So a statistical reviewer should assess the sensitivity of the analyses to different approaches.

      Pre-registration of fundamental research in a clinical context is quite uncommon for intracranial data, owing, for instance to the time needed to accumulate data, or to the uncertainty of cortical sampling location in a given participant. Nevertheless, in the current study, sample size is much higher than in typical intracranial studies (usually 5-20 participants), in fact much higher than most typical Cognitive Neuroscience research. The same is true for the number of recording contacts (>10000 site here), and the number of trials considered for analysis. Each participant had a minimum of 164 face trials and an average of 262 trials (i.e. an average of 3.2 stimulation sequences of 82 trials), which is higher than most standard human electrophysiological studies.

      In the revised manuscript, to unsure that we have the maximum available power, and because our hypothesis is independent of hemisphere, we collapsed data across hemispheres for all analyses. We nevertheless provide analyses split by hemispheres as supplementary material.

      In addition, since onset latency estimations depend on the methods used, we now report onset latencies from 4 different methods (2 statistical and 2 non-statistical).

      (2) Frequentist NHST is used to claim lack of effects, which is inappropriate, see for instance:

      Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: A guide to misinterpretations. European Journal of Epidemiology, 31(4), 337-350. https://doi.org/10.1007/s10654-016-0149-3

      Rouder, J. N., Morey, R. D., Verhagen, J., Province, J. M., & Wagenmakers, E.-J. (2016). Is There a Free Lunch in Inference? Topics in Cognitive Science, 8(3), 520-547. https://doi.org/10.1111/tops.12214

      Please see reply to the next comment (3).

      (3) In the frequentist realm, demonstrating similar effects between groups requires equivalence testing, with bounds (minimum effect sizes of interest) that should be pre-registered:

      Campbell, H., & Gustafson, P. (2024). The Bayes factor, HDI-ROPE, and frequentist equivalence tests can all be reverse engineered-Almost exactly-From one another: Reply to Linde et al. (2021). Psychological Methods, 29(3), 613-623. https://doi.org/10.1037/met0000507

      Riesthuis, P. (2024). Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science, 7(2), 25152459241240722. https://doi.org/10.1177/25152459241240722

      We thank the reviewer for pointing this out. In the revised manuscript we conduct and report a thorough examination of equivalence using ROPE and Bayesian factor to support the lack of differences between regions. We did not use TOST procedures as these have low power and require huge samples be meaningful (Riesthuis, 2024). Instead we relied on Bayes factor and descriptive proportion in ROPE. Equivalence bounds and region of practical equivalence (ROPE) were defined to account for physiological variability corresponding to a small effect (Cohen’s d = 0.199, i.e. standard in equivalence testing) and axonal conduction delays between regions (i.e. ATL is further away from ‘early visual cortex’ than OCC). These analyses, now reported in the result section of the revised manuscript, along with effect sizes, confirm the hypothesis of concurrent onset latencies across VOTC.

      Riesthuis P. 2024. Simulation-Based Power Analyses for the Smallest Effect Size of Interest: A Confidence-Interval Approach for Minimum-Effect and Equivalence Testing. Advances in Methods and Practices in Psychological Science 7.

      Detailed methods are reported as well:

      “In addition, to statistically assert whether onset latencies measured across main VOTC regions (OCC, PTL, ATL) were consistent with a concurrent (parallel) face-selective activation, we used Bayesian equivalence testing, relying on two separate metrics: (1) the percentage of differences in region of practical equivalence (ROPE), and (2) the Bayes factor using Cauchy prior. Equivalence bounds for ROPE were defined by combining two components: (1) a component of physiological variability and (2) a component reflecting expected delays in response onset between regions attributed to neural conduction delay, given the differential distances separating early visual cortex (EVC) from posterior face-selective regions (e.g. IOG) vs. anterior regions (ATL) and assuming signal mainly travels between VOTC regions through major postero-anterior axis fiber bundles of the Inferior longitudinal fasciculus (ILF) or the inferior fronto-occipital fasciculus (IFOF). Physiological variability corresponded to expected measurement noise and between-subject variability. Physiological equivalence bound was obtained by multiplying a Cohen’s d of 0.2 (conventional threshold for a negligible effect) with the (pooled) between-subjects variability in onset latency (computed for each region using a jackknife procedure and a correction factor of [N-participants – 1] to the jackknife standard deviation). For conduction delay bounds, the expected latency difference between regions under parallel activation depends on: (1) the distance between each region (considering the estimated origin of the signal is the same for all regions - EVC), and (2) neural conduction velocity. For inter-region distance we determined, for each region, the 5 and 95 percentile of the Talairach y-coordinate distribution and defined the maximal distance bounds as the distance between the y-coordinate corresponding to 5% of region 1 (e.g. OCC) to the coordinate corresponding 95% of region 2 (e.g. PTL). This resulted in the following maximum distance values: OCC-PTL: [54]mm; PTL-ATL: [54]mm; OCC-ATL: [83] mm. For conduction velocity, we used a constant value of 3.5 m/s, based on median axonal conduction velocity for cortico-cortical connections (Lemarachal et al., 2022; Van Blooijs et al., 2023), which was more conservative than using a range of values (e.g. 1.7 to 5.3 m/s based on Lemarachal et al., 2022). Maximum expected conduction delay was computed as [maximum distance / conduction speed] (e.g. for OCC to PTL: 0.054 / 3.5 = 15 ms). Resulting equivalence bounds (ROPE) were asymmetrical given than one region (e.g. OCC) is always closer to the source (EVC) than the other region (e.g. PTL) and were defined as [-1*physiologial_bound +1*physiologial_bound+max_conduction_delay]. For instance, using the z-score method to measured onset latencies, ROPE was [-12 to 27] ms for OCC to PTL, meaning that under equivalence, OCC can be activated up to 27 ms earlier than PTL (maximum conduction delay + noise), while allowing for some instances where OCC activates later (up to 12ms, due to noise only).

      For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (4) The lack of consideration for sample sizes, the lack of pre-registration, and the lack of a method to support the null (a cornerstone of this project to demonstrate equivalence onsets between areas), suggest that the work is exploratory. This is a strength: we need rich datasets to explore, test tools and generate new hypotheses. I strongly recommend embracing the exploration philosophy, and removing all inferential statistics: instead, provide even more detailed graphical representations (include onset distributions) and share the data immediately with all the pre-processing and analysis code.

      Data will be shared upon publication of the manuscript (see OSF repository in https://osf.io/2qzym). While we agree the dataset is large and could be explored in many ways, we do not consider the current study to be exploratory in nature. While our measurements could have turned out to clearly support hierarchical processing in human VOTC, our point in this manuscript is that the evidence derived from this large dataset unequivocally points instead toward concurrent activation of face-selective regions along the VOTC from IOG to antFG+ (i.e. a ~90 mm portion of cortex), with potential small variability accounted for by variability in axonal conduction velocity, signal-to-noise ratio or simple physiological variability. Other likely sources of variability such as type and density/size of fiber bundles across regions cannot easily be modeled with the current data set.

      (5) Even if the work was pre-registered, it would be very difficult to calculate p-values conditional on all the uncertainty around the number of participants, the number of contacts and the number of trials, as they are random variables, and sampling distributions of key inferences should be integrated over these unknown sources of variability. The difficulty of calculating/interpreting p-values that are conditional on so many pre-processing stages and sources of uncertainty is traditionally swept under the rug, but nevertheless well documented:

      Kruschke, J.K. (2013) Bayesian estimation supersedes the t test. J Exp Psychol Gen, 142, 573-603. https://pubmed.ncbi.nlm.nih.gov/22774788/

      Wagenmakers, E.-J. (2007). A practical solution to the pervasive problems of p values. Psychonomic Bulletin & Review, 14(5), 779-804. https://doi.org/10.3758/BF03194105 https://link.springer.com/article/10.3758/BF03194105

      All analyses and preprocessing stages are identical between regions and the number of trials is large enough not to be a constraining factor. As indicated above and below, we now report detailed equivalence testing and effect sizes and recomputed all statistics, taking into account the structure of the data as suggested by the reviewer.

      (6) Currently, there is no convincing evidence in the article to clearly support the main claims.

      Bootstrap confidence intervals were used to provide measures of uncertainty. However, the bootstrapping did not take the structure of the data into account, collapsing across important dependencies in that nested structure: participants > hemispheres > contacts > conditions > trials.

      Ignoring data dependencies and the uncertainty from trials could lead to a distorted CI. Sampling contacts with replacement is inappropriate because it breaks the structure of the data, mixing degrees of freedom across different levels of analysis. The key rule of the bootstrap is to follow the data acquisition process, and therefore, sampling participants with replacement should come first. In a hierarchical bootstrap, the process can be repeated at nested levels, so that for each resampled participant, then contacts are resampled (if treated as a random variable), then trials/sequences are resampled, keeping paired measurements together (hemispheres, and typically contacts in a standard EEG experiment with fixed montage). The same hierarchical resampling should be applied to all measurements and inferences to capture all sources of variability. Selectivity and timing should be quantified at each contact after resampling of trials/sequences before integrating across hemispheres and participants using appropriate and justified summary measures.

      The authors already recognise part of the problem, as they provide within-participant analyses. This is a very good step, inasmuch as it addresses the issue of mixing-up degrees of freedom across levels, but unfortunately these analyses are plagued with small sample sizes, making claims about the lack of differences even more problematic--classic lack of evidence == evidence of absence fallacy. In addition, there seem to be discrepancies between the mean and CI in some cases: 15 [-20, 20]; 8 [-24, 24].

      In light of the reviewer’s comment, we recomputed all timing analyses using a stratified hierarchical approach to evaluate confidence intervals (using bootstrapping), statistical comparisons (using permutation tests) and equivalence testing.

      This is what we wrote in the revised methods:

      “The first two timing parameters of face-selective response, onset and offset latencies, were quantified per main VOTC region using a hierarchical bootstrapping approach to respect the nested structure of the data (region > participants > contacts > trials). For each bootstrap iteration and each region, we first sampled participants with replacement. Within each sampled participant, we then sampled contacts and then trials within sampled contacts, with replacement. Resampled trials, then contacts within participants, then participants within a region, were successively averaged to obtain a bootstrapped region-level response from which we derived onset latency (4 different methods) and offset latency. We obtained bootstrap distributions of onsets/offsets using 2000 bootstrap iterations per region, allowing to compute the median and 95% confidence interval for these 2 parameters.”

      Then, later about permutation tests:

      “Statistical significance of latency differences between main VOTC regions was assessed using a hierarchical permutation test. We use a stratification approach to partition participants into a paired set (i.e. participants that had recording contacts in the two regions compared) and unpaired set (participants with contacts in a single region). For paired participants, the region labels were randomly swapped within subject (i.e., exchanging the signals from the two regions), thereby preserving all participant-, contact-, and trial-level structure while breaking the association between region and latency estimates. For unpaired participants, participants were randomly reassigned between regions while preserving the original group sizes, to generate pseudo-groups under the null hypothesis of no regional difference. In each permutation, signals were averaged across trials, then contacts, then participants within each permuted group and latency was computed and stored from the resulting region-level signals. We performed 10000 permutations to obtain a distribution of regional differences of latencies under the null hypothesis and determine the p-value as the fraction of the null distribution larger or smaller than the observed (non-permuted) difference.”

      And then about equivalence testing:

      “For each pair of region compared, we used a bootstrap procedure to (1) define the percentage of differences between regions that fall within the ROPE, (2) compute the bayes factor using a Cauchy distribution (scale = 0.5) to estimate the proportion of the prior distribution in ROPE, and the bootstrap distribution to estimate the proportion of the posterior distribution in ROPE. The bootstrap distribution was obtained using a hierarchical stratified bootstrap procedure that naturally respects the nested structure of the data, that accommodates for unequal numbers of participants, contacts, and trials across regions, as well as partially overlapping participants samples across regions.

      For each bootstrap iteration, with first sample participants with replacement within each stratum (i.e. paired vs unpaired participants samples). For paired participants, sampling was performed jointly across regions to preserve the dependency structure, whereas unpaired participants were sampled independently within each region. Within each sampled participant, contacts were then resampled with replacement, and within each contact, trials were resampled with replacement. For paired participants, trial resampling was performed using identical trials across sampled contacts with a participant to preserve trial-level covariance. Resampled trials were averaged at the contact level, contact-level signals were averaged within participant, and participant-level signals were averaged to obtain a region-level response. Onset latency was then estimated from this averaged signal for each region using one of the 4 methods defined above (‘HFB response timing parameters). This procedure was repeated across 2000 bootstrap iterations to obtain a distribution of latency estimates for each region that respects the structure of data at iteration-level. Latency differences between regions were computed at each iteration, yielding a bootstrap distribution of differences which was used to compute percentage of differences in ROPE and posterior distribution for the Bayes factor.”

      (7) Three other issues related to onsets:

      (a) FDR correction typically doesn't allow localisation claims, similarly to cluster inferences: Winkler, A. M., Taylor, P. A., Nichols, T. E., & Rorden, C. (2024). False Discovery Rate and Localizing Power (No. arXiv:2401.03554). arXiv. https://doi.org/10.48550/arXiv.2401.03554

      Rousselet, G. A. (2025). Using cluster-based permutation tests to estimate MEG/EEG onsets: How bad is it? European Journal of Neuroscience, 61(1), e16618. https://doi.org/10.1111/ejn.16618

      In fairness, we do not understand or share the reviewers’ concern here. Hundreds of fMRI or EEG studies use FDR or cluster tests to make inference about spatial or temporal location. We use FDR correction in one of the onset latency estimation method and only consider one-sided differences. Other methods in the revised manuscript do not use FDR correction.

      (b) Percentile bootstrap confidence intervals are inaccurate when applied to means. Alternatively, use a bootstrap-t method, or use the pb in conjunction with a robust measure of central tendency, such as a trimmed mean.

      Rousselet, G. A., Pernet, C. R., & Wilcox, R. R. (2021). The Percentile Bootstrap: A Primer With Step-by-Step Instructions in R. Advances in Methods and Practices in Psychological Science, 4(1), 2515245920911881.

      Again, we are not sure what the reviewer’s is referring to. The confidence intervals are computed on latency estimates from bootstrapped waveforms. In the revised manuscript, these waveforms are obtained by averaging (i.e. mean) resampled trials, resampled channels, resampled participants. A trimmed mean could not be applied in this condition, except perhaps when averaging across trials. But then the trimmed mean would have to be applied separately at each time sample which would disturbed within-, or between-trial, variability.

      (c) Defining onsets based on an arbitrary "at least 30 ms" rule is not recommended:

      Piai, V., Dahlslätt, K., & Maris, E. (2015). Statistically comparing EEG/MEG waveforms through successive significant univariate tests: How bad can it be? Psychophysiology, 52(3), 440-443. https://doi.org/10.1111/psyp.12335

      The rule of contiguous significant points is a heuristic that many researchers have used successfully to avoid spurious detection due to temporal autocorrelation. While we are aware that more sophisticated methods exist to correct for autocorrelation, such as cluster-based approaches, it is not directly usable since it requires comparing 2 conditions. The approach described in Piai et al., 2015 is interesting but incorrect as well since it relies on split-half simulations, which reduced signal-to-noise ratio, resulting in over estimated correction to be applied. In our revised manuscript, we rely on multiple methods to estimate onset latency, some of which not relying on this heuristic. Moreover, we apply plausible physiological constrains to our latency estimates, such as rejecting any onset before 40 ms after stimulus onset.

      (8) Figure 5 and matching analyses: There are much better tools than correlations to estimate connectivity and directionality. See for instance:

      Ince, R. A. A., Giordano, B. L., Kayser, C., Rousselet, G. A., Gross, J., & Schyns, P. G. (2017). A statistical framework for neuroimaging data analysis based on mutual information estimated via a Gaussian copula. Human Brain Mapping, 38(3), 1541-1573. https://doi.org/10.1002/hbm.23471

      (9) Pearson correlation is sensitive to other features of the data than an association, and is maximally sensitive to linear associations. Interpretation is difficult without seeing matching scatterplots and getting confirmation from alternative robust methods.

      We rely on Pearson correlation because this replicates the method used in Kadipasaoglu et al., 2017. It is also a widely accepted measure of (linear) relationship (in our situation we did expect linear or near linear relationships) in the literature. To address the reviewers concern, in the revised manuscript we nevertheless report, as supplementary material (Figure S10), the same functional connectivity analyses performed using the methodology and code provided in Ince et al. (2017). The results of this analyses are extremely similar to the results using Pearson’s coefficients.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 6, the response onset latencies are rendered in a smoothed manner on the brain surface. However, with this smoothing, variability between electrodes cannot be seen, and it would be better visualized in color, rendered in each electrode.

      Latency estimates computed at individual channels are noisy, which is why we do not report individual channels latencies but rather rely on averaging signals across contiguous channels, either across whole regions (Figure 4) or across smaller volumes as in Figure 6.

      (2) Onset latencies of 60 seconds seem extremely early compared to literature typically citing evoked responses with a latency of ~170ms. It would help if some additional sanity checks were shown, such as showing the latency of early visual responses. This would help with relative comparisons.

      In the revised manuscript the earliest median latency is 95 ms, which is in line with previous intracranial electrophysiology literature (e.g. Jacques et al., 2016; Jacques et al., 2022 ; https://pubmed.ncbi.nlm.nih.gov/26212070/; https://pubmed.ncbi.nlm.nih.gov/36074548/). The 99% confidence intervals can result in earlier latencies both due to some participants showing early responses and noise in latency estimates. Also please keep in mind that latency estimates are usually earlier when combining data across channels/participants compared to individual channels simply due to differences in SNR or across participants (see e.g. Kadipasaoglou et al., 2017).

      The reviewer indicates “…to literature typically citing evoked responses with a latency of ~170ms.”. We are assuming that they refer to the face-selective N170 ERP component measured on the scalp in EEG. Even with this ERP component, the face-selective response usually starts around 120-130 ms after stimulus onset (e.g. Rousselet et al., 2008; Jacques, Retter and Rossion, 2016; https://pubmed.ncbi.nlm.nih.gov/18831616/; https://pubmed.ncbi.nlm.nih.gov/27138205/) at scalp level. With the same highly sensitive paradigm as used here in EEG, we have systematically shown latency onsets of face-selective activity shortly after 100 ms (e.g., Retter et al., 2020; also Quek & Rossion, 2017) not accountable for by low-level visual cues (i.e., not present for phase-scrambled stimuli; Rossion et al., 2015; Or et al., 2019). Our latency onsets are also in line with spiking activity recorded with the same approach in the LatFG (Laurent et al., 2026) https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5955677

      (3) Figure 6B shows that the variability of latencies in the ATL is larger than the variability of latencies in the PTL. It would be helpful to evaluate whether, rather than in mean onset latencies, there is a change in variability in onset latency along the VOTC.

      This is an interesting point. However, it is difficult to evaluate since SNR is reduced in the ATL compared to OCC or PTL (Jacques et al., 2022). As a result, any measured modulations in the variability of onset latencies along VOTC may simply reflect changes in the precision of latency estimation driven by SNR variability.

      (4) Line 415 typo: 'there appears to be no delay' instead of 'there appear to be no delay'.

      We thank the reviewer for their careful reading of our manuscript. This has been corrected.

      (5) The discussion states in lines 455-457 that "a large proportion of neuronal populations in anterior VOTC regions exhibiting similar activity to different face images independently of the context in which they appear". However, this claim about similar activity to different faces should be evaluated and tested at the single-trial level. In addition, it is not clear how context was varied in the experimental design.

      Context is variable because each face image appears directly after a different object image (or object images) in the sequence. We have shown also in previous studies with this paradigm in EEG that the time-course of face-selective responses is similar across base frequencies (3-15 Hz) unless the rate is too fast, and whether an orthogonal or explicit face categorization task is used (Retter et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32119982/). Note that we do not claim that activity is identical across images but similar – if it was not (largely) similar, the averaged response would be jittered and low.

      (6) Line 516-517: DTI does not provide evidence for whether connectivity is direct or not, and what the directionality of connectivity is between two areas. This sentence should therefore state "..., suggest independent connections between early visual cortex and face-selective regions...".

      This has been rephrased.

      (7) Line 555 in the discussion, the definition of low-level visual should be expanded to include other early visual areas that have been demonstrated to respond earlier than VOTC (e.g. Martin et al., 2019, JNeurosci https://doi.org/10.1523/JNEUROSCI.1889-18.2018), to avoid the suggestion that V1 directly projects synaptically to all of VOTC (e.g. Markov et al., 2014, Cerebral Cortex, https://doi.org/10.1093/cercor/bhs270).

      We are not proposing that V1 directly projects directly/synaptically to all of VOTC, i.e., without other low-level retinotoptic areas involved; only that face-selectivity in the association cortex is not organized hierarchically. We have revised this sentence.

      (8) Line 585, for the sentence: "with temporal synchrony strengthening their connections", evidence or citations should be provided.

      Citations have been provided.

      (9) It is not clear what is meant in the paragraph starting in line 571: do the authors suggest that top-down signals are not necessary for fast recognition of clear views of faces, or additionally argue that these top-down signals are not necessary for detecting ambiguous or degraded inputs as faces?

      Exactly: That top-down (i.e., descending) signals may contribute but would not be necessary for fast recognition of clear views of faces AND for detecting ambiguous or degraded inputs as faces.

      (10) No statement was provided on data or code availability.

      Data will be made available on a repository upon publication (see https://osf.io/2qzym).

      Reviewer #2 (Recommendations for the authors):

      (1) FDR correction: which one? Please provide a reference.

      We now provide a reference, both in the results and methods: Benjamini and Hochberg, 1995.

      (2) In the introduction, this statement is too strong: "arguably the most familiar and ecologically valid stimulus". It is unclear how static 2D representations of faces are the most familiar and valid stimuli. Could you rephrase this? What about other very familiar stimuli like letters, words and biological motion?

      This statement is not about static 2D images of faces, but faces in general (in their natural environment). We do consider human faces (in general, not restricted to laboratory context) to be indeed the most familiar and ecologically important stimulus, both from an ontogenetic and phylogenetic perspective, unlike written material.

      (3) About the questioning of a strict temporal hierarchy, this EEG reference comes to mind: Foxe, J. J., & Simpson, G. V. (2002). Flow of activation from V1 to the frontal cortex in humans. Experimental Brain Research, 142(1), 139-150. https://doi.org/10.1007/s00221-001-0906-7

      As confirmed by Foxe et al. ’s (2002) paper to which the reviewer is referring to, there is indeed ample evidence that areas in the dorsal stream or frontal cortex (e.g. FEF) are activated very soon after V1 and before many ventral stream regions (e.g. Lamme and Roelfsema, 2000; https://pubmed.ncbi.nlm.nih.gov/11074267/). While Foxe et al.’s 2002 is highly valuable, it can hardly be compared with our current study which looks specifically into ventral stream areas which are largely indistinguishable using scalp EEG as in Foxe et al. ’s paper.

      (4) Regarding statistical significance, there is no such thing as a "trend". The threshold for a trend should have been pre-registered and applied to both sides of the magical boundary, for instance, with matching conclusions for a "trend toward non-significance (p=0.04)". P values near 0.05 provide weak support against the null. I would suggest leaving it at that. Nothing special happens at 0.05.

      This no longer appears in the revised manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study provides evidence for our understanding of HIV transmission dynamics by age and sex in Zambia during the PopART trial; by combining phylogenetic and individual-based mathematical modelling (IBM), it adds depth to the epidemiological literature and may inform more strategic allocation of HIV prevention resources in sub-Saharan Africa. The authors employ two complementary and well-established methodologies (phylogenetics and IBM), and this dual approach is a notable strength. However, the evidence supporting key conclusions is incomplete, with several claims insufficiently substantiated by the data presented. Improvements in data presentation (e.g., quantification of qualitative statements, statistical estimates, and clearer description of results) would substantially strengthen the paper.

      We thank the editor and reviewers for their positive comments. We have revised the manuscript in response to the points raised, as described below.

      First of all, we would like to summarise what we have changed regarding the presentation of summary statistics throughout the text. We agree that many of the statements in the original submission tended towards being qualitative. This was the result of shying away from presenting two separate estimates, with different ways of quantifying uncertainty, in the text. The phylogenetics could be presented as mean and confidence interval, while the IBM would need some measure of centrality (mean or median) and the highest density interval for a summary statistic (e.g. the mean age gap) as it varies over the posterior. These are not directly comparable. We have now changed this to present both where appropriate, with cautionary note about the difference between the CIs and HDIs (lines 257-260).

      We also were somewhat arbitrary regarding where we chose to summarise the posterior in the IBM or look at the best-fitting single simulation, and where we presented the mean as opposed to the median. We have done a considerable overhaul of what is presented in this revision:

      (1) We always present the posterior summary unless the level of detail is such that summarising uncertainty over the posterior is not feasible (e.g. in figures 3, 4 and 5). In the latter case we still use the best-fitting IBM replicate.

      (2) In the main text we always present the mean. For the phylogenetics the summary statistics are mean and confidence interval. For the IBM this is the posterior mean, and 95% HDI, of the mean of a particular statistic as calculated in each of the 1000 IBM replicates. For example, each replicate will have its own distribution of male source ages which have a mean value. These means also vary over the posterior, and a mean of them is calculated, as well as the HDI interval to represent posterior uncertainty. This “mean of means” may be a slightly confusing piece of terminology at first glance, but it allows us to properly capture posterior uncertainty in a way we mostly avoided in the first submission.

      One result of 1) above is a change to figure 6. It is now summarised over the posterior, with the result that time trends that were previously not evident become clear. This changes our conclusions slightly (lines 529-537) but it should be noted that the magnitudes of the trends remain small.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This manuscript describes the results of phylogenetic and epidemiological modeling of the PopART community cohorts in Zambia. The current manuscript draft is methodologically strong, but needs revision to strengthen the take-home messages. As written, there are many possible take-away conclusions. For example, the agreement between IBM and phylogenetic analysis is noteworthy and provides a methodological focus. The revealed age patterns of transmission could be a focus. The effects of the PopART intervention and the consequences of a 1-year disruption could be a focus. It is important, though, that any main messages summarized by the authors are substantiated by the evidence provided and do not extrapolate beyond the data that have been generated. I recommend that the authors think deeply about what the most important, well-supported messages are and reframe the discussion and abstract accordingly.

      We have rewritten the abstract, and also made changes to the discussion in order to centre our message around the contribution of particular of demographic groups to transmission, and how, with that contribution revealed, such groups can be selected for specialised interventions.

      Strengths/weaknesses by section:

      (1) ABSTRACT

      The Abstract summarizes qualitative findings nicely, but the authors should incorporate quantitative results for all of the qualitative findings statements.

      The abstract in the revision is extensively revised, and contains quantitative estimates throughout, from both methodologies where appropriate.

      The ending claim is not substantiated by the modeling scenarios that have been run: "targeted interventions for demographic groups such as under-35 men may be the key to finally ending HIV." It is straightforward to run this specific scenario in the model to determine whether or not this is true.

      Our modelling framework is not set up to model the “last mile” of HIV elimination, notably as it has no component for MSM or FSW transmission, and we do not feel that we could confidently present results regarding it. As a result, this statement has been greatly softened in the new abstract (lines 75-78).

      The authors should add confidence intervals to the quantitative metrics, such as the 93.8% and 62.1% incidence reduction.

      These have been added.

      (2) RESULTS

      The authors should check the Results section for any qualitative claims not substantiated by the analyses performed, and ensure the corresponding analyses are presented to support the claims.

      The Results and Methods describe the model's implementation of the PopART intervention differently. The Methods describes it as including VMMC, TB, and STI services, while the Results only mentions intensified HIV testing and linkage.

      This is a slight misreading of the text. That paragraph in the Methods is describing the trial itself, not the modelling framework.

      A limitation of the model is that HIV disease progression is based on the ATHENA cohort in the Netherlands, which is a different HIV subtype (B) than the one in the research setting (C). The model should be configured using subtype C progression data, which have been published, or at least a sensitivity analysis should be conducted with respect to disease progression assumptions.

      The available literature does not suggest a significant difference in progression between subtypes B and C, and we have added text and citations to this effect (lines 699-701).

      In Table 2, the authors should consider adding a p-value to establish whether or not IBM and phylogenetics estimates are different.

      We have done this; the appropriate test was a posterior predictive check. See lines 261-263, 575-579 and 805-814.

      (3) DISCUSSION

      The literature review and comparison of study results to previously published phylogenetic studies is very nice. The authors could strengthen this by providing quantitative estimates with CIs for a more scientific comparison of the study results vs. prior studies, perhaps as a table or figure.

      We have expanded the discussion on this point (lines 504-527). We considered adding a table, but the existing literature that directly answers the questions we ask is quite limited and fragmentary. For example, Monod et al do not present a complete treatment of age gaps. The literature using regression analyses to identify predictors of HIV prevalence or incidence related to partner age is extensive, but those results are not directly comparable to ours.

      The authors state that due to "the narrow geographical catchment area... The results should not be automatically extrapolated to apply to other SSA settings." The authors should exercise this caution when comparing the results to studies in South Africa and elsewhere.

      We have made more explicit acknowledgements of these limitations (lines 598-600).

      There are many other limitations to the analysis, including some mentioned above, that are not acknowledged. The authors should think carefully about what the most important limitations are and acknowledge them honestly at the end of the Discussion section.

      The limitations paragraph has been revised (lines 598-605).

      Reviewer #2 (Public review):

      Summary:

      The authors analyzed PopART data to better characterize the age and sex-specific heterosexual HIV transmission dynamics in Zambia, with the goal of allocating resources.

      Strengths:

      Important analysis to hone in on the key driver of HIV transmission in Zambia, which hopefully can be used to tune prevention efforts to maximize effect while limiting required resources. Two analytic approaches were used, and while the phylogenetic data were markedly more limited, they mirrored the simulated epidemic. The authors did a nice job reviewing the limitations of the data and the analyses. The authors did a nice job of providing analyses to support their goals and hypothesis, and this work may have more impact now that resources in SSA for HIV prevention and treatment may become more scarce

      Weaknesses:

      To increase the impact and utility of this work, it would be helpful to parse the analysis just a bit further to estimate the roles of undiagnosed vs diagnosed and untreated subpopulations on this transmission. PopART is a multifaceted intervention, but the cost, effort, and approach to reengagement in care vs testing/treatment can be quite different.

      We have now provided stratified results by diagnosed and non-diagnosed status of the source, as well as an overall summary of the proportion of undiagnosed sources by age and sex. See lines 305-310, 539-547, and table 3.

      Recommendations for the authors:

      Reviewing Editor:

      We commend you for conducting a rigorous and comprehensive study titled "The age and sex dynamics of heterosexual HIV transmission in Zambia: an HPTN 071 (PopART) phylogenetic and modelling study" that significantly advances the understanding of HIV transmission dynamics in sub-Saharan Africa. The study utilizes an innovative dual-methodology approach integrating individual-based mathematical modelling (IBM) and pathogen phylogenetics to characterize heterosexual HIV transmission patterns by age and sex during the PopART trial in Zambia.

      This manuscript reports on HIV transmission dynamics in Zambia using data from the PopART study, combining individual-based modelling and phylogenetic analysis. The use of two independent methodologies enhances confidence in the consistency of the findings and enables robust cross-validation. The work addresses an important topic in HIV prevention, particularly in settings where resources may become more constrained, and offers insight into potential demographic targets for intervention.

      However, several aspects of the manuscript limit its current impact. The main take-home messages are diffuse and not clearly presented. Some conclusions in the abstract and discussion appear to go beyond the scope of the presented data. For instance, the claim that targeting under-35 men may be key to ending HIV is not directly tested in the modelling scenarios and should be reframed or removed unless supported by new analyses. Furthermore, important quantitative details, such as confidence intervals, p-values, and precise age group estimates, are lacking in key sections (e.g., the Abstract and Results).

      The authors are encouraged to clearly identify and communicate their central findings, ensure all claims are fully supported by their analyses, and make the data more accessible to readers by adding detailed, quantitative summaries where needed.

      The following are our recommendations to the Authors:

      (1) Clarify Study Objectives and Central Messages

      Reframe the abstract and discussion to highlight a clear, well-supported set of main findings.

      Avoid overgeneralized or unsubstantiated claims, especially those not directly tested by your model (e.g., the effectiveness of targeting under-35 men).

      As stated above, we have revised this text accordingly.

      (2) Support Qualitative Claims with Quantitative Data

      Provide numerical results, including effect sizes and confidence intervals, wherever qualitative trends are mentioned.

      For example, restate: "The largest gaps for female recipients were among the youngest" as "... in the age group XX-YY with OR = Z.Z (95% CI: A.A-B. B)."

      As mentioned at the top of the review, we have overhauled the treatment of summary statistics extensively, and now give confidence or highest density intervals throughout the text.

      (3) Improve the Results Section

      Check that all claims are supported by the analyses, and ensure figure references are accurate.

      The statements that went beyond what was supported, notably about ending the epidemic by targeting young men, have been removed. The typo in table references has been fixed.

      Annotate Figure 6 with trendline coefficients and p-values where applicable.

      The takeaway message of figure 6 has now changed and we no longer see no trend, just a minor one.

      Revise Figure 4 for clarity or consider replacing it with a tabular format.

      We would prefer to keep the current figure 4, as we have not found any clearer way to illustrate the patterns, which are the consequence of the phenomenon observed in figure 5. We have put more explicit descriptive text in the discussion, linking the two figures (lines 470-476).

      (4) Address Potential Bias and Model Assumptions More Rigorously

      Explain sampling bias in IBM and phylogenetics (e.g., how the 355 high-confidence phylogenetic pairs were selected).

      The reviewer comment regarding the 355 pairs was based on a misapprehension; we used all the pairs we found using the phyloscanner pipeline. There are no sampling bias issues involved in the IBM as every individual in the simulations is considered. Appendix 2 includes some sensitivity analysis results if the procedure used to find the 355 is changed.

      Discuss how the use of subtype B disease progression data from the ATHENA cohort may impact results in a subtype C setting. A sensitivity analysis would strengthen this.

      Subtype B progression data was used in the absence of any appropriate data from subtype C, but the literature does not suggest any major difference between the two (lines 699-701).

      (5) Include More Detail on Undiagnosed Populations and ART Effects

      Estimate the roles of undiagnosed and untreated subpopulations in driving transmission.

      As mentioned above, this analysis has been added.

      Clarify mechanistically how ART might influence age gaps in transmission dynamics.

      This now is clarified in the introduction (lines 127-129).

      (6) General Improvements

      Provide p-values where comparisons are made (e.g., in Table 2).

      Use consistent terminology and definitions across Methods and Results.

      Add more discussion on limitations, especially regarding generalizability to other SSA settings.

      All of these have been inserted as previously mentioned.

      By addressing these points, the manuscript would present a more coherent narrative and a stronger, evidence-based contribution to the field. We appreciate you all for your fantastic effort and hope you will reflect the feedback in your final paper.

      Reviewer #1 (Recommendations for the authors):

      Thank you for the opportunity to review this interesting manuscript.

      In the public review, I have recommended that the authors should incorporate quantitative results for all of the qualitative findings statements. As one example, I would recommend that "We found the largest gaps for female recipients were among the youngest of those recipients" is re-written as "The largest gaps for female recipients were in the age group XXX-YYY with OR=ZZZ (XXX-YYY)." such as odds ratios, and specific outcome definitions including ages. To give one more example: "immediate increase in the average age at transmission of both sources and recipients" could be rephrased as "increase in the average age at transmission by XXX (YYY-ZZZ) years for sources and XXX (YYY-ZZZ) for recipients over [TIME PERIOD]."

      We hope the revisions we have made to the statistical presentation are satisfactory as a response to this request.

      Again in the public review, I recommended checking the Results section for any qualitative claims not substantiated by the analyses performed, and ensuring the corresponding analyses are presented to support the claims. An example is: "Trends are minor or non-existent in the former two variables." - please annotate Figure 6 (assuming the authors meant to reference Figure 6 and not 7 here?) to show over what period trendlines were fit and provide the coefficient and CI. To support the stated claim even more strongly, a p-value might be apt with a null hypothesis of a slope of zero.

      Please check the numbering on all figure references in the text, as some appear to be misnumbered. E.g., where the text refers to Figure 7, I believe the authors meant to reference Figure 6.

      The change to how we handled the statistics has changed the message of figure 6 (which is now figure 7) and rendered this somewhat moot. We have checked that all figure and table references are now correct.

      Figure 3 is very nice, but if the axes were flipped on one panel, it would make them easier to compare, and then adding some statistics to assess whether the patterns are the same or different when a man vs woman is the source.

      We have flipped the axes here.

      Figure 4 was too complicated for me. I could not follow the Sankey flows because there is too much going on and overlapping. Consider revising to make it easier to digest... perhaps to table format?

      As mentioned above, we would prefer to keep this figure, but we have situated it better in the text.

      Reviewer #2 (Recommendations for the authors):

      A few points that would improve the clarity and the strength of the manuscript

      (1) There is a need to clarify more about how the IBM and phylogenetic data does not suffer from sampling bias. For e.g.,

      Line 205: What proportion of the transmissions modeled in the IBM from Zambia?

      All of them. We confined the analysis of the IBM to the Zambian communities from which phylogenetic data was acquired (lines 755-758).

      Line 217: What proportion of the phylogenetic pairs (cherries) suggesting transmission were the 355 that had high confidence in directionality. How do these pairs compare to the others

      There was no identification of “cherries” involved in picking these pairs; the phyloscanner procedure does not use that step. We confined our analysis solely to the pairs for which we did identify a direction of transmission; that is the 355. Appendix 2 includes a sensitivity analysis involving varying the parameters by which these were identified.

      (2) I appreciate the authors noting that MSM transmissions are unlikely to be playing a role in this cohort, as noted in previous work by the group. However, systematic undersampling of men is common in other study cohorts of HIV. While the MSM and heterosexual networks may be relatively distinct, undersampled men who are bridging the networks could impact the estimates. Can the authors use the time to diagnosis analysis (HIV phyloTSI) to estimate rates of undiagnosed men and women?

      We feel that this is beyond the scope of this work. The phylogenetics dataset in its totality could be used for this purpose (although it is probably highly biased towards undiagnosed individuals due to the considerable majority of samples coming from the healthcare facilities). However, we concentrate here solely on the subset involved in our probable transmission pairs, which is fairly small. Extending the scope to an exploration of the full dataset would seem like a separate study, which we do have plans to do.

      We have used the IBM for this question instead (lines 303-321), however, as MSM transmission was not modelled, it is also not ideal for answering this question. Ultimately we feel that the way these studies were implemented makes it an unsatisfactory tool for answering the MSM question, important as it is.

      (3) Expanding on the point above, in other settings, transmission to young men has been associated with partnerships with older men, and if these young men then transmitted to young women, would we see a similar effect as noted in these models (assuming the young men were less well sampled).

      Our previous work (Hall et al., 2024) suggested no excess of identified male-male pairs in the phylogenetics dataset which might suggest cryptic male-to-male transmission. The age disparities would be worth exploring had this been found, but is curtailed by the lack of it.

      (4) Related to the point above, is there an estimate of the populations (age and sex) that are undiagnosed in the IBM model? Can this be teased out... is transmission from men to women more likely 2/2 lack of diagnosis... or lack of engagement in care?

      We have explored results by diagnostic status as it pertains to age and sex, but we feel that moving on to a more general exploration of the role of diagnosis and lack of engagement in care is again going beyond the scope of what is already a long paper.

      (5) I'm still not fully clear as to why ART might affect age gaps. Can this be explained in more detail?

      See lines 127-129.

    1. Author response:

      We are pleased that the reviewers viewed the core demonstration (that Dscam mutually exclusive splicing is preserved in a vector and that exon alternates can be replaced with genes of interest) as a solid foundation for the system. We agree that the manuscript would be strengthened by clearer quantitative characterization of expression, additional controls for fluorophore imaging, improved presentation of the figures, and more precise wording about the current scope of evidence. In a revised manuscript, we plan to address these points by adding or clarifying quantitative expression analyses, including S2 cell validation data, adding appropriate imaging controls where available, revising claims about 12-transgene expression to distinguish design capacity from direct experimental demonstration, and improving figure labels and legends throughout.

      We also plan to expand the discussion of PXGS limitations, including cell-type dependence on Dscam splicing machinery, possible position effects, and gene size considerations. Finally, we will improve Methods reporting by adding resource identifiers, cell culture quality control information, statistical design details, and data/code availability statements where appropriate.

      We appreciate the opportunity to revise the manuscript and believe these changes will make the strengths and limitations of PXGS clearer to readers.

    1. Author response:

      Reviewer #1 (Public review):

      Strengths:

      Overall, the computational model tested in the paper is novel and interesting.

      The demixing framework represents an appealing hypothesis that deserves further investigation.

      The current paper provides new empirical data showing that the target stimuli with the same absolute noise level can be either repelled from or attracted to non-target items, depending on the relative noise levels. The observation that biases depend on the relative noise levels is by itself an interesting one, and is consistent with the prediction of the demixing model.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      While this manuscript contains interesting new experimental observations and theoretical ideas, it has several substantial problems in its current form, which limit the conclusions that can be drawn. The description of the computational model is too brief. The key modeling assumptions need to be better motivated and explained. As the computational models generate different predictions in different regimes, it is a bit difficult to evaluate how well the experimental data support the model at a more quantitative level. Also, the results focused on studying the biases in the behavior; it is unclear whether the model can fully explain the behavior data (such as error distributions or behavioral precision).

      We agree that the model description should be expanded and that quantitative agreement with the data should be assessed more thoroughly, and we plan to address this during the revision. In the initial version of the manuscript, we aimed to highlight the qualitative agreement of the data with the novel and counterintuitive predictions by the model. While the reviewer is correct that the model "generates different predictions in different regimes," the particular predictions we test (the interaction between the target noise level and the parity of the target and non-target noise levels in Experiments 1-3, and the effects of non-target noise when the target noise is held constant in Experiment 4) hold across regimes (Figure S1 shows this for the former prediction). We aim to further expand on this point in the revision.

      Major concerns:

      (1) Concerns/suggestions regarding the computational modeling

      The current paper seeks to test the predictions of the demixing-based computational model proposed in reference 22. There are several problems with the modeling component in the current paper.

      (1a) The description of the model is too brief and difficult to understand. Although the model was proposed in reference 22, it would still be beneficial to provide more details of the model so that readers can understand and appreciate the strengths/limitations of the model.

      The generative model and the inference procedure could be better explained to better link the model to the behavior. In particular, how was the observer's behavioral report in each trial modeled? This requires more explanation because currently the demixing procedure estimates four parameters for a given trial, yet for a given trial, only one behavioral report was produced (e.g., current Experiment 1), or two reports were produced sequentially (e.g., current Experiment 2).

      We will provide more details about the model and how it was fitted to the data. Please note that the model parameters were fitted per subject and condition, not per trial: 2 hue noise parameters,  and , corresponding to the noise of the target and non-target item across 4 noise combination conditions, plus a shared identifiability noise, , across conditions, determining the discriminability of the items along the identifying dimension. This strongly limits model flexibility as only 3 parameters (including the shared  across conditions) are used to create the bias curve for each subject in each condition.

      (1b) Key modeling assumptions need better justification.

      One such key assumption is that on a given trial, each stimulus triggers many samples (or approximately, an entire response distribution), rather than a single sample. This assumption deviates substantially from prior work on ideal observer models. It was not clear whether this assumption is realistic. For the type of stimuli used in the current experiments, perhaps one can argue that each pixel corresponds to one sample of brain activity, thus collectively each stimulus should trigger many samples of activity in the brain. If this were to be the case, it would have two implications. First, the noise parameter in the model should be directly related to the magnitude of the stimulus noise. Thus, one should be able to plug these experimentally-controlled parameter values into the model to directly generate predictions about the biases. Second, when using stimuli with no stimulus variability (e.g., simple grating stimuli), the predicted biases should change. However, it wasn't clear whether this would hold experimentally, i.e., using gratings would lead to different biases or no biases.

      If the variability of the samples for a given stimulus involves neural noise, it would be useful to justify why it is reasonable to consider that many samples were generated per stimulus.

      We are grateful to the reviewer for raising this point, and we will provide more details on it in the revision. In brief, we believe that it is the standard ideal observer assumption of one sample per trial that is unrealistic and works only in cases when there is a single signal source, so that the samples can be simplified to a single average. Consider that determining a stimulus value is a similar problem for an ideal observer to the one that a researcher who aims to decode neural data from populations of neurons (or fMRI voxels) has to solve. Different populations of neurons would provide responses that match different stimuli – in essence, creating different samples in an ideal observer framework. Thus, even without external noise, the demixing problem would be present when there is more than one stimulus, but internal noise is much more difficult to control, so in our experiments, we used multi-colored stimuli.

      (1c) As mentioned in (1b), the model assumes that on each trial, a large number of samples was generated. It would be useful to study and report how the prediction would change when the number of samples generated per stimulus is small. In particular, what happens when each stimulus only generates one measurement? This might be useful for interpreting previous experiment results with grating stimuli.

      This is an interesting point that we aim to address in the revision.

      (1d) Reference 22 studies how the predicted biases depend on the d-prime of the identifying dimension and found that the pattern of the biases varies substantially depending on the information available for the identifying dimension. However, the current paper didn't really discuss this important point. It is also unclear what parameters the authors used for the d-prime of the identifying dimension. Was it fitted directly to the data? The Methods section has some description on the "identifiability dimension", but it was a bit obscure.

      Intuitively, when the d-prime of the identifying dimension is very large, the demixing problem becomes irrelevant. In this case, there should not be any biases induced by demixing. In the case of the d-prime for the identifying dimension is 0, the problem should reduce to the simplified 1-d problem studied in reference 22. If my reading of reference 22 was correct, they reported different conclusions. It would be useful to clarify these points.

      We are grateful for the suggestion to expand the discussion of this point and will do so in the revision. The reviewer is correct that for very large d-prime in the identifying dimension, the demixing problem solution is trivial. However, the 2D case does not resolve to the 1D case when d-prime reaches zero. This is because the identifying dimension is still used to identify which item to report—unlike in the 1D case, when the reported dimension is the same as the identifying one. Consider what happens if the observer in our task does not remember at all which stimulus was left and which was right. It would report the other item in 50% of cases, leading to a strong attractive bias.

      In any case, the d-prime of the identifying dimension appears to be a key parameter. It would be great to constrain this parameter using the empirical data. When the d-prime of the identifying parameter is small, the observer would easily confuse the probed stimulus with the other stimulus in a given trial. This should lead to poor task performance. Thus, it may be possible to directly estimate the value of the d-prime of the identifying dimension based on the observer's performance, and then use this parameter to generate model predictions accordingly.

      We apologize for the confusion. We constrain the discriminability of items in the "identifying" dimension using the  parameter that determines the noise in that dimension for both items. The means in this dimension are fixed at an arbitrary value, as means and noise are interchangeable when considering discriminability. We will revise the description of the fitting procedure accordingly. Regarding the use of the same values in predictions, while possible, we prefer to keep predictions separate from fitting to avoid them becoming postdictions. The curves for the fitted model in Figure 2 already illustrate what the model predicts under the fitted parameter values.

      (1e) The current model assumes that a large number of samples are generated per stimulus and the brain can manipulate this information to perform the demixing task. It was well documented that visual working memory has a capacity limit (i.e., it can only hold information about a few items); this discrepancy needs to be clarified or addressed.

      We are grateful to the reviewer for raising this point, which we will address in the revised discussion. Briefly, we believe that the number of samples in the ideal observer model does not correspond directly to the working memory “slots”.

      (2) How well the computational model can explain the experimental data remains not entirely clear

      The authors show that there exists a parameter regime that can qualitatively explain the experimental finding. They also show that it is possible to fit the model to the data to explain the bias patterns. However, given that the model is flexible, it would be stronger if the authors could show that the same parameters that explain the biases could also explain other aspects of the behavior, for example, the magnitude of the errors.

      It would also help if the authors could report the best-fitted parameters from the experimental data. From these parameters, one can simulate synthetic data and apply the demixing model to see if the error distribution of the simulated observers is indeed similar to the experimentally measured error distribution. That way, one can check whether the fitted parameter explains the observer's behavioral performance beyond the biases.

      We are grateful to the reviewer for raising this point. We both agree and disagree with the reviewer here. The predictions reported come from an earlier paper describing the model (ref. 22). In our opinion, this represents a pure hypothesis-driven approach, where a prediction is formulated first and then tested with subsequently collected data. The model we test is normative, not descriptive; its goal is not to fit the data as closely as possible, but rather to make predictions about internal brain mechanisms. We do not suggest, for example, that demixing is the sole source of biases, so the resulting bias pattern might differ significantly from the predictions. That the model fits the data is, therefore, an additional bonus. At the same time, we agree that it is interesting to test whether the model can explain other parameters of the data. Note that our current fitting procedure was not geared toward this; we optimized the model to explain only the bias curve. In the revision, we aim to test whether the model can also explain the error variability.

      In other words, the model is not well constrained in the way it was tested in the paper. But it should be possible to improve it. First, if the noise parameter in the model is determined by the stimulus variability, one can determine it directly based on the external noise in the stimuli (discussed also in 1b) and see what prediction it leads to. Second, from the behavioral data, it may be possible to estimate the noise for the identifying dimension. Doing so will help better constrain the model.

      External noise accounts for only a portion of the total noise, as evidenced by behavioral errors. Even for a single item, the total noise consists of the amount of information the observer samples from the stimulus, the variability of these samples (external noise), and early (applied to each sample) and late (applied after integration) internal noise. Therefore, external noise alone might not constrain the model in the right regime. Regarding the identifying-dimension noise, as noted above, we do constrain it with the data. However, we aim to explore these points further in the revision.

      Other comments:

      (1) How does the model account for the swap errors? I am not sure I understood the way how the swap errors were treated in the paper. To me, substantial swap errors seem to be a consequence of having low d-prime values for the identifying dimension; that is, if there is only little information to discriminate the identity of the two stimuli, swap errors would be large. However, this possibility didn't seem to be mentioned in the paper.

      We apologize for the confusion. We will further clarify and perhaps reassess the treatment of swap errors in the revision. The model itself produces swap errors when the stimuli sources are misidentified.

      (2) Since the solution of the demixing problem was obtained using a numerical procedure based on EM. It would be useful to check whether the initialization has affected the biases obtained.

      Indeed, this is a valid point, and it's why we use a multi-initialization strategy. For each simulation of a single trial sample set (e.g., 100 random samples), we use a large number of initial points (50 in the initial submitted manuscript) to ensure the obtained EM solution is truly optimal. Additionally, we conduct a large number of simulated trials (10,000 for each parameter combination) to ensure the accuracy of the bias distribution we obtain.

      Reviewer #2 (Public review):

      Summary:

      This manuscript investigates the origins of inter-item biases in visual working memory. The authors proposed a computational model where overlapping memory signals are disentangled, inducing memory biases that depend on relative noise levels across items. The key theoretical advance is the prediction that bias direction depends not only on absolute memory noise but on the relative noise levels of target and non-target representations. Using four experiments with color mosaics whose color variability manipulates memory precision, the authors report that biases reverse as a function of relative noise in a manner predicted by the model.

      Strengths:

      The manuscript is clearly written and theoretically motivated. The experiments are well designed and provide converging evidence for a distinctive and non-intuitive prediction of the proposed model. I found the central result compelling: independently manipulating target and non-target noise leads to qualitatively different bias patterns, consistent with the model's prediction that relative noise is a key determinant of bias direction.

      We are grateful for the positive evaluation of the model and the empirical observations.

      Weaknesses:

      The main limitation is that the evidence establishes consistency of the data with the proposed Demixing Model, but does not demonstrate that the model provides a unique explanation of the data. Although the manuscript argues that dominant theories struggle to account for the observed reversals, no formal comparison with alternative computational frameworks is presented. In addition, model fitting results are reported only briefly, making it difficult to evaluate fit quality at the level of individual observers.

      We agree and we aim to provide a comparison with alternative models and an expanded description of the fitting results in the revision. Note, however, that the majority of existing models are descriptive, while we believe that as a normative model, the Demixing Model should be compared with other normative models, thus limiting the selection of competitors significantly.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      In this paper, the authors provide a systematic investigation of structural brain differences associated with congenital aphantasia (self-reported lifelong absence of voluntary visual imagery). Specifically, the authors analysed a structural neuroimaging dataset involving 18 individuals with aphantasia and 18 visualizers to test two competing hypotheses: (1) that aphantasia reflects alterations in visual pathways and early visual cortex, and (2) that it instead reflects differences in higher-order frontotemporal and cingulate systems. To test these hypotheses, the authors employed multiple analysis approaches (e.g., cortical morphometry, tractometry, graph-theoretic network analysis).

      They report structural differences between the two groups in frontotemporal and cingulate systems. In contrast, they found no reliable group differences in early visual cortex or major visual tracts. On this basis, they propose that aphantasia is primarily associated with differences in higher-order systems supporting integration and conscious access to internally generated representations, rather than with deficits in sensory visual representations themselves.

      Strengths:

      (1) The present work addresses an important gap in the mental imagery literature, providing a systematic investigation of structural neuroimaging differences in congenital aphantasia. By showing that structural differences between aphantasics and visualizers are mainly concentrated in frontotemporal and cingulate systems (rather than in visual cortex), it makes an important step toward a better understanding of individual differences in mental imagery and provides a set of candidate regions for future mechanistic work.

      (2) A key strength of the study is the multimodal approach employed to address the main research question, integrating tractometry, functional region-of-interest (fROI)-based tractography, graph-theoretic network analysis, and surface-based cortical morphometry, which provide a converging assessment of structural differences between aphantasics and visualizers.

      (3) The complementary use of Bayesian analyses alongside NHST to assess evidence for null results is a further strength of this work.

      Weaknesses:

      (1) A weakness of this work is related to aspects of the framing and, in particular, what can be confidently inferred from the results. The framing of existing accounts of aphantasia in the Introduction appears limited in that it reduces the views on aphantasia to two options (sensory strength account versus conscious access account) without acknowledging a third distinct position, namely that aphantasia reflects a specific deficit in the voluntary generation of imagery (Milton et al., 2021; Zeman et al., 2015, 2020; Whiteley, 2021; Cavedon-Taylor, 2022). Like the conscious access account, the view that aphantasia involves a deficit in the generation of sensory representation also speaks against the hypothesis of reduced sensory strength of internally generated representations. This third view could be acknowledged/discussed as it also maps quite well onto the presented results.

      (2) Relatedly, I think the main weakness of the paper concerns the interpretation of results being restricted to a lack of "conscious access". The paper frames its findings as mainly evidence for a conscious access failure, the view that visual representations are generated by aphantasics but cannot be consciously accessed. However, the structural findings are equally consistent with a voluntary generation failure, especially since the same higher-order regions examined can also be implicated in the top-down generation and control of imagery. The authors themselves initially define aphantasia as "lifelong absence of voluntary visual imagery". Given the nature of structural imaging data (as opposed to functional data), it is not possible with the present study to distinguish between a lack of generation versus a lack of conscious access. As such, examining this alternative interpretation appears appropriate, and it would considerably strengthen the paper. Structural MRI alone is not sufficient to dissociate imagery generation from conscious access, as these are fundamentally functional questions.

      (3) Some inconsistency and lack of clarity around the specific choice of regions/networks, which could be better motivated and explained. E.g., the "core imagery network" analysed in the white-matter connections analysis was derived from a previous 7T study (with which the sample partially overlaps) and is not necessarily the network most commonly associated with visual imagery in the literature (e.g., see Dijkstra et al., 2019; Pearson, 2019). It is, for instance, unclear why V1 was examined in the cortical thickness analysis but not in the previous one, given that both analyses are related to the visual pathway hypothesis. Related to this, in the graph-theoretic analysis, the rationale for network selection is inconsistently established in the Introduction. The attention and salience networks do have some grounding in the Introduction through the mention of specific regions such as FEF and anterior insula, though these are discussed as individual regions rather than as networks. However, the default mode network receives no motivation in the Introduction. More explicit elaboration on these choices would be appropriate.

      (4) The interpretation provided in the Discussion tends to oversimplify what is in fact a heterogeneous and rich set of structural findings into a relatively coherent mechanistic account. The observed differences are spatially and directionally variable across tracts, cortical regions, and metrics: e.g., FA is reduced in the UF and posterior interparietal corpus callosum but increased in the dorsal cingulum; cortical thickness is reduced in aPFC but increased in medial temporal regions, and so forth. The Discussion acknowledges this in part (e.g., proposing increased dorsal cingulum FA as potentially compensatory) but does not address the directional heterogeneity systematically. The authors could discuss more explicitly what the opposing directions of effects mean for their overall interpretation. Relatedly, some parts of the Discussion link specific structural findings to specific imagery processes in ways that go beyond what the current data can support. The authors could more clearly distinguish between what the structural data show and what functional interpretations are taken from prior work.

      We will add two recent in-press Cortex papers to the Discussion. One provides lesion-based double-dissociation evidence against V1 as a necessary causal substrate of visual imagery. The other shows that aphantasic individuals can display visualizer-like oculomotor patterns during mental map exploration despite reporting little or no imagery vividness. Together, these studies help clarify our interpretation of our null V1 findings and structural effects in higher-order brain regions, which are consistent with aphantasia involving altered integration or access rather than a primary V1-dependent imagery deficit.

      Reviewer #2 (Public review):

      Summary:

      This paper addresses whether congenital aphantasia reflects an alteration of visual representations themselves, or rather of the systems that allow internally generated representations to reach conscious experience.

      Strengths:

      The study is novel and ambitious. The authors combine several complementary structural MRI approaches in a rare and well-characterised population, and the convergence of the findings toward frontotemporal and cingulate systems, with relative sparing of early visual cortex and major visual pathways, is particularly interesting because it could affect the way visual imagery is modelled and tested experimentally and clinically.

      Weaknesses:

      Overall, I found the manuscript conceptually and methodologically strong. My main concern regards the interpretation of the anatomical findings, rather than the findings per se. The authors discuss their results within a rich cognitive framework. However, the current dataset does not appear to include independent behavioural or neuropsychological measures that would allow the proposed cognitive interpretation to be tested in the same participants. As a result, the manuscript sometimes moves quite rapidly from 'these structural differences involve systems associated with higher-order control, salience, conscious access' to 'these structural differences may explain the cognitive mechanisms of aphantasia'. I agree that this is the most interesting interpretation, and probably the right one to explore. Although plausible, it remains indirect. The authors already acknowledge this point when discussing memory, affective control, and semantic processing. However, the same logic should be extended to the interpretation of the full set of findings. For example, if the salience/anterior insula findings are interpreted in relation to access to internally generated representations, it would be useful to know whether aphantasic participants also differ behaviourally on tasks tapping interoception or related aspects of internal monitoring. I appreciate that collecting additional behavioural data may not be feasible at this stage, especially given the difficulty of recruiting participants with such a specific manifestation. However, I think it should be acknowledged more explicitly in a dedicated limitation paragraph.

      We thank the reviewer for this thoughtful and constructive comment. Lack of introspective report of voluntary imagery is arguably the defining signature of aphantasia. This motivated us to primarily interpret our anatomical findings in a broader cognitive context of higher-order control, internal monitoring, and conscious access in aphantasia. We expect that a reliable behavioural test measuring imagery sensitivity and accessibility would allow us to direct link these findings to individual imagery ability. Nevertheless, to our best knowledge, this kind of test on imagery is still missing. Instead, our findings point to some plausible structural signature or brain regions that may be related to conscious imagery, which motivate future studies to examine their direct or causal roles. We agree with the reviewer, future studies should test the relationship between these anatomical structures and the accessibility to internal representation, together with related aspects of internal monitoring. We will therefore add a dedicated paragraph to discuss the plausible cognitive mechanisms during the revision.

      Reviewer #3 (Public review):

      Summary:

      The authors investigate the structural brain basis of congenital aphantasia, a condition characterised by a lifelong absence of voluntary mental imagery. They test two competing accounts: one predicting structural differences in early visual pathways, the other predicting differences in higher-order frontotemporal and cingulate systems. To do this, they combine four complementary structural imaging approaches: white-matter microstructure profiling along anatomically defined tracts, tractography seeded from functional regions of interest, whole-brain structural network analysis, and cortical thickness mapping. The main finding is that white-matter differences are selective for frontotemporal and cingulate pathways and absent in early visual pathways, which the authors interpret as support for the higher-order account.

      Strengths:

      The multi-modal design is a genuine strength: running four independent analyses increases the chance of detecting real effects and of identifying false positives that appear in only one stream. The statistical choices within each analysis are appropriate. Permutation-based correction with a threshold-free method is well-suited to the tract-level comparisons. The use of Bayes factors to quantify evidence for null results, rather than simply reporting non-significant tests, is particularly valuable here, since the absence of visual pathway differences is central to the argument. The robustness checks across multiple brain parcellations for the network analysis strengthen confidence in those findings.

      Weaknesses:

      The main limitation concerns the relationship between two of the analysis streams. The measure used to weight structural connections in the network analysis is calibrated to match fiber density estimates derived from the same diffusion signal that drives the white-matter microstructure differences. If the two groups differ in tissue organisation in certain pathways (which the microstructure analysis suggests they do), that difference will feed into both measures. The authors should acknowledge this dependency when discussing convergence across analyses.

      More broadly, the imaging metrics used throughout (measures of fiber organisation and weighted connection counts) reflect what the diffusion model captures from the tissue and cannot be directly read as measures of axon number or connection strength. This is a known limitation of the field, but it is relevant to the strength of structural claims made in this paper.

      The network analysis is presented without comparison to a null network. Without this, it is hard to know whether the node-level differences reflect specific network topology or simply follow from overall differences in connectivity weight or density between groups.

      The study runs four separate discovery analyses on the same 36 participants, each corrected within itself but with no control across analysis streams. At 18 participants per group, this is exploratory work. Some of the language used in the abstract and discussion, like "first comprehensive characterization" and "selective structural phenotype", reads as more definitive than the data support at this sample size. Framing the results as hypotheses to be replicated would make the paper stronger.

      The paper frames the results as distinguishing between two competing accounts. The positive evidence for the higher-order account is clear. The absence of differences in visual pathways is a different kind of result: it means such differences were not detected in this sample, not that visual pathways are uninvolved. The discussion at times moves toward that stronger conclusion, which the data do not support.

      The cortical thickness analysis finds one cluster in the predicted direction, while the other analyses each return multiple effects. One cluster in a whole-brain search with 18 participants per group is not strong evidence and should not be presented as equivalent to the other results.

      Effect sizes are reported without confidence intervals throughout. With 18 participants per group, the uncertainty around those estimates is large, and confidence intervals would give readers a more accurate sense of what can be concluded.

      We are grateful to the Reviewer for the constructive and thoughtful assessment of our manuscript. In response to the reviewer’s comments, we will revise the manuscript to clarify the dependency between diffusion-derived analysis streams, to state more explicitly the biological limits of diffusion MRI metrics, to add a null-network sensitivity analysis for the clustering coefficient findings, to include confidence intervals for reported effect sizes, and to temper the interpretation of the cortical thickness result. We will also revise the Abstract and Discussion to better reflect the exploratory nature of the study and to frame the findings as hypotheses requiring replication in larger independent samples. We believe that these revisions will make the manuscript more balanced, transparent, and appropriately cautious, while preserving the central conclusion that congenital aphantasia is associated with structural differences centered on higher-order frontotemporal and cingulate systems.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      In their manuscript, Arjun et al. investigate the role of the histone acetyltransferase Gcn5 in the control of drosophila blood cell homeostasis in the larval lymph gland. They use gcn5 zygotic mutants as well as targeted knock-down and over-expression of Gcn5 in various lymph gland populations to show that these modulations impact (in a rather haphazard manner) niche cell number, blood cell progenitor maintenance, plasmatocyte differentiation, crystal cell differentiation or DNA damage accumulation. Their results suggest that Gcn5 controls autophagy and they show that decreasing the expression of the autophagy machinery increases blood cell differentiation. Using drugs to modulate the mTOR pathway, they conclude that Gcn5 levels are regulated by mTOR but that the impact of this pathway on blood cell homeostasis can override Gcn5 function.

      While the authors did a lot of experiments and good quantifications of the blood cell phenotypes, many results do not make much sense or do not bring valuable information about Gcn5 mode of action. Several conclusions of the manuscripts are not backed by solid data (e.g. that Gcn5 action is mediated by TFEB and the autophagy machinery) and different aspects of the literature are not well taken into consideration. Some results (such as the validation of the knockdown and overexpression of Gcn5) seem flawed. There are some concerns about the results obtained with gcn5 zygotic mutants and an interpretation of the phenotypes observed upon manipulation of Gcn5 expression in different cell types is missing.

      We have now performed several experiments to address the comments raised by the reviewer and have also provided possible explanation of the phenotypes in cases where it was lacking.

      Important revisions are needed to improve the quality of the manuscript and confirm the authors' findings.

      Reviewer #2 (Public Review):

      Summary:

      Drosophila hematopoiesis has been shown to be governed by a number of signaling pathways such as JAK/STAT and Dpp. This important study shows the role of nutrient sensing and autophagy in determining blood cell differentiation. The authors show that General control non-derepressible 5 (Gcn5), a histone acetyltransferase affects blood cell differentiation. Gcn5 also negatively regulates autophagy through its effector TFEB which directly regulates autophagy genes. The authors also show that mTORC1 modulates Gcn5 levels and through it, TFEB activity thus acting as a fine-tuning mechanism that maintains optimal levels of autophagy.

      Strengths:

      The main strength of the work lies in the interesting finding that cellular metabolic processes such as autophagy have a direct role in blood cell differentiation and has the potential to be of interest to those working on vertebrate haematopoiesis as well. The report has generated intriguing data, using promoters specific for sub-sections of the lymph gland, that different cellular subsets of the lymph gland contribute differently towards haematopoiesis, but this is not followed up in detail and the final conclusions are derived from a combination of whole lymph gland perturbations as well as those from specific promoters.

      Weaknesses:

      (1) Gc5 seems to be expressed throughout the lymph gland but modulating it in the subsections does not have the same result. It is very striking that the knockdown of Gcn5 in the prohemocyte population does not have an effect on differentiation whereas overexpression does. The modulations of Gcn5 in PSC also have variable effects across hemocyte subpopulations which is not explored in the manuscript.

      We have now explained and discuss why Gcn5 modulation could be affecting the PSC size. Please check Discussion section Paragraph 1 line 10 onwards.

      Interestingly, also the domain deletion constructs show a differential effect on blood cell differentiation when altered solely in the prohemocytes which is not explained.

      Currently, with our observations all that we can comment about that data is that expression of domain deletion mutants causes aberrant hematopoiesis indicating a dominant negative phenotype since they are expressed in the wild type genetic background. Beyond this, we will be exploring mechanistically how these domains are functioning during hematopoiesis in future studies. We have already described the dominant negative effect in the text: Discussion Section Paragraph 3.

      While Gcn5 can be seen in all sections of the lymph gland in the first figure, under the HHLT-Gal4 and Hml-Gal4, Gcn5 looks cytoplasmic and almost completely excluded from the nucleus strikingly unlike Gcn5 expression under the Collier-Gal4 and Dome-Gal4.

      We have now revised Figure 1 and have only included the images with Collier-Gal4 and Dome-Gal4 which clearly shows both the niche cells, Dome-positive progenitors and Dome-negative cells of the primary LG lobe essentially showing that Gcn5 is expressed throughout the primary LG lobe. In Fig. 1C-F’, Gcn5 expression is both in the nucleus and cytoplasm as this molecule shuttles between cytoplasm and nucleus. The staining pattern with the other Gal4 could be due to problems in the immunofluorescence protocol and acquisition parameters. We have now removed those images from Figure 1. Please check revised Figure 1.

      The rest of the experiments in the manuscript are done with multiple promoters, with autophagy flux measured by modulating Gcn5 with a pan hemocyte promoter, but the mTORC1-Gcn5 axis is explored using chemical modulators which affect the whole of the lymph gland (Fig7) or using two pro-hemocyte promoters (Fig8).

      We have used a pan-hemocyte promoter for the autophagy analysis to investigate if Gcn5 regulation over autophagy is a hemocyte specific effect which we indeed see. We have removed the western blot data now in the revised manuscript where we looked at Atg8 and p62 levels in whole larval lysates when Gcn5 was perturbed using hemocyte driver as the results were puzzling and difficult to comprehend given the complete absence of a p62 band in Gcn5 knockdown conditions. Also, it’s worth noting that Hml-Gal4 is also active in the LG hemocytes. We did 2 alternate promoters for prohemocytes to cross-validate some of our results and the chemical modulators experiment was done since effects like mTOR inhibition/nutrient sensing effects are systemic and hence such modalities were employed.

      (2) The knockdown of Gcn5 seems to affect the gland size (A compared to B and C). Since mTORC1 is a central regulator of cell size, it is possible that some of the effects seen in these knockdowns are potentially through mTORC1 affecting size suggesting that the signalling axis between mTORC1 and Gcn5 might not be a one-way axis as suggested in Figure 9. Also, this would mean that in experiments where absolute cell counts of crystal cells or niche cells are used to assess blood cell differentiation, further analysis to consider total cell numbers in the lymph gland would strengthen the manuscript.

      It is a possibility that Gcn5 perturbation could be affecting lymph gland size although we have not seen any consistent trend that would point towards this phenotype either upon knockdown or over-expression. We believe Gcn5 controls blood cell differentiation phenotypes strongly via mTORC1. But in order to answer reviewer’s comment we have now re-analyzed our crystal cell differentiation data particularly and quantitated it and represented it as crystal cell differentiation index for dome-Gal4 specific Gcn5 modulation and for the data with genetic modulation of mTORC1 pathway. Please see Fig 3P and S10J for the revised analysis.

      (3) A genetic manipulation of mTORC1 specifically in the pro hemocytes would strengthen the role of mTORC1 in the pathway rather than the chemical modulation which affects the whole of the lymph gland.

      We thank the reviewer for their useful critique. We have now addressed this concern and we have genetically perturbed the mTORC1 pathway in the progenitors using both abrogation of TORC1 via depletion of Tor or Raptor or by activation using over-expression of Rheb. We have now included this data as Supplementary figures – Fig S10 and S11 and have described it in the results section. Please see results section “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      The abstract could clearly be improved. It does not make a clear presentation of what is new in the manuscript. The conclusions that Gcn5 function in the lymph gland is mediated by the autophagy machinery and the acetylation of its non-histone target TFEB are not grounded and purely circumstantial. The implication of mTOR and nutrition in drosophila larval blood cell homeostasis has already been studied but not mentioned here. Most of the time the authors do not provide any possible explanation about the phenotypes they observe and how they fit with the current literature. Several pieces of results are of serious concern.

      We would like to thank the reviewer for their feedback. We have revised the abstract and have incorporated the insights obtained from our study. We have now included relevant literature that talks about the implication of mTOR and nutrition in Drosophila larval blood cell homeostasis (Please see Introduction section Paragraph 2 in the manuscript). We have also noted the input of the reviewer on many phenotypes lacking any description of a possible explanation. We have worked on results section to provide possible explanation and speculation wherever relevant.

      In the introduction, the authors do not provide an up-to-date and accurate presentation of the field. For example, they could use much more recent and comprehensive reviews since Evans et al. 2003. (eg. MID: 30733377 or 35887113). Their choice for signaling pathways involved in Drosophila blood cell progenitors seems very much biased for lead author self-citation rather than more directly related citations. It is surprising too that the authors failed to mention a series of publications on Akt/mTOR and nutrient sensing impact on drosophila larval blood cells (PMID: 22951642 ; 22911822 ; 22407365 ; 22510984). Along the same line, there are already several reports on autophagy genes implicated in Drosophila hematopoiesis and blood cell functions (PMID: 23406899; : 33560224 ; 20498061 ; 37623416). The introduction on GCN5 is a bit of a catalogue and should be streamlined- citing a recent review would be useful (PMID: 32735945). Again, the authors fail to cite publications showing that Gcn5 levels can be modulated by nutrition (PMID 27022023; 27874008) and they do not mention that amino acids starvation or mTOR inhibition leads to a decrease in GCN5 activity / TFEB acetylation (ref 40). Taking into account all the missing information, the novelty of the present manuscript is strongly decreased.

      We would like to thank the reviewer for the detailed suggestions on including the relevant literature that are appropriate and relevant to be mentioned in the context of the observations in our manuscript. We have now included these references and have cited them as per reviewer’s suggestions. Please check Introduction section paragraph 2.

      Results

      While there is little doubt that Gcn5 is expressed in the entire primary lobes based on Fig 1C-F, the quality of the staining in G-J (especially H, J) is really poor and essentially looks like non-specific background with no clear signal in the nuclei. Better images should be presented. The conclusion of the paragraph ("all cellular populations of the LG") and title of Fig.1 are not fully accurate as the authors do not provide evidence that Gcn5 is also expressed in posterior lobes.

      As per reviewer’s suggestions, since the images in Fig 1A-D’ clearly show that Gcn5 is expressed in the entire primary LG lobe in PSC cells, MZ and CZ; we have removed panels E-H’ which lacked clear nuclear signal. Fig1A-D’ clearly show the nuclear staining pattern of Gcn5. We have also modified the conclusion of the paragraph to say that Gcn5 is expressed in cellular populations of the primary lymph gland lobe accordingly.

      Concerning, Fig S1 and Fig 2, while the analysis seems technically sound, the results are puzzling. The lack of P1 differentiation in gcn5 null heterozygotes is very surprising. The authors should check that this stock does not carry a mutation in nimC1 (for details see: PMID: 23899817) and use other plasmatocyte differentiation markers to confirm their observation (also with the different allelic combinations). I'm also concerned by the levels of plasmatocyte differentiation and crystal cell number in the control line (notably in S1H), which seem very low (and quite variable for P1 as there is a notable difference between S1H and Fig 2H). Moreover, the analysis of the allelic combinations gives rather incoherent results: PCSC cell numbers are affected only in null/hypomorph, whereas differentiation (NimC1 and Hnt), as well as DNA damage, was only increased in hypomorph homozygotes. The authors propose no hypothesis to explain these observations.

      We have now repeated these experiments with the E333st null allele by placing it on a different balancer and we observe homozygotes that are alive till late third instar/early pupal stage as shown before by Carre et al., 2005. We have now included these revised results on the plasmatocyte differentiation status of the E333st heterozygotes and homozygotes (See Fig 2 and Fig S1). We do find P1 positive cells in the E333St heterozygotes unlike earlier. Plasmatocyte and crystal cell numbers in the control line always shows some level of heterogeneity. We have included the wild type control individually with those respective mutants during the experiment hence drawing a cross comparison across two different experiments would not be appropriate. We have now explained the observations obtained on PSC cell numbers (Discussion section paragraph 1). Experiments to check all hematopoietic aspects of the gcn5 null have been done after changing the balancer line and the null mutants overall show a decrease in PSC size and a widespread increase in hemocyte differentiation which could be due to a systemic effect due to various signalling pathways being affected which needs to be investigated and is beyond the scope of this study. This has also been discussed in the Discussion section Paragraph 1.

      Although a side-by-side comparison would have been better suited, it seems that the homozygotes or trans-heterozygotes do not have stronger phenotypes than the heterozygotes as far as crystal cell and DNA damage are concerned, which is rather unexpected. Besides the authors should introduce why they look at DNA damage.

      We agree with the reviewer that for the crystal cell and DNA damage phenotype the homozygotes or trans-heterozygotes do not have a stronger phenotype as compared to the heterozygotes alone but since these are whole animal mutants there could activation/inactivation of various signalling pathways and systemic effects that would be difficult to account for and comprehend here which needs to be investigated further. The only conclusion that we draw from these observations is that Gcn5 is required for maintaining blood cell homeostasis. Regarding DNA damage, we have now included the rationale and supporting literature for why we have studied DNA damage in the context of Gcn5. Please see result section 2 paragraph 1.

      Importantly too, the authors failed to obtain gcn5 E333st/E333st (null) larvae, whereas Carre et al. originally reported that E333st/E333st individuals are viable until the late third instar larvae. I suspect that the stock they use carries additional mutations that need to be eliminated by back-crossing it to control flies for several generations. Of note too, a recent report showed that a deletion of gcn5 (generated by CRISPR) does not prevent adult emergence, challenging the conclusion that gcn5 expression is absolutely required for fly development (PMID: 37545086).

      The reviewer is right in pointing out that E333st homozygotes survive until late third instar as reported by Carre et al.,2005. We have procured the null allele again and used another balancer to obtain homozygotes and we were able to get homozygotes that survived till late third instar as reported earlier. We have now included new data from these homozygotes for all hematopoietic aspects and heterozygotes particularly for plasmatocyte differentiation Please see Fig 2 and Fig S1 and corresponding results section 2 of the manuscript.

      Concerning the validation of Gcn5 knock-down and overexpression: the results are highly dubious. In Fig S2B (hml>Gcn5 RNAi), there is virtually no Gcn5 signal in the primary lobes but hml is normally expressed only in the cortical zone. How is it possible? Similarly, the western blot (which is really too much cropped around the bands of interest) does not show any signal in the hml>Gcn5 RNAi lane (not even some background. According to the Methods section, the western was performed on whole larvae extracts; hml-mediated knock-down can not wipe out its expression in all the tissues. As for the overexpression, flag immunostaining in hml>Gcn5-flag is mostly cytoplasmic (S2E), which doesn't make sense and does not fit with S2C (Gcn5 immunostaining).

      Hml-Gal4 is a pan hemocyte driver and its expression is not limited to the CZ of the primary lymph gland lobe (Banerjee et al., 2019) and recent single cell sequencing data corroborate this that Hml domain is not limited to the cortical zone (Yarikipati and Bergmann, 2026). GFP driven by Hml-Gal4 is spread out across the primary LG lobe which could explain the phenotype of no Gcn5 signal obtained in the immunofluorescence experiment. Regarding the western blotting experiment which was performed on whole larval extracts, we were also puzzled by lack of Gcn5 bands in these lysates upon depleting Gcn5 using Hml-Gal4. We need to systematically probe further to understand expression of Gcn5 in other tissues and organs. We have now removed the western blot data as the data obtained cannot be comprehended at the moment. Regarding the FLAG staining experiment – the staining gave us a cytoplasmic pattern and since Gcn5 is known to shuttle between the cytoplasm and nucleus it is possible that the anti-FLAG staining detected the Gcn5 localizing in the cytoplasm. It is difficult to draw a direct comparison here between the images S2C and S2E as both are different antibodies.

      The initial analysis of Gcn5 level modulation in the prohemocytes, PSC or Hml+ cells is mainly descriptive and the authors do not elaborate on possible explanations based on the current literature.

      We have added a possible explanation wherever required for these respective results on Gcn5 modulation in prohemocytes, PSC and Hml positive hemocytes. Please see result section 3 where we elaborate on possible explanation for the phenotypes observed.

      The structure/function analysis of Gcn5 is based on overexpression of truncated mutants in the prohemocytes using the tep4-GAL4 driver and monitoring PSC cell, prohemocyte maintenance, plasmatocyte and crystal cell differentiation as well as DNA damage. As the overexpression of the full-length protein was made with a different driver (Dome), it is difficult to interpret the data. Nevertheless, no clear message emerges from this analysis and the authors do not reach any conclusion. Thus, the interest of these experiments remains limited.

      The structure-function analysis was largely done to understand which of the domains of Gcn5 upon over-expression results in a dominant negative like phenotype and our analysis shows that expression of some of these domain mutants results in a dominant negative phenotype in the wild type genetic background which we have now stressed upon in the text. However, further mechanistic understanding and in-depth analysis of each of these domains of Gcn5 warrants further separate investigation and is beyond the scope of this study. Please see the end of result section 4 for conclusion and possible explanation.

      The authors then analyze autophagy markers (in hml>Gcn5 LOF or GOF). Contrary to their say, hml-GAL4 is not a pan-hemocyte marker. It would have been interesting to ensure that the effects observed on Atg8 and Ref(2)P in the lymph gland are cell-autonomous- as expected for a direct role of Gcn5 on this pathway. Again, it is very surprising that p62 is not detected in the western blot on whole larval extracts when Gcn5 is knocked down in Hml+ cells only (Fig 5D). Moreover, quantifications on multiple samples will be needed to validate the increase/decrease of p62 and Atg8 as detected by western blot. As for the RT-qPCR (Fig S5), according to the Methods sections, they were made on adult blood cells but this is not explicit in the result section.

      We have corrected the text and mentioned Hml-Gal4 as a hemocyte specific Gal4 shown earlier as Gal4 marking both embryonic and larval hemocyte population (Goto et al., 2003, Yarikipati and Bergmann, 2026). Regarding the Atg8 and Ref (2)P blots – yes, it is surprising to us too that the p62 is not detected in the larval lysates when Gcn5 is depleted using Hml-Gal4. However, this result was consistent over the replicates performed and needs to be further studied. Since this phenotype of complete absence of p62 in larval lysates upon Gcn5 depletion cannot be comprehended and explained, we have removed the western blot data from the figure and have just retained the immunofluorescence data and have also quantified the Atg8 and p62 puncta per cell and included this data in Figure 5, Graphs D and E. For the qRT-PCR we have now included a description in the corresponding results section. Please see result section – result 5 under “Autophagic flux in the Drosophila blood cells is negatively regulated by Gcn5”.

      The knock-down of TFEB or several autophagy genes in the prohemocytes (tep4-GAL4) leads to a rather convincing increase in plasmatocyte and crystal cell differentiation. It would have been interesting though to quantify prohemocyte maintenance, PSC cell number, and DNA damage. Also, the authors should have performed Gcn5 GOF/LOF experiments with the same driver (they present tep>Gcn5 RNAi in Fig 8 but without the proper controls).

      We have now included data for prohemocyte index (Figure S8M) upon knockdown of TFEB and other autophagy genes along with PSC cell number, DNA damage (Supple Fig S8) in the revised manuscript. Please see corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe” for the description of the results.

      The use of chloroquine should be better described. How long was the treatment? Did the authors observe an effect on autophagy in the lymph gland? Chrorloquine also affects lysosomal pH, so it remains to be demonstrated that the effects observed here are only autophagy-related.

      We have now written a detailed protocol for the treatment in the methods section and also mentioned the treatment time which is 16 hours in the results. We have included data to validate the effect of Chloroquine on autophagy by p62 and Atg8 staining in the LG and have quantitated the data (Refer Supple Fig S9) and the corresponding results section titled “Genetic and chemical ablation of autophagy boosts blood cell differentiation in the primary lymph gland lobe”

      Similarly, the use of drugs to activate (3BDO) or inhibit (Rapamycin) mTOR should be better controlled. More generally, given the promiscuous roles of mTOR (and autophagy) in the larvae, tissue-specific manipulations would be better suited.

      We have now perturbed mTOR pathway genetically by activation and in-activation and have studied the effect on blood cell differentiation. Please see Figure S10 and the corresponding result section titled “Chemical or genetic modulation of mTORC1 activity controls blood cell differentiation” where we discuss the results of genetic perturbation of mTOR pathway.

      Actually, as pointed out above, it has already been shown that modulation of Akt/TOR in hemocytes or amino-acid deprivation affects blood cell homeostasis (see above). The authors should definitely discuss how their results fit with the literature on this subject.

      We have added relevant literature in the introduction section and have also discussed how Gcn5 could fit into this context of nutritional sensing and control of hematopoiesis. Please check revised Introduction section paragraph 2. Also, check discussion section in last paragraph where we have discussed role of Gcn5 in nutrient sensing.

      Again, Gcn5 levels need to be quantified using multiple samples (Fig 7M, N) before concluding.

      Sorry for not including the quantitation earlier but we have now included the quantitation for the blots presented in Fig. 7 M and N.

      Finally, the authors show that 3BDO still induces an increase in blood cell differentiation when gcn5 is knocked-down in tep4+ cells and that Rapamycin still represses differentiation when Gcn5 is overexpressed in Dome+ cells. They conclude that mTORC1 overrides the effect of Gcn5. This seems a far-reaching conclusion given the available evidence.

      We have now toned down the conclusion that we make to accommodate other possibilities which we have been unable to test here currently.

      In particular, in the conditions used, the authors do not necessarily assess the activity/requirement for Gcn5 and mTORC1 in the same cell population.

      Other comments and suggestions:

      The discovery of the SAGA complex is not Grant 1999 but 1997 (PMID: 9224714).

      Ref 30 is not appropriate -nothing to do with HAT.

      GCN5 not only acetylates TFEB but also Atg7 (PMID: 28594263) to limit autophagy.

      Thank you so much for these suggestions. We have made the necessary amendments in the references.

      In the results section, the first paragraph is largely a repetition of the introduction. The same is true for most paragraphs in this section. A shorter (hypothesis-driven) introductory sentence would be more adequate.

      We have now taken the suggestion into consideration and made the necessary change in the results section throughout the manuscript.

      Fig 1: it seems that there is a higher accumulation of Gcn5 in a few cells in the cortical zone. This may correspond to crystal cells and could be easily confirmed.

      We have now checked this aspect. Please see supple fig S5 where we co-stain lozenge-GFP cells containing LG with Gcn5 to check for the accumulation. However, we do not see any accumulation in the Lozenge-positive crystal cells.

      Figure 3: the authors should also quantify the proportion of progenitors (dome>GFP+) in the different conditions.

      We have now done this and added it to the Figure. Please see panel N in Figure 3 and Figure S8M.

      Figure S3: how do the authors explain that Gcn5 knockdown in the PSC reduces plasmatocytes differentiation (but does not affect PSC cell number or crystal cell differentiation)? What could be the origin of the increase in DNA damage (essentially in CZ)? How do they explain that Gcn5 over-expression increases PSC size but does not affect (reduce?) blood cell differentiation?

      These observations need to be investigated further. We currently have no answer to these comments. The signals that are produced by the PSC could be affected due to which we observe these phenotypes like an effect on plasmatocyte differentiation and an increase in DNA damage whereas no effect on PSC cell numbers or crystal cell numbers which needs to be studied further. Also, in the case of Gcn5 over-expression in PSC we do not know how the increased size of PSC controls differentiation. This would need further experimentation and since this paper is not about the role of Gcn5 in PSC exclusively, we will look into this in our future studies. These aspects will be studied in our future follow-up studies as it is beyond the scope of the current manuscript.

      Figure S4: how do the authors explain the non-cell autonomous increase in PSC cell number upon Gcn5 KD/GOF in hml+ cells? How do they explain the increase in crystal cell number in Gcn5 GOF? Is it really cell-autonomous (i.e. all the Hnt+ cells are Hml+?)?

      We have discussed how Gcn5 depletion or over-expression in HmlΔ cells could affect PSC cell numbers. Please see discussion section, paragraph 1. Regarding the crystal cell phenotype - We have now tested if the increase in crystal cell numbers is cell autonomous by driving Gcn5 over-expression using a crystal cell specific driver and we find that the increase is cell-autonomous. Please refer to Supple Fig S5.

      The discussion is lengthy and should be reduced. It does not appropriately consider the current literature.

      We have tried to reduce the length of the discussion and have also added relevant references as per recommendations of the reviewer.

      Reviewer #2 (Recommendations For The Authors):

      (1) In general, it is not clear why in some of the experiments Tep-Gal4 is used to modulate proteins in prohemocytes while in others Dome-Gal4 is used.

      There is no particular reason. These Gal4’s have been used interchangeably as both label the hematopoietic progenitor population. Although recent single cell sequencing data has identified subsets within the progenitors namely core progenitors marked by tep4 largely and dome being a distal progenitor marker (Cho et al.,2020, Girard et al.,2021), in our study perturbations in Gcn5 using either of the Gal4’s results in a similar phenotype.

      (2) Considering alteration in lymph gland size (Figure 2), the number of positive cells should be analysed in relation to total cell numbers or s4ize.

      Although we do not find any visible differences or defects in the overall LG size in various genetic conditions discussed in this manuscript, we have done so for the plasmatocyte differentiation where we have represented it as plasmatocyte differentiation index (relative to the size of primary LG lobe) throughout the manuscript. We have now done this for crystal cell numbers too for critical genotypes in this manuscript and have represented it is as crystal cell index for example please see Figure 2O, 3P, S5G, S10J where these graphs have now been added.

      (3) Figure 1A G-I' does not look like mCD8 GFP expression, but rather cytoplasmic GFP.

      We have made the change in the figure and the corresponding text accordingly.

      (4) One of the main conclusions in the manuscript is that Gcn5 affects autophagy (Figure 5). Here, the puncta need to be quantified (relative to total cell numbers).

      Thank you for the suggestion. We have now quantitated the p62 and Atg8 positive puncta per cell and have represented it as panel D and E in Figure 5.

      (5) Figure 5 D and E show p62 and Atg8 total protein levels in the larvae when Gcn5 is modulated only in the hemocytes. It is surprising that there is a complete reduction in p62 levels across the whole larvae when Hml gal4 is used for the knockdown.

      Yes, we observe a complete absence of p62 in whole larval lysates when Gcn5 is depleted using Hml-Gal4 and we see this across replicates. This result is indeed puzzling to us and difficult to comprehend as to why a hemocyte specific driver would result in such a dramatic change hence we have decided to remove the western blot data as it is difficult to draw a solid conclusion from. We have retained the immunofluorescence data which shows a consistent alteration in autophagy upon Gcn5 perturbation using Hml-Gal4 and we have now included the quantification for the number of p62 and Atg8 positive puncta per cell for the IF data.

      (6) The beta-actin levels in the western blots in Figure 5 are highly oversaturated and do not represent loading control adequately. Also, it looks like there is substantially more total protein in 5D 3rd lane where Gcn5 is overexpressed.

      Thank you for pointing this out. We have loaded equal amount of protein in all the wells so we are unsure why the actin bands look over-saturated. We have now removed the western blot data from this figure as the data is puzzling and difficult to comprehend given a total absence of p62 in whole larval lysates in Gcn5 depletion conditions using Hml-Gal4. Hence, we are just retaining the immunofluorescence data.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      It is unclear to what extent the model's success relies on the way non-decision time is formalised in the model. In the proposed PDG model, non-decision time is decomposed into separate visual encoding, saccadic execution, and manual execution components. Several values (assumed or recovered) do not match known physiological or behavioural ranges. This is a common issue in the literature, and the authors may want to address it in light of broader work discussing what non-decision time consists of in both manual and saccadic actions (e.g., Bompas et al., 2024, Non decision time: the Higgs boson of decision, Psychological Review).

      In particular, the "saccadic execution" parameter appears far too long and too variable to reflect merely execution; instead, it likely includes decisional components. This would make more sense since manual and saccadic planning essentially rely on distinct brain areas, hence it seems unrealistic that crossing a single threshold would trigger both manual and saccadic execution. Similarly, recovered manual non-decision times are substantially longer (though not more variable) than expected motor execution durations for button presses. These patterns suggest that parts of what the model treats as non-decision time are likely decisional in nature, although perhaps related to "action decision" rather than the "value-based decision" of interest to the authors. To what extent these two processes neatly follow each other or overlap could be usefully considered.

      We have added a paragraph to the Discussion explaining how our model’s estimates of sensory and motor latencies relate to corresponding values inferred from physiology or behavioral manipulations (e.g., Bompas et al., 2024). Specifically, we write:

      “The key assumption of the PDG model is that there is a delay between the moment a choice is internally committed and the moment it is externally reported with a key press. Because eye movements are typically faster than manual responses (𝜏<sub>e</sub> < 𝜏<sub>m</sub> in our simulations), this delay creates a window during which gaze can already be directed toward the covertly chosen item before the response is formally registered. We do not interpret these non-decision latencies as irreducible physiological minima for moving the eyes or pressing a button (Bompas et al., 2025). Rather, they are inferred indirectly by fitting an additive non-decision-time parameter to the behavioral data, which we decompose into a sensory delay (𝜏<sub>s</sub>) and a manual execution delay (𝜏<sub>m</sub>). Values of 𝜏<sub>e</sub> are then chosen so that the model reproduces the observed magnitude of the behavioral effects. This estimation procedure has important limitations. Some participants show relatively “flat” chronometric functions: response times vary little with value despite otherwise normal psychometric performance. Such patterns likely reflect processes not explicitly represented in the model, including procrastination, reduced motivation, task-unrelated thought, or noise in item ratings. Within a drift-diffusion framework, however, these cases are accommodated by assigning a long non-decision time together with a short evidence-accumulation period (Table S1). Consequently, some estimated non-decision times are substantially longer than would be expected if they represented only sensory and motor delays. A further limitation is conceptual. We model non-decision time as occurring either before or after evidence accumulation, whereas in reality decisional and non-decisional components are likely temporally interleaved (Graziano et al., 2011). This simplification may also inflate the recovered latency estimates. With these caveats in mind, sensory and oculomotor delays on the order of 300 ms remain broadly plausible, although they likely lie near the upper end of a realistic range. The estimated eye-movement latency is especially long. For instance, in monkeys trained to report simple perceptual decisions with a saccade, roughly 100 ms elapses between the threshold-crossing signal in parietal cortex (or the superior colliculus) and the executed eye movement (Roitman and Shadlen, 2002; Stine et al., 2023). Crucially, however, varying the assumed non-decision latencies across a reasonable range does not alter the qualitative predictions of the model (Fig. 8).”

      Further, we have added a parameter sensitivity analysis. Importantly, although the magnitude of the predicted effects depend on the non-decision latencies, the qualitative aspect of these predictions do not (new Figure 8). Specifically, (i) the increasing tendency to look at the ultimately chosen item as time elapses (new Fig. 8A), (ii) the lack of an interaction between the last-fixation bias and overall value (Fig. 8B), and (iii) the absence of an effect of choice consistency on Δdwell (Fig. 8C) are all findings that are independent of 𝜏<sub>e</sub>.

      Reviewer #2 (Public review):

      The paper focuses on analyzing the Krajbich 2010 data, but shows that the second effect replicates in many other datasets. A more principled approach, in which both effects are analyzed and presented for all datasets, would be more convincing. The results should then be shown together for clarity/readability.

      Following this suggestion (and the reviewer’s elaboration in the private comments to the authors), we have substantially restructured the manuscript. Both aDDM predictions are now presented together (new Fig. 2), and Figs. 3–4 test these predictions across multiple food-choice datasets. In doing so, we no longer treat the data from Krajbich et al. (2010) separately, and we extend the analysis of the last-fixation–choice association (MELFB) to additional datasets. We note that the same datasets could not be used in both Figs. 3 and 4, as some lack information on the final fixation required for the MELFB analysis. Nevertheless, results are highly consistent across datasets and align with findings from a recent study by Ting & Gluth (2025), which independently identified and examined one of our key predictions; this work is now cited in the revised manuscript. Finally, to reduce redundancy, we have consolidated all aDDM variants and optimal models into a single figure (new Fig. 10).

      Similarly, it would be nice to show to what extent the models' predictions depend (not depend) on using the best-fitting parameter values (are there any parameter settings under which the two effects are not predicted?)

      The key predictions of the model depend on the difference between the manual (𝜏<sub>m</sub>) and eye-movement-related (𝜏<sub>e</sub>) latencies. We have now added a parameter-sensitivity analysis to show how the model predictions depend on this difference. The new analysis shows that while the quantitative predictions do depend on the precise latency values, the results are qualitatively similar across values of 𝜏<sub>e</sub> (new Figure 8).

      Reviewer #3 (Public review):

      There was limited discussion about why one might allocate attention post-decision. I would have appreciated more discussion on the potential functional consequences or implications of post-decision gaze.

      Thank you for this suggestion. We added a new paragraph to the discussion (paragraph #2), where we argue that it is sensible for a decision maker to direct the gaze to the chosen item once a covert choice commitment has been made, as the benefits of attending to a stimulus do not end with the decision itself. Specifically we now write:

      “Instead, these observations are better explained by a post-decision account of the gaze-choice association that is, one in which gaze shifts to the selected item after a covert commitment to a choice. We argue that directing gaze to the chosen item after a covert choice commitment is sensible, as the benefits of attending to a stimulus do not end with the decision itself. In naturalistic settings, for instance, selecting a food item is typically followed by the action of reaching toward it, where visual attention supports spatial localization and motor planning for the upcoming action. Although participants in our computerized task did not physically act on their choices, these sensorimotor processes are likely highly automatized and may still be engaged by default, even when not strictly required. Beyond motor preparation, post-decisional attention may also serve additional functions, such as facilitating sensory anticipation of the reward, supporting metacognitive evaluation of the decision, and contributing to value updating for future choices. From this perspective, a degree of attentional “stickiness” whereby the chosen item remains preferentially attended after commitment could emerge as an effectively optimal policy once these post-decisional processes are taken into account. Moreover, a specific feature of the task design may further reinforce this tendency: in the snacks paradigm, the unchosen item typically disappears from the screen immediately after a response is registered. It is therefore plausible that directing gaze to the chosen item after commitment partly reflects anticipation of the imminent disappearance of the unchosen option. To disentangle these mechanisms, it would be interesting for future work to test whether this attentional bias persists when the chosen item, rather than the unchosen one, is the stimulus that disappears upon response.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Major Comments:

      (1) Framing of the modelling approach

      The manuscript would benefit from acknowledging the known limitations of DDM-based frameworks, especially given that the entire study is conducted within these constraints. The introduction highlights successes of the DDM, but the manuscript does not mention any of its conceptual or empirical limitations.

      We are unsure about what specific limitations the reviewer has in mind, but we have added a paragraph to discussion mentioning some limitations, like the inflation of the non-decision times and the difficulty of interpreting the fit parameters (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (2) Dependence on non-decision time assumptions

      The alternative model's explanatory power appears to rely heavily on assumptions regarding the decomposition of non-decision time: fixed visual encoding (𝜏<sub>s</sub>= 0.3 s), manual non-decision time (𝜏<sub>m</sub>; two free parameters), and saccadic execution (𝜏<sub>e</sub>; fixed parameters μ<sub>e</sub> = 0.35, σ<sub>e</sub> = 0.11).

      - 𝜏<sub>e</sub> is substantially longer and more variable than typical saccadic execution times, suggesting it likely incorporates decisional components.

      - Estimated 𝜏<sub>m</sub> values are approximately twice as long as known manual execution durations.

      - σnd is more plausible, implying that variability is captured correctly but mean durations are not.

      Together, these points raise the possibility that portions of what the model treats as non-decision time are in fact part of a (action) decision process. Only then does it make sense to assume that Tm is usually larger than Te. If Tm and Te were truly execution delays, then Tm would always be larger than Te.

      You may find it helpful to consider the framework in Bompas et al. Psych Review (2024), which discusses in detail what non-decision time is likely to comprise across effectors.

      Thank you we have added (i) a sensitivity analysis showing that our results are robust to changes in the specific value used for the eye movement related latencies (new Fig. 8), and (ii) a new paragraph in Discussion addressing the issue of the mismatch between our parameter estimates and the manual and saccadic execution times (Paragraph #5 of Discussion: “The key assumption of the PDG model is that there is...”).

      (3) Code availability.

      The authors should consider sharing all relevant code and data publicly.

      We agree, we now share the code and data on GitHub and indicate so in the revised manuscript.

      Minor Comments:

      (1) Lines 74-77. These are not worded as predictions but as questions; one tests predictions, but answers questions. I feel it would be clearer to stick to predictions (like in the abstract), and the introduction could benefit from explaining these predictions in a bit more detail (I found it difficult to get my head around these predictions from the intro text only).

      We rewrote the section in the introduction where we provide a gist of the model predictions (last paragraph of Introduction). We agree with the reviewer that the previous explanation was not clear.

      (2) It is confusing that panel B appears to the left of panel A in Figure 2.

      We agree. We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed and the panels follow a more logical order.

      (3) Figure 3C - remove MATLAB toggles.

      Yes, thanks.

      (4) Figure 5A shows the proportion of left choices, but the text and legend refer to right choices.

      Good catch, thank you.

      Reviewer #2 (Recommendations for the authors):

      This may appear self-serving, but the authors seem to be unaware of some highly relevant work from our group. Most importantly, in a recent publication (Ting & Gluth, 2024, JEP General), we have already looked at the dependency of the last- (or final-) fixation bias on overall value in value-based (VB) and perceptual (P) decisions. In VB, we found a negative effect; in P we did not find a significant effect. This is largely consistent with the current results, showing a negative but not significant trend. Another relevant work is Gluth et al. (2020, Nat Hum Behav), where we extended the aDDM by assuming that the probability to fixate on an option is a function of the accumulated evidence for that option. It would be interesting to know whether this assumption changes the predictions of the aDDM. Finally, we just published a new theory on how people search for information to make efficient value-based decisions (Gluth et al., in press, Psychol Rev; https://osf.io/preprints/psyarxiv/3qzak_v2). Although this theory focuses on multi-attribute choices, it can be applied to "simple" choices, too (by assuming that there is only one attribute = value). Interestingly, while the model also mispredicts a (slight) increase of the last-fixation bias with overall value, it correctly predicts the independency of the dwell-time advantage effect on choice consistency as well as the small increase of the effect with RT (attached here is a figure to show this: [https://elife-rp.msubmit.net/elife-rp_files/2026/01/22/00149589/00/149589_0_attach_9_477122. pdf], and the match with the empirical data shown in Figure 3B and 12 is striking). In general, the model shares many features of the Callaway and Jang models, but does not need to assume a biased value prior, which the authors suggest is responsible for the misprediction of the second effect. I leave it up to the authors to discuss this new theory, but I wanted to point this out.

      Thank you for pointing this out; these are all relevant points and studies.

      We now note that the first of our predictions has recently been identified and tested by Ting and Gluth (2025).

      We also considered extending the manuscript with a variant of the model proposed by Gluth et al. (Psychological Review, 2026). In fact, we attempted to fit this model to the Krajbich et al. (2010) dataset under the assumption that the duration of each sampling epoch is a free parameter. We find this model very interesting. However, in our current implementation it appears to make the same qualitative prediction as the aDDM, namely that ΔDwell depends on choice consistency (see Author response image 1).

      Given this, we have decided not to include these results in the manuscript. It remains possible that with further development particularly with a more realistic specification of fixation durations (e.g., allowing them to depend on value) the model could account for the full set of observed effects. We think this would be best addressed in a separate study.

      That said, we do find the model promising, as it provides a better account than most of the alternative models we explored for the patterns shown in panels D, H, and I.

      Author response image 1.

      Fits of a variant of the MACS model (Gluth et al. 2026) to the data of Krajbich et al. (2010).

      The paper would benefit substantially from restructuring. The aDDM's predictions are provided first, together with the empirical data, and then the optimal models are discussed. But Figure 2 shows all of this together. Later, the new (PDG) model is elaborated, and its predictions are shown. Towards the end of the results, variations of the aDDM and combinations of aDDM and PDG are shown in a series of figures (8-11), followed by a last figure showing one of the tested effects in other datasets. All of this feels pretty much thrown together without a clear structure. For instance, the aDDM and the optimal models could be described together (or the optimal models get a separate figure). The additive variants could be described earlier. And some figures could be put into the supplement. And the empirical results of the different studies could be shown together.

      We fully agree with this suggestion. We have now restructured the manuscript along the lines proposed by the reviewer (see the more detailed explanation of the restructuring in our response to the public comments).

      I strongly suggest avoiding the term "influence" in the y-axis of Figure 2, upper row, as it implies causality. Similarly, in line 182, the term "causal influence" is used in the context of the Callaway model, but as far as I know, this is not what the model assumes.

      We replaced the y-axis label with “Association of last dwell with choice (β)”

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 2 - Panel labels for A and B are reversed?

      We have restructured the manuscript (following the suggestion of another reviewer), and now Figure 2 has changed.

      (2) Does 3C include a .pdf screenshot?

      Thank you, it’s a Matlab bug on Mac. I guess they want us to switch to Python -:)

      (3) Figure 4 - It would be helpful if the green line were defined in the figure legend.

      Added

      (4) The effect size in 5B looks much more dramatic than in 2B(A?) - Is this for one example subject as opposed to all subjects? Please clarify what is different about the data.

      We are no longer showing the psychometric functions in Figure 2.

      (5) Line 252 - they say they compared the probability of choosing the right item (Fig. 5B) by the y-labels of that figure, which are all p(choose left).

      Yes, corrected now.

      (6) In general, they reference the subpanels of Figure 5 out of order, which causes the reader to jump around. They might consider reordering the panels of the figure so they follow the ordering of descriptions in the text.

      We agree, we have rearranged the figure panels to follow the ordering of the descriptions in the text.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      (1) Alternative mechanisms for performance differences.

      The authors assume that the difference in performance between the low-switch (LS) and high-switch (HS) frequency conditions is explained by a change in the "leakiness" of integration. However, several other mechanisms could potentially explain this effect:

      (1) Temporal Uncertainty: Integration might start later in the HS condition, leading to lower performance.

      (2) Reduced Efficiency: Integration could be less efficient in the HS condition (i.e., lower signal-to-noise ratio) without a change in the leak parameter itself.

      (3) Evidence Contamination: Motion information from the adapting stimulus in the HS condition may be integrated rather than ignored, which might be the case since the transition from the adapting to the test stimulus is not externally cued.

      To distinguish between these alternatives, I suggest two possible analyses. First, a formal model comparison could be performed, though I acknowledge this may be inconclusive in the absence of response-time data. Second, an analysis of motion energy kernels could be revealing; the leak hypothesis makes the specific prediction that for long test stimuli, early samples should contribute more to the choice in the LS condition than in the HS condition, relative to late samples.

      We thank the reviewer for raising these important points. We agree that we cannot definitively identify the algorithmic underpinnings of the behavioral effects we report and have made substantial revisions to the manuscript to be clearer about what is supported and what is speculative in our claims. Most importantly, we agree that we do not know if the context-dependent differences in how accuracy depends on viewing time are based on adjustments to a leak or to something else (e.g., a saturating non-linearity, as we identified in Glaze et al, 2015, that is separate from the leak itself), which we cannot resolve with this dataset, even with more formal model comparisons. We therefore:

      Changed the wording throughout the manuscript to refer to changes in leakiness as just one of several possible sources of the behavioral differences. We also added this point to the list of “limitations” (and possible future directions, including using motion-energy kernels, which would require us to use lower-coherence test stimuli) in the Discussion (L487-493).

      Added a new figure panel (Fig. 2D), a new Extended Data figure (Extended Data Fig. 3), and additional explanatory text (L168-175) that collectively describe the behavior in more detail, including quantifying a “crossover” dynamic similar to what we reported previously (Glaze et al, 2015).

      Added new explanations (L152-163) and analyses (Extended Data Fig. 9) indicating that the monkeys used some information from the end of the adapting stimulus to inform their decisions, which accounts for the patterns of choices at the shortest viewing durations.

      Indicate that the context-dependent differences in the slopes of the psychometric functions (and complementary analyses based on “raw” accuracy measures as a function of binned viewing duration) rule out the temporal uncertainty and evidence contamination explanations, but are consistent with effects on the temporal dynamics of the decision process (L175-179).

      (2) Independence of neural and pupil-linked signals.

      The authors take the lack of session-wise correlation between context-dependent contributions from neural and pupil terms as evidence that these two signals provide independent contributions to the behavioral effect. However, could this lack of correlation simply be a result of high variability or noise in these estimates? The data shown in Figure 7B suggests that measurements are very noisy, which might obscure a potential relationship.

      We agree that the lack of session-wise correlation between neural and pupil terms cannot be taken as definitive evidence of independence. We have both softened the language around the claim (L368) and added a sentence to the Discussion (L464-468) acknowledging that this lack of correlation may reflect underlying noise and/or variability rather than true independence of the underlying mechanisms.

      Reviewer #1 (Recommendations for the authors):

      (3) The neural data analyses rely fundamentally on "switch" trials (Figures 3-5). It might be informative to also examine "non-switch" trials to see if there are specific neural markers indicating the exact moment the motion stimulus becomes behaviorally relevant. Given that this may fall outside the primary focus of the paper, it is up to the authors whether to pursue this line of inquiry.

      We thank the reviewer for this suggestion. We agree and have added new analyses of data from non-switch trials (Extended Data Fig. 9), which show some effects of stimulus information from the adapting epoch on the monkeys’ choices, as we detail below in response to related comments from the other reviewers.

      Reviewer #2 (Public review):

      Aspects of the behavioral analysis would benefit from a tighter connection between theoretical claims about evidence accumulation and the empirical features of the psychometric functions. For example, the rightward shifts observed across adapting conditions are interpreted as consistent with a reset of accumulation on switch trials, but similar patterns could also arise from failures to detect the test stimulus on a subset of trials, leading responses to default to the final adaptor direction. Likewise, changes in psychometric slope and asymptote are attributed to differences in evidence accumulation without explicit modelling or consideration of alternative explanations.

      Clarifying how specific features of the psychometric functions map onto distinct components of the decision process will strengthen the link between the theoretical framework and the behavioral data.

      We agree and have made substantial revisions to address these important points. Specifically, we added a new figure panel (Fig. 2D), new Extended Data Figures (3 and 9), and several lines of explanatory text (L152-179) that collectively describe the behavior in more detail, including clarifying that: 1) for the shortest viewing durations, the monkeys’ decisions were informed by information from the adapting stimulus, which accounts for generally lower accuracy on LSF (longer exposure to the final adapting direction, thus more accumulated evidence for that direction before processing the switch) vs. HSF (shorter exposure to the final adapting direction, thus less accumulated evidence for that direction before processing the switch) switch trials; and 2) as viewing duration increased, the rate of rise of accuracy versus viewing duration was higher for LSF vs. HSF trials, implying differences in the process of evidence accumulation. As detailed in our response to a similar comment from Reviewer 1, above, we are now careful to temper our claims about the specific computational basis (e.g., a leak or other form of nonlinearity) for these differences.

      We also de-emphasized our treatment of the asymptotes of the psychometric functions. In principle, these regimes could give insights into leakiness (which can limit the total amount of information that can be accumulated) and lapses (which are measured at the asymptotes). In practice, however, the long-duration trials that constitute the asymptotes were relatively under sampled (to promote the unpredictability of the offset of the stimulus, which we believed was the more important consideration when designing the experiment), yielding unreliable estimates.

      A slight concern is the lack of a consistent analytical approach for relating behavioral changes to neural and pupil-linked measures. Different sections of the manuscript rely on different behavioral metrics-such as differences in accuracy within a selected stimulus-duration range (e.g., Figure 5C) or psychometric slope differences (Figure 6C) without clear justification for these choices. The analytical approach likewise varies between simple correlational analyses (Figure 5C, Figure 6C), pseudo-experimental group comparisons (Figures 5D, E), and the inclusion of neural or pupil terms in the behavioral psychometric regression model (Figure 7B). While each metric and approach may be defensible in isolation, adopting a more consistent framework will help convince readers that the reported effects are robust and not contingent on the selective choice of metric or analysis.

      We thank the reviewer for this thoughtful critique and agree that the rationale for our choice of behavioral metrics and analytical approaches could be stated more clearly. We have added text to the relevant sections of the Results (L247-251) clarifying these choices. In particular:

      The neural analyses (Figures 3D-E, Figure 4, Figure 5D-E) focused on preferred-motion switch trials, because: 1) low switch-frequency non-switch trials provide an additional 800 ms of exposure to the final adapting-stimulus motion direction relative to high switch-frequency non-switch trials, which confounds comparisons of context-dependent evidence encoding between conditions, and 2) MT neurons exhibit minimal responses to null motion (although note that we also included analyses based on ROC area, which is computed from both preferred- and null-motion switch trials, to account for possible contributions of null-motion responses; Figure 5A-C). Thus, to ensure a meaningful comparison between neural and behavioral measures, we used behavioral accuracy on switch trials as the relevant metric in Figure 5C-E, rather than psychometric slope, which is estimated across both switch and non-switch trials.

      The pupil analyses (Figure 6) focused on a time window preceding test-stimulus onset, representing the arousal state around when the decision process started, and included both switch and non-switch trials. Thus, for these analyses we used psychometric slope, which is estimated across both switch and non-switch trials.

      We used several different analyses to compare and contrast the neural-behavioral and pupil-behavioral relationships because they provide complementary and useful insights. The correlational analyses in Figures 5C and 6C characterize session-level relationships between neural/pupil signals and behavior. The group comparisons in Figures 5D–E provide a complementary visualization of the same relationship. The model-based approach in Figure 7 then allows direct quantification of the trial-wise contributions of each signal to behavior within a common framework. Importantly, the conclusions drawn from each approach converge on the same interpretation, which we believe speaks to the robustness of the reported effects.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 2 legend. Description of 'running average (5-trial window)' is unclear - presumably this is a running average in stimulus space rather than across trials.

      We thank the reviewer for flagging this ambiguity. We have updated the legend (L136-137) to clarify that the running average is computed across trials sorted by test-stimulus duration.

      (2) L158. Difficult to establish an asymptotic performance level for HSF conditions within the stimulus duration range tested.

      We have removed the reference to asymptotic performance and replaced it with a discussion of performance on longer-duration switch trials in the context of the newly added Figure 2D.

      (3) L515 Equation 1. While this is a standard formulation of lapse rate in psychometric functions, the construction here in terms of switch probability is not standard. Given the task and training, it seems more likely that on lapse trials, the animal will respond according to the last adapted direction (rather than randomly switch/stay with equal probability).

      We thank the reviewer for this point. We agree that it is possible that on at least some of the “lapse” trials the monkeys may respond according to the final adapting-stimulus direction rather than choosing randomly. However, we cannot distinguish those alternatives using this task design. We include a statement to this effect in Methods (L569-571).

      To explore the idea further, we refit the behavioral data using separate upper and lower asymptotes corresponding to lapse rates on switch and non-switch trials, respectively. Across monkeys, there were no significant differences between upper and lower lapse rates for either low (Wilcoxon signed-rank test for equal medians: p = 0.15, Cohen's d = -0.13) or high switchfrequency (p = 0.07, Cohen's d = -0.16) conditions. So, at the very least, there was no evidence for lapse-like errors driven by switch- (or non-switch-) specific defaults to the final adapting direction.

      (4) L256. Statistical significance of attenuation is not directly tested here.

      We have replaced "were attenuated" with "we did not identify any reliable context-stability differences" (L297) to accurately reflect what was directly tested without implying a statistical comparison between groups of sessions that was not performed.

      (5) L429. Does the increase in explanatory power warrant the increased complexity of the model here?

      We thank the reviewer for raising this important point. We used Tjur's pseudo-R<sup>2</sup> because it does not increase by default with added model complexity, making it more conservative than other R<sup>2</sup> measures in this respect. Tjur's pseudo-R<sup>2</sup> is a coefficient of discrimination, and as such its value increases only when additional terms improve the model's ability to separate predicted probabilities across response outcomes. Thus, the observed increases in explanatory power when adding neural or pupil terms reflect real improvements in discriminability rather than an artifact of model complexity. We have added a brief clarification of this point to the Methods (L662-664).

      Reviewer #3 (Public review):

      The task design may not be optimal. While the amount of time the monkey is exposed to each motion direction during the adapting stimulus is matched, it's hard to know if the reduced MT responses to the test stimulus are truly due to the greater frequency of switches during the HSF adapting stimulus or because the monkeys have been exposed to more repetitions of the stimulus. It's increased sensory adaptation in either case, but it makes it problematic to interpret this as temporal context-dependent adaptation specifically. I think this could potentially be partially addressed by an analysis that is in the paper, but could potentially be emphasized/fleshed out more, specifically the results shown in Figure 4D that seem to show that most of the reduction in neural response for adapting units occurs between the first and second stimuli.

      The reviewer raises an important point. The number of stimulus repetitions and switch frequency are confounded in the experimental design, making it difficult to attribute context-dependent differences in MT responses to the temporal pattern of switches rather than to accumulated repetitions. We also note, as the reviewer acknowledges, the observed differences reflect sensory adaptation either way. Figure 4D does offer relevant evidence, suggesting that a majority of the change in neural response occurred with just one stimulus repetition. This finding complicates an interpretation where adaptation scales with the number of stimulus repetitions. We have added several lines to the Results about these points (L231-233).

      The pupillometric analysis seems to be an indirect way of assessing whether the accumulator itself might be modulated by temporal context, but the link could be made clearer. The authors show that context-dependent behavior is related to pupil size, which is related to arousal/neuromodulation, but it would be helpful to have some idea of what neural mechanisms underlying adaptive decision-making are actually impacted by this neuromodulation. Lacking neural data to address this question (e.g., from a brain region proposed to be involved in the accumulation process), at least more discussion of this would be helpful. Essentially, I'm unsure of how to interpret the pupil results: the argument that temporal context affects instantaneous evidence encoding in MT that then drives the accumulator is very clear, but I am a bit confused about what, mechanistically, I should think about the effect of neuromodulation doing.

      We thank the reviewer for this thoughtful comment and agree that the mechanistic interpretation of the pupil results could be made clearer. We acknowledge that we cannot directly identify the neural mechanisms underlying the arousal-related contributions to adaptive evidence accumulation from pupil data alone, given that pupil size is an indirect and imperfect proxy for neural (e.g., LC-NE system) activity. However, we can offer some informed conjecture and have added to the Discussion (L469-482) in an effort to elaborate on possible mechanisms.

      Reviewer #3 (Recommendations for the authors):

      (1) Abstract could be retooled - does not emphasize the pupillometry/arousal results very much, and they are presented more as a control than an independent result.

      We agree and have revised the Abstract accordingly.

      (2) Do all neural/pupil analyses use only switch trials? Sometimes the figure captions do specify only switch trials, but not everywhere. It would be helpful to specify either in the Methods or at the beginning of each figure caption that all subplots show switch trial results. Also, if you do always use switch trials, it would be useful to see in the Supplement how the non-switch trial results differ from switch trials. It seems like they may in interesting ways based on the behavioral results (supporting a reset of evidence accumulation on switch but not non-switch trials).

      We thank the reviewer for flagging these important points. We have added a justification for switch trials (L186-190) as well as clarification about which trial types were used for which analyses (L246-249) and information about trial types to relevant figure captions. We have also added a new Extended Data figure (Extended Data Fig. 9) examining relationships between neural activity and behavior on non-switch trials. As inferred by the reviewer, behavior on non-switch trials is consistent with the use of information from the adapting stimulus.

      (3) In Figure 3C, 5B, etc, when computing firing rate for the test stimulus (50-500 ms), are differently sized windows used to compute the rate for different test stimulus durations (since some will be <500 ms)? Or are only trials where the test stimulus duration is > 500 ms used for this analysis?

      We thank the reviewer for raising this point. To clarify, the 50–500 ms window does not reflect a fixed window applicable for all trials. Rather, neural activity from 50 ms after test-stimulus onset through test-stimulus offset was included for each trial, with 500 ms serving as the upper bound for trials with longer durations (> 500 ms). We have clarified this in the Methods (L607-610) to avoid ambiguity.

      (4) I think it might be better to be consistent with the time windows used for analysis; specifically, to choose either the 50-500 ms window used in Figures 3, 4, and 5B, or the 200- 400 ms window used for the remaining analyses in Figure 5.

      We agree that using the same window for all of the analyses would improve consistency, but not doing so provides advantages that we believe take precedent and now describe in more detail. The broader 50–500 ms window used for Figures 3, 4, and 5B was chosen to characterize MT neural activity over a relatively large a time window, ensuring that every trial contributes to each estimate. Because test-stimulus durations were drawn from a truncated exponential distribution (100–1200 ms), restricting these analyses to the 200–400 ms window would have excluded the substantial proportion of trials with durations <200 ms (but would yield similar figures and conclusions). The narrower window used in subsequent analyses allows us to focus on the conditions that exhibited the biggest modulations of neural activity when comparing them to behavior.

      (5) Similarly, provide justification for using only trials ending 375-600 ms after test stimulus onset for the behavioral correlations. It seems reasonable to choose a subset of test stimulus durations where the monkeys' behavior is greater than chance but less than ceiling, but it would be good to specify this so that it doesn't seem arbitrary.

      We agree and have added text to make this important point (L249-251).

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      Okabe and colleagues build on a super-resolution-based technique they have previously developed in cultured hippocampal neurons, improving the pipeline and using it to analyze spine nanostructure differences across 8 different mouse lines with mutations in autism or schizophrenia (Sz) risk genes/pathways. It is a worthy goal to try to use multiple models to examine potential convergent (or not) phenotypes, and the authors have made a good selection of models. They identify some key differences between the autism versus the Sz risk gene models, primarily that dendritic spines are smaller in Sz models and (mostly) larger in autism risk gene models. They then focus on three models (2 Sz - 22q11.2 deletion, Setd1a; 1 ASD - Nlgn3) for time-lapse imaging of spine dynamics, and together with computational modelling provide a mechanistic rationale for the smaller spines in Sz risk models. Bulk RNA sequencing of all 8 model cultures identifies several differentially expressed genes which they go on to test in cultures, finding that ecgr4 is upregulated in several Sz models and its misexpression recapitulates spine dynamics changes seen in the Sz mutants, while knockdown rescues spine dynamics changes in the Sz mutants. Overall, these have the potential to be very interesting findings and useful for the field. My major concerns from the initial manuscript, especially regarding cherry picking and circularity have been addressed with revised analytical approaches. I have some remaining minor comments.

      (1) The comparison between two wild-type samples versus wild-type-mutant samples is helpful - I think this could be added to the manuscript.

      As suggested, we added the figure comparing two wild-type samples against wild-type mutant samples as Supplementary Figure 2. 

      (2) For results of time-lapse imaging - please spell out in the results section the direction of change (lines 270 - 277).

      As suggested, we added the direction of change (an increase in the turnover rate) to the text (page 12, lines 270-271).

      (3) Using linear mixed effect models for statistical analysis is a significant improvement. While a sample size (n) of mice = 3 is not ideal, I think given the multiple different mouse lines used and intensity of analysis, this is probably the best that can be done, although further validation in larger samples eventually is to be hoped for.

      We appreciate the reviewer for recognizing the effort required to collect data across multiple mouse lines.

      (4) The revised text is much improved, but I still think the authors should be upfront somewhere in the text that the schizophrenia-associated genes can only confer biased risk for schizophrenia (and that the clinical phenotype can also include autism). As I said before, I think this is the best we can do and I agree with their choices, but it is important not to overstate the link. The differences they see make it clear that these are still relevant distinctions.

      As suggested by the reviewer, we further modified the discussion related to the comparison between ASD- and schizophrenia-associated mouse models (pages 23-24, lines 508-522).

      “The nanoscale features of dendritic spines in mouse models of Nlgn3<sup>R451C/(y or R451C)</sup>, Syngap1<sup>+/−</sup>, POGZ<sup>Q1038R/+</sup>, and 15q11-13<sup>dup/+</sup>, which we classified as being related to ASD, are highly heterogeneous. This heterogeneity may reflect the broad clinical spectrum of ASD, which ranges from mild impairments in social skills to severe intellectual disability. Accordingly, these four mouse models may represent distinct subgroups characterized by different degrees or forms of hippocampal dysfunction. Notably, among the ASD-related models, 15q11-13<sup>dup/+</sup> showed population-level spine properties closer to those found in the 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> mouse models. Although we classified 22q11.2<sup>del/+</sup> and Setd1a<sup>+/-</sup> as schizophrenia-related models, both 22q11.2 deletion syndrome and Setd1a haploinsufficiency in humans are also associated with ASD, suggesting substantial overlap in the genetic risk factors underlying ASD and schizophrenia. Further systematic analyses linking rare genetic variants to synaptic phenotypes in mouse models may provide important insights into the mechanisms underlying both shared and disorder-specific synaptic alterations in neurodevelopmental and psychiatric disorders.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) I would suggest that it might be preferable to use the word 'neuropsychiatric' rather than 'mental' in the title.

      As suggested, we modified the manuscript title.

      (2) I think it would be clearer to say that DEGs are listed if present 'in three or more models' rather than >2 (I appreciate the latter is mathematically clear, but can easily be read as 2 or more if reading fast). This is changed in the figure legend, but I suggest it is also changed in the main text (line 352-3)

      As suggested, we changed the main text to incorporate "in three or more models" (page 16, line 352).

      (3) Please add to Methods (line 557) that 'control cultures were prepared from littermate embryos....'

      As suggested, we added the phrase "control cultures were prepared from littermate embryos" (page 26, line 559).

      (4) Sorry to add something, but please could the authors add a definition of how they calculate spine turnover (and add units to the y axis of Figure 5A-C)?

      As suggested, we modified the y-axis of Figure 5A-C (% as unit) and added the method of calculating spine turnover rate in the text (page 36, lines 808-811).

    1. Author response:

      We would like to thank the editors for their interest in our work and the three referees for their time and careful reading of the manuscript. The reviewers have provided a series of helpful suggestions that we discuss in this provisional reply and will seek to address in the revised version of the manuscript.

      The main concern raised is that the bacterial community we refer to as a biofilm may instead correspond to a cell aggregate. Following the passing of Prof. Kevin Wood, in whose lab the experimental work was carried out, our ability to perform additional experiments is limited. Nevertheless, we plan to wash and fluorescently stain the extracellular matrix before imaging to measure the extent to which the observed bacterial community is an attached biofilm. In the meantime, we would like to highlight the work of Wen Yu et al. [1], in which E. faecalis biofilms were grown in 96-well plates under antibiotic stress. In particular, one of the strains of E. faecalis used in this article was OG1RF, the same strain used in our study. Crystal violet staining was used to quantify biofilm biomass, providing evidence for biofilm formation under those conditions. While we recognize that the experimental setup differs from ours and that the OG1RF sample used did not contain fluorescent and resistance plasmids, these results nevertheless support the expectation that OG1RF will readily form biofilms.

      Reviewer #1 (Public review):

      The mechanistic interpretation could, however, be clarified further by more explicitly emphasizing the competing timescales associated with detoxification, growth, and resource limitation. The current results suggest that when resistant cells are initially abundant, detoxification occurs rapidly relative to growth, allowing the population to approach carrying capacity after relatively few doublings, whereas slower detoxification at lower resistant fractions may permit greater expansion of sensitive cells once antibiotic concentrations decline. Additional direct measurements of antibiotic concentrations over time would also strengthen the connection between the experimental system and the modeling framework by testing whether the detoxification dynamics assumed in the model are quantitatively appropriate, although this seems very plausible.

      The timescale of drug degradation is an important system metric. We appreciate the referee’s suggestion to quantify antibiotic concentration over time. We plan to perform experiments in which samples are collected from the culture at fixed time intervals. After removing the bacteria from the samples via centrifugation, serial dilutions of the supernatant will then be spotted on a lawn of sensitive cells to measure the antibiotic efficacy at each time point.

      Reviewer #2 (Public review):

      One clarification the author should make is on the biofilm growth process. Specifically, could staining experiments be performed to demonstrate the secretion of the extracellular matrix? Just by looking at Figure 1b, it is hard to say. It remains a question whether the biofilm culture simply contains unstructured clusters rather than real biofilms (that are usually structured).

      We agree with the referee that additional evidence would strengthen our study. As noted above, we will perform additional experiments to demonstrate the presence of an attached biofilm.

      Reviewer #3 (Public review):

      The observed results are tied very closely to the experimental setup of adding antibiotics very close to the time of inoculation, but this connection is not discussed. [...] The mathematical model is used to confirm the result that no spatial components are needed to describe the results; however, this is mostly linked to the initial setup of the experiment, where antibiotics are added at the time of inoculation, and no biofilm could form before the outcome of the antibiotic-cell interactions was concluded.

      The experiment was designed to address how coupled planktonic and biofilm populations develop in the presence of antibiotics, which we will more explicitly discuss in the revised manuscript. We do agree that investigating how mature biofilms and their planktonic populations respond to antibiotic stress is an exciting direction for future studies. However, we believe that is beyond the scope of the our study on the development of coupled populations. We will be sure to explicitly identify this limitation in our revisions.

      The described ‘population inversion’ effect is better described as frequency-dependent selection for resistant cells, but frequency-dependent selection is not discussed.

      The reviewer is correct that this ‘population inversion’ is a frequency-dependent (perhaps also density-dependent) effect, and we should have situated it within that broader ecological framework. We will use this terminology in our revisions. We do want to acknowledge that the late Dr. Kevin Wood was fond of this phrasing to describe the reversal of the dominant strain, which is not necessarily true for frequency-dependent effects. Although we do not know for certain, we suspect this was a play on the ‘population inversion’ term used in quantum physics, used to describe a system in which its excited state (high energy) population unexpectedly outnumbers its ground state (low energy) population.

      The authors claim that biofilm and planktonic bacteria are protected equally by the presence of resistant bacteria; however, Figure 1a and b seem to clearly show that the proportion of sensitive cells is higher in the planktonic cells compared to biofilm cells when started from an equal frequency inoculum, meaning this is not always the case.

      If the reviewer is indeed discussing Figures 1a and 1b, these are not comparable as the starting fractions differ. On the other hand, if the reviewer was talking about Figures 2a and 2b (which is more clearly discussed by looking at Figures 2c and 2f), we agree that it appears that planktonic communities tend to have a slightly greater frequency of sensitive cells than the biofilms. We will be sure to highlight this observation and possible explanations in our revisions. However, given the uncertainty in these observations, we do not believe the differences are sufficient to alter our overall conclusion that final resistant fractions in biofilm and planktonic populations are quantitatively similar. Furthermore, the no drug treatment shows the same trend, which suggests it’s an effect of different growth dynamics of these two strains at high density rather than driven by the protective effects of resistance cells.

      Confocal microscopy was used to quantify the relative proportion of antibiotic-resistant and sensitive cells in the biofilm; however, it is unclear if the entirety of the Z stacks was used to determine these proportions. This is also the case for the analysis of whether the sensitive/resistant cells are non-randomly distributed in the biofilm: it is unclear whether the vertical distance between cells was taken into account.

      The entirety of the Z stack was used to measure the final resistant fraction in the biofilm. On the other hand, we used only the densest slice of the Z stack to calculate the correlations. The correlations follow the same trend when calculated over less dense slices, but as density decreases, noise increases, so such plots did not bring more clarity to our conclusions and were not included in the manuscript. Additionally, only horizontal correlations (over a slice) were calculated because consecutive Z-stack slices were imaged with a 2.5 µm spacing. Given that the average cell diameter is approximately 1 µm, calculating vertical correlations may miss neighboring cells located between imaged slices, making such measurements unreliable. We will clarify the points raised by the reviewer in the results section and add more detail to the imaging methods section in the revised manuscript.

      References

      (1) Wen Yu, Kelsey M. Hallinen, and Kevin B. Wood. “Interplay between Antibiotic Efficacy and Drug-Induced Lysis Underlies Enhanced Biofilm Formation at Subinhibitory Drug Concentrations”. In: Antimicrobial Agents and Chemotherapy 62.1 (Dec. 2017), 10.1128/aac.01603–17. doi: 10.1128/aac.01603-17. url: https://journals.asm.org/doi/10.1128/aac.0160317 (visited on 01/11/2026).

    1. Author response:

      We thank the reviewers for their careful reading of our manuscript and for providing positive, constructive feedback. In particular, we thank the reviewers highlighting the several strengths of our study.

      To address the reviewers’ major concerns, we will revise the presentation of our main findings (specifically data/animal vs data/ROI), provide more clarity in the Results, Methods, and Discussion sections, and modify the title to better reflect these nuances.

      Additionally, we will perform the following new experiments:

      (1) Astrocytic Marker Validation: To further confirm comparable astrocyte cell counts between the CTRL and Upf2-cKO conditions, we will perform immunostainings using Aldh1L1, Sox9, or S100b instead of GFAP.

      (2) NMD Candidate Validation: To validate top candidate NMD target transcripts, we will perform immunostainings or qRT-PCR for Gabbr2, Adora1, S100b, or Cldn9.

      (3) Sample Size Expansion: To strengthen the morphological and PSD-95 quantifications, we will increase the sample size (N) by incorporating additional animals.

      (4) Mechanistic Timeline & Phenotype Linkage: We value the reviewer’s comment regarding the timeline of morphological and Ca<sup>2+</sup> phenotypes. To gai insight into whether these phenotypes are independent or linked, we will perform 3D reconstructions in CalEx conditions to assess whether Ca<sup>2+</sup> restoration rescues astrocyte morphology in Upf2-cKO mice. This will allow us to determine if increased Ca<sup>2+</sup> activity is upstream of the morphological alterations. Taken together, we believe that incorporating these manuscript revisions will strengthen the clarity and conclusions of our work. We thank the reviewers for their time and careful evaluation of our study.

    1. Author response:

      We sincerely thank the editors and reviewers for their overall positive assessment and constructive feedback on our manuscript detailing the nanoscale organisation of βII-spectrin of the membrane-associated periodic skeleton (MPS) in mouse sciatic nerve axons. Their perspective and comments will help refining the manuscript.

      A common comment by the reviewers relates to the description of the characteristic longitudinal periodicity of the MPS. We value these comments, which we believe are motivated by the fact that the longitudinal periodicity of the MPS is undoubtedly the most studied and prominent feature of the MPS in cultured neurons. However, the main goal of the present project was to describe how βII-spectrin is organised in the transverse axis of individual segments of the MPS in nerve tissue. This is why we utilised cross-sections of the sciatic nerve, hence achieving the best resolution possible in that plane, at the expense of the resolution in the axial axis. Furthermore, this study clearly shows that the transverse morphology of axons, and thus of the MPS, of neurons in the tissue is highly irregular, in comparison to cultured neurons. This imposes an extra challenge to observe correlated longitudinal structures when the observation length is limited, as in our studies. Nonetheless, to improve this aspect of the manuscript, we will revise our data and previous evidence, clarify the methodological trade-offs made, and make our interpretations more accurate.

      Additionally, we will clarify several imaging- and definition-related inquiries, including tests for insufficient staining, the interpretation of βIII-tubulin staining, the assessment of axon–glia boundaries, and consistency in the use of terms like ‘clusters’ and ‘periodicity’, among others.

      We believe these and other revisions will substantially strengthen the manuscript and comprehensively address the reviewers' feedback.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In the wild, bacteria can be found in a wide range of metabolic states, including states in which they are resource-limited. Because phages heavily rely on the infected cell's molecular machinery to replicate, it is natural to wonder how phage-bacteria interactions depend on the metabolic state of the cell. In this work, Marantos et al. investigate specifically how the rate of infection of 5 different phages changes between cells grown in energy-rich conditions and cells grown in energy-depleted conditions. Their results clearly show that 4 out of the 5 phages studied display a significant reduction in infection rate in cells that are energetically depleted and provide a potential explanation for this observation by looking into the mechanisms that these phages use to irreversibly infect their host cells.

      The work also tries to explain the observation using a mathematical/mechanistic model that describes infection as the sequence of two steps, where a phage first needs to bind to a cell receptor, from which it can potentially unbind, and then irreversibly infects by injecting its genome. While the model is sensible from a mechanistic perspective, the experimental evidence that supports how each model's rate is affected by the cell metabolic state is weak, as only ratios of these rates can be inferred from the data.

      Reviewer #2 (Public review):

      Summary:

      The authors investigate the dependence of phage adsorption rates on host metabolic state, using 5 coliphages that differ in their infection cycles and host receptors. They find that four of the 5 phages showed significantly reduced infection under low metabolic states, with phages that generally have weaker adsorption being more strongly affected by low metabolism. The authors complement their findings with a 2-step infection model where phages can disengage from their hosts after initial adsorption. The paper illustrates the power of standardized experimental protocols for quantitative trait comparisons and highlights the dependence of phage infection success on host physiology.

      Strengths:

      The paper is well written and clearly structured.

      The experiments are well-designed, and particularly commendable is the diligent use of control scenarios to allow for quantitative comparison between phages. This standardized protocol will be valuable for the entire phage community.

      The authors convincingly show the impact of host physiology on phage adsorption success. This dependence has so far mainly been considered for intracellular phage replication, and the paper shows that host physiology has to be taken into account at all steps of phage infection.

      Weaknesses:

      There are some concerns about the experimental setup and which conclusions can be drawn from it:

      Before phage infection, bacterial cultures are grown to exponential growth, washed, and then resuspended with glucose or arsenate-azide for 10min. It is however, questionable that 10 minutes is enough to simulate high and low metabolic states realistically. 10 minutes seems to be quite short to go from exponential growth to a low metabolic state, given the transcriptional memory of previous environments. It seems more likely that the population will be quite heterogeneous, with cells in various states of transition towards low metabolic states.

      While we agree with the reviewer that during metabolic transitions there may be a period in which the population is heterogeneous, with cells in different stages of transition toward a low metabolic state, the 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyper diffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027). We have also corrected the DOI for this reference in the manuscript. Furthermore, the ATP pool of log-phase E. coli turns over several times per second (Holms et al., Arch. Mikrobiol. 1972, http://dx.doi.org/10.1007/BF00425016). We therefore assumed the bacteria were energy depleted after 10 minutes. We have clarified this point in the revised manuscript.

      Given that arsenate and azide inhibit cellular metabolism, i.e., have antimicrobial effects, cells might not just downregulate metabolism but also activate the stress response, and this causes some of the observed effects on phage adsorption. Therefore, the 'low metabolic state' of the cells in this paper could mean that cells are starved or that they are stressed or both.

      The reviewer is correct. We don’t exclude indirect effects. However, as nutrients were removed from the bacteria by washing and energy metabolism was inhibited by the addition of arsenate and azide, we assumed a stress response requiring biosynthesis would be unlikely to occur.

      The abundance of receptors could change between the high and low metabolic media conditions and contribute to the observed differences in adsorption, while the authors seem to assume in their model that the initial adsorption rate always remains the same.

      We do not think that the observed differences in adsorption are explained by a change in receptor abundance. In a previous study using the same experimental protocol as in the present work, phage λ was compared to the metabolically insensitive mutant λh (Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119). If the lower adsorption in the low-metabolic condition were caused by a reduced number of receptors, then λh should also have shown a lower adsorption rate under the same condition. Instead, no measurable effect on λh adsorption rate was observed. We therefore conclude that the effect is not explained by changes in receptor number on the timescale of the experiment. We have clarified this point in the revised manuscript.

      Reviewer #3 (Public review):

      Summary:

      Marantos et al. showed that for some coliphages, the energetic state of the bacterial host cell has a strong impact on whether phage infection is initiated. The authors drew this conclusion from the observation that there are more free phages remaining in the medium after infection of arsenate-azide-treated cells as compared to after infection of untreated cells. These data were analyzed and reported both as ratios of the treated vs. untreated conditions and using a mass-action kinetic model of phage-cell collision in the infection mixture. The data supported the findings that for four phages infecting Escherichia coli bacteria, namely, phages λ, ɸ80, m13, and T6, the phages are less likely to initiate infection if the host bacteria are energy-depleted. However, for phage T5, the authors found that their infection propensity is not impacted.

      Strengths:

      The data presented by the authors clearly supported the principal conclusion of the study ("Viral commitment to infection depends on host metabolism"). The five phages chosen by the authors represent different viral lifestyles and infection mechanisms, highlighting the potential applicability to other Escherichia coli phages. Finally, the authors successfully used a classic mass-action model of phage-cell collision to interpret their data. The simplicity of their experimental assay, combined with the use of this mathematical model, offers other investigators who study phage-bacterial interactions in other contexts a potentially useful toolkit to examine infection in general, and specifically, the dependence of phage infection on the host's metabolic state.

      Weaknesses:

      (1) The authors isolated and measured the numbers of free phages in the medium after infection of bacteria under different treatments. These measurements were analyzed in two different ways: (1) simply as ratios (corrected/normalized using different controls), and (2) fitted using a simple mathematical model. I have concerns regarding both analyses.

      (1.1) For the first method, having different time points at which the sample of each phage is collected critically complicates data interpretation. As one incubates the phage-bacteria mixture for a longer time, more infection occurs, and the number of phages collected from the mixture decreases. Therefore, the different incubation time forfeits the goal of "a systematic and quantitative comparison across different phages [...]", just as the authors self-criticized. Conceivably, the authors could have used the shortest measurement time for all phages (i.e., 10 minutes, as for phage λ). Alternatively, the authors could have applied a systematic criterion such as half (or any other fraction) of the latent period of each phage, which would still "maximize the incubation period while ensuring that manipulations were completed before the first infection cycle concluded". In my view, the seemingly arbitrary measurement time for each phage renders the entire first analysis very challenging to interpret. It also goes against the author's proposition that the protocol was "standardized" or "consistent". It is not clear what the readers are supposed to take away from this first analysis, or rather, which evidence, finding, or conclusion the manuscript would lose if the authors only presented the modeling-based analysis.

      (1.2) The second method of analysis sought to remove the dependence of the measurements on time. I completely agree with this goal, and the findings extracted from this analysis significantly contributed to the merits of this manuscript. However, the authors achieved this goal using a single time point for each phage to calculate the infection rate (η). As shown in Figure S3, each of the phage depletion curves is anchored by only one data point (note that the P(t)/P(0) = 1 at t = 0 is assumed, not measured). This goes against the typical way this collision model is used in the literature, where a time series is measured and used to fit the model (e.g., DOI 10.1007/978-1-60327-164-6 18, or more recently, PMID 39700139). This practice in the current manuscript reduced the robustness of the inferred η values. This problem is exacerbated by assumptions used by the authors in formulating this model. For instance, the authors used a constant value for the bacterial concentration, B, because "bacterial growth and lysis were negligible" (lines 135-136). However, considering that the bacteria were cultured at 37oC in a very rich medium (first in YT broth, then in 2% glucose), the measurement times of 20, 30, and 55 minutes are most likely one or a few generations of bacterial growth and division.

      Related note: I suggest that one of the panels in Figure S3 should be moved to the main text, since it is critical to the second method of analysis.

      We would like to clarify that the manuscript does not present two separate methods, but rather one method presented in two steps: a first step with results that are directly tied to the experimental measurements and show whether the effect is present for each phage, followed by a second, analytical step that makes the results comparable across phages.

      The first step presents the ratios because they directly reflect the measurements performed in the experiment and allow the reader to see the effect of the metabolic state for each phage in contrast to its control. We agree that these ratios are time-dependent and therefore not suitable for quantitative comparison between phages. Their purpose is to illustrate the experimental outcome and to show that the effect is present (or absent) on a per-phage basis not to compare magnitudes across phages.

      We then follow this with the second step, allowing the reader to follow the logic of the analysis. The analytical step that follows does not represent a second method, but a continuation of the same analysis. Here, we remove the time-dependence specifically in order to make comparison of the effect across phages possible, by connecting our results to standard measures such as the adsorption rate η. Importantly, P(0) is measured for every phage in every experiment. The only modeling assumption used (a standard one in the field) is the exponential form for the decay in free phage number, which naturally yields P(t)/P(0) = 1 at t = 0.

      Regarding the reviewer’s concern that bacterial growth may not have been negligible over the relevant time window, we note that recent work on rich-to-minimal growth lags in E. coli reports substantial delays before growth resumes after nutrient downshift. One 2023 study (Wu et al., Nature Microbiology 2023, https://doi.org/10.1038/s41564-022-01310-w) considering wild-type E. coli shows in Fig. 2c a lag of up to about 2 hours after a shift from MOPS minimal medium with 0.2% glucose plus 18 amino acids to the same medium without amino acids. Another 2023 study (Zhu and Dai, Nature Communications 2023, https://doi.org/10.1038/s41467-023-36254-0) examining both rel+ and rel− strains reports a growth lag of about 49 minutes for rel+ and more than 5 hours for the relA deletion strain. While these conditions are not identical to ours, they support the general point that growth does not immediately resume after such shifts. We therefore think it is unlikely that, following transfer from YT, the cells underwent one or a few full generations during the time window of our adsorption measurements.

      On the related note: Following the comments of all reviewers on Figure S3, we have decided to remove it to avoid confusion.

      (2) The data were able to distinguish phages that successfully infected bacteria and those that remained free in the medium, and the authors appropriately interpreted the data as such throughout the Results section. However, in the Discussion (starting from the very first sentence, line 172), the authors used terms that include "adsorption" and "entry" more interchangeably (for example, see the three sentences in lines 310-313, for "viral entry efficiency is shaped by [...]", then "adsorption kinetics modeling"). I do not see how the authors' data could distinguish between adsorption (the phage particles attaching to the outside of the cell) and entry (the phage DNA being injected into the cell). Conceivably, any phage particles that irreversibly attach to a cell but do not yet inject their genome into the cell would still be removed from the medium and therefore not quantified. Another example: in lines 189-191, the authors interpreted that "[...] when the bacterium is in a low metabolic state, the phage does not bind irreversibly to the host", but how do the authors eliminate the case of no phage binding (i.e., the reversible step) to begin with?

      We agree with the reviewer that our use of the terms adsorption, entry, and infection should have been more careful. Our experiment can only identify the irreversible commitment of phage to a host cell. We have therefore revised the text to refer consistently to phage commitment.

      Similarly, in lines 283-293, how do the authors delineate whether energy depletion would increase the k_off term or decrease the k_inj term, because either would result in more free phages in the medium as observed in the data? I believe that the writing of the Discussion, as it stands now, is doing a disservice to the conclusions presented in the Results section.

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: if energy depletion leads to reduced commitment, this can arise either because k_off increases, because k_inj decreases, or because both change, as long as k_off/k_inj becomes larger in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also leads to the trade-off now discussed in the revised manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become significantly larger in the inactive case. Depending on whether this is achieved through changes in k_off or k_inj, the cost of discrimination appears either as slower commitment or as additional energy dissipation. We agree that the previous wording overstated the mechanistic interpretation, and we have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (3) The authors presented an argument that performing infection of all five phages in the same condition is an advantage, allowing for comparison across different phages. While this goal is a completely valid one, it is difficult to reconcile that with the fact that different phages require different optimal conditions for successful infection. For instance, phage T5 famously requires Ca2+ for successful infection into the host bacterium (and later successful replication); see PMID 13174489. However, all infections were performed in TMG, which lacks Ca2+. Perhaps the absence of T5 dependence on the host metabolism is because the infection condition used by the authors was not optimal for T5 to begin with? Similar arguments could be made for other phages.

      Our study alone cannot eliminate that possibility. However, we have cited multiple previous studies, for example references citing Braun et al., showing that T5 remains insensitive to the host metabolic state under different buffer conditions. We therefore believe it is unlikely that the lack of metabolic dependence we observe for T5 is simply due to suboptimal infection conditions.

      (4) Whereas the manuscript examined five coliphages, only phage T5 and phage λ were discussed extensively. I believe some discussion points for these two phages need clarification.

      We focused our discussion on the phages T5, λ and φ80 because these are the phages for which similar effects have been reported previously in the literature. This allowed us to connect our findings directly to existing work and to discuss mechanistic hypotheses in a meaningful comparative framework. For the remaining phages, to our knowledge no prior studies have examined their behavior under comparable metabolic conditions, and therefore a similarly detailed discussion would have been speculative. Nevertheless, all five phages are treated equally in the presentation of the experimental results and in the quantitative comparison of adsorption rates.

      (4.1) Phage T5: The data obtained by the authors show that the infection rate of phage T5 is not impacted by the metabolic state of the host cell. Considering that the authors used the terms "infection", "adsorption", and "entry" interchangeably to refer to the irreversible commitment of a phage to a host cell (see point 2), this discussion regarding phage T5 lacks one critical literature context: DNA entry of phage T5 is known to occur in two phases (first-step transfer and second-step transfer). Critically, the second step can only occur if phage proteins encoded by the phage DNA transferred in the first step are expressed (see PMID 10577483 and the cited papers therein). In that context, metabolic poisoning of the host bacteria should have impeded T5 infection. The authors should comment on this point.

      As the reviewer pointed out, our usage of the terms infection, adsorption, and entry should have been more careful. Our experiment can only identify irreversible commitment of phage to a host cell. For T5, we expect that this irreversible commitment already occurs upon first-step transfer of phage DNA. As a result, even if second-step transfer is impeded under metabolic poisoning, our method would not resolve that effect. We have added this clarification to the revised manuscript.

      (4.2) Phage λ: The experiment using phage λ in this current study shares many resemblances to that in Brown et al. 2022. That feature alone is not a problem, but at many places in the text, the writing is ambiguous as to whether it is discussing the results in Brown et al. 2022 or in the current manuscript. I am giving three examples below, but this is not exhaustive: (i) Lines 67-69, there is no Brown et al. 2022 reference immediately after "a mutant phage variant (λh) could bypass this dependency [...]" (not just in the previous sentence); (ii) Line 228 should clearly say "Our previous findings suggested that phage λ is capable of [...]", since it concerns Brown et al., 2022, not the current study; and (iii) Lines 245-246, there is no Brown et al., 2022 reference immediately after "we observed that a mutant variant [...] even energy-depleted host" (without a reference, it reads like the authors "observed" that finding in this current manuscript).

      The reviewer is right. In those places, the text was ambiguous as to whether it referred to the present study or to Brown et al. (2022). We have now inserted the reference at the relevant points and revised the wording where needed to make this distinction explicit.

      Also, regarding phage λ: The discussion between line 230 and line 249 is very interesting, but since it concerns the differences between λ PaPa and Ur-λ, the authors should consider mentioning and discussing a very relevant recent study, PMCID: PMC6312755.

      We agree that the study by Guan et al. is very relevant and interesting. However, our point in this part of the Discussion is only to clarify that we used λ PaPa and not the originally isolated λ strain. We have therefore limited the discussion here to that distinction.

      (5) Control experiments, or references to prior studies, are needed to support that the As/Az treatment at this concentration and duration (at least 10 minutes) is sufficient to deplete the metabolic state of the cell. For instance, this can be shown by impeded or null cell growth, arrested motility (using a standard swimming assay), or a fluorescent reporter for the energetic state of the cell.

      The 10-minute treatment used here was chosen based on prior work showing that arsenate–azide rapidly inhibits cellular energy metabolism and is sufficient to eliminate the hyperdiffusion of the λ receptor (Winther et al., Biophysical Journal 2009, http://dx.doi.org/10.1016/j.bpj.2009.06.027) where the effect was assessed by monitoring the rate of movement of the λ receptor on the bacterial surface. We have clarified this point in the revised manuscript.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      As mentioned earlier, I found the paper interesting and addressed an important and significant knowledge gap.

      My biggest concern is about the interpretation of the experimental data in light of the two-step model. In particular, around line 286, it is stated "k_inj is more sensitive to metabolic state than k_off". Assuming k does not depend on metabolic state, which is a fair assumption, the equation for eta only depends on the ratio between k_inj and k_off and not on the individual parameters separately. Consequently, there is no way of saying which one of the two is more affected by metabolic state, unless the model already assumes that k_off is not influenced by metabolic state. The results could equally be explained by k_inj decreasing in metabolically depleted cells, or k_off increasing in such cells. If this is an assumption of the model, this should be clearly stated and not reported as a consequence of the data, as it is at the moment. Also, how does this mathematical model connect to the fitting function used in Figure 2b?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: discrimination requires that k_off/k_inj be larger for inactive hosts than for active hosts, such that commitment is specifically reduced in the inactive case. Put differently, what matters is not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This introduces a trade-off: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination requires this ratio to become significantly larger for inactive hosts. If this is achieved through changes in k_off, discrimination comes at the cost of slower commitment by allowing more time to leave; if it is achieved through changes in k_inj, it can preserve fast commitment to active hosts but requires additional energy dissipation in order to actively modulate commitment. We have therefore revised the text accordingly to frame the argument in terms of this trade-off, rather than attributing the effect specifically to k_inj. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      I have a related experimental criticism. The kinetic model presented assumes an exponential decay of free phage, which is a commonly used assumption in the phage literature. Given that the phage types used in this study lyse relatively slowly, it would be good to actually see adsorption curves, in which free phage is measured at different time points between inoculation and lysis. This data would not only provide useful evidence for the kinetic model, but it should also replace what is now in Figure S3, which consists of fitting one experimental point with one line. As it currently stands, Figure S3 is not useful actually misleading.

      We appreciate the reviewer’s point. We agree that adsorption curves, in which free phage is measured at different time points between inoculation and lysis, would provide a stronger basis for evaluating the kinetic model. However, we do not have the resources to perform these additional experiments within the scope of the present study. Following the comments of all reviewers on this point, we have therefore decided to remove Figure S3 to avoid confusion.

      Finally, it is not clear to me why the quantity "Ratio" has been chosen to be presented in Figure 1, rather than the ratio of estimated adsorption rates eta'/eta, which is much more intuitive for a phage study and contains the same information. I would recommend switching to this choice, unless there is a clear rationale for why the quantity "Ratio" is more useful/effective. Showing eta'/eta would also increase the readability of Figure 1, as it would move the y-axis to a logarithmic scale and better visualize values around 1.

      We used “Ratio” in Figure 1 to illustrate the experimental design, controls, and measured quantities directly, as it more transparently reflects the data collected. In the second part of the analysis, where we compare time-independent adsorption rate estimates, we have presented the corresponding values of η′/η as suggested.

      Minor comments:

      (1) Introduction

      Line 31: "... such as nutrient limitation, fluctuating temperatures, and variable energy availability" - if drawing a distinction between energy availability and nutrient limitation, please make explicit what this distinction is. Energy availability seems like a natural consequence of nutrient availability.

      While energy and nutrient availability are often linked in E. coli, they represent distinct physiological constraints. Nutrient limitation refers to the lack of essential biosynthetic precursors such as nitrogen, phosphorus, or amino acids. Energy availability, in contrast, reflects the cell’s ability to generate ATP and reducing equivalents through metabolic processes. For example, under anaerobic conditions, E. coli may have ample nutrients but limited energy production due to the lower efficiency of fermentation compared to aerobic respiration. Thus, energy limitation can occur independently of nutrient limitation.

      (2) Results

      (a) Whole Section: Please label equations.

      All equations have now been labelled in the revised manuscript.

      (b) Lines 105 to 114: As stated in Major Comments, I think the clarity of the paper would be improved by introducing the relative adsorption rate here and dropping the concept of Ratio entirely. However, if the authors wish to use Ratio, I would recommend the following:

      Lines 105 to 109 are confusing to read because of the number of connectives: "... ratio of free viruses from permissive AND resistant hosts respectively TO the free viruses in buffer under energy-depleted AND energy competent conditions". This would be clearer if each quantity were given an algebraic symbol, and RP, RR, and Ratio were defined through formal algebra, rather than mixed mathematical and sentence notation.

      This section has been rewritten for clarity. We now introduce explicit algebraic symbols and define the quantities formally, which removes the ambiguity present in the sentence-only description while retaining the intended meaning.

      The chemical names "arsenate" and "azide" should appear in the body of the text before they appear abbreviated in an equation. Please state at this point that these are both metabolic inhibitors, as it is not immediately clear what role they play or why you are using them.

      The text has been updated to introduce arsenate and azide by name before the abbreviations are used, and we now explicitly note that they act as metabolic inhibitors.

      On line 114, the authors helpfully provide an interpretation of Ratio = 1. It would be useful to provide at the same time interpretations of Ratio >1 and <1, perhaps 2 and 0.5 specifically?

      We have added brief explanations illustrating the interpretation of Ratio values greater than and less than 1, including examples of 2 and 0.5.

      I would consider giving this quantity a more interpretable name than Ratio. This quantity represents how much a bacteriophage preferentially adsorbs to metabolically active cells, so perhaps "Selectivity" or "Adsorption Bias"?

      We intentionally retained the generic term “Ratio”, as this quantity reflects an intermediate experimental measure used to describe the process rather than a newly defined metric. Its purpose is to bridge the experimental observations and the subsequent quantification of effects on the adsorption rate (η).

      (c) Lines 117 to 122: the authors sometimes refer to ratios explicitly, "average ratio of around 1.6" and other times say e.g., "a greater than 3 times increase in viral particles". Using more consistent language (saying "Ratio" every time) would be clearer.

      We have standardized the terminology in this section and now refer to all fold-changes consistently using “Ratio” to avoid ambiguity.

      (d) Figure 1

      Phages λ and T6 look like they have ratios less than 1 for resistant cells? If this is true / if the ratio is statistically significantly below 1, please comment.

      The ratios for λ and T6 are not statistically different from 1. The apparent deviation is within the standard error of the mean. To make this clearer, we have added the corresponding p-values to Table S2 in the Supplementary Information.

      Ratios near 1 are difficult to distinguish from 1, especially in panels A and D. Using a logarithmic scale on the y-axis would make the plots more readable.

      Because the values in these panels are not statistically different from 1, changing to a logarithmic scale would not alter the interpretation. We therefore retained the current axis scaling to reflect that there is no meaningful deviation from 1 in these cases.

      The data corresponding to individual experiments have no error bars. Given that the number of free virions was determined by plaque assay, which carries an intrinsic sampling error, this uncertainty should be reflected in the plots.

      We thank the reviewer for this important comment. Because plaque assays have compound sources of stochastic variation, assigning a per-measurement error bar would risk implying false precision. For this reason, we present the values from each biological replicate directly, and the uncertainty is represented in the statistical summary across replicates. Specifically, for each phage and condition we show the three independent experimental measurements and report the mean along with the standard error of the mean. This approach allows us to represent biological variability without implying a precision that cannot be accurately quantified at the level of single plaque counts.

      Similarly, the average value does show error bars, but it is not stated what these error bars correspond to: standard error in the mean, standard deviation of the sample, or combined uncertainty?

      The caption has been updated to state that the error bars represent the standard error of the mean.

      The resistant bacteria seemed to have ratios close to 1 in all cases. Is this because very few virions adsorbed under both energy conditions?

      Resistance is commonly associated with a lack of a surface receptor for the phage (or generally an entry pathway). We use the resistant bacteria as a control group for the effect of the conditions on adsorption. For resistant bacteria, the Ratio should be 1 since virions do not adsorb under both energy conditions. Any slight variations from 1 should come from sampling errors or small heterogeneity in the population.

      (e) Figure 2

      Please comment on what the error bars here represent. Error bars in Figure 2 A seem to permit negative (or at least zero) values of relative adsorption rate for phages m13 and T6, possibly implying an overestimate of the error? If it is the case that multiple values used to calculate the mean are far apart, possibly showing the values individually through a superimposed swarm plot would be clearer.

      This point is now addressed in the Supplementary Information, where we clarify how the error bars were calculated.

      (3) Discussion

      (a) Line 189: "high metabolic state" is imprecise. Say "energy-competent" to be consistent with earlier language.

      To maintain continuity with earlier terminology, we now include “energy-competent” in parentheses alongside “high metabolic state,” while retaining the original phrasing for readability.

      (b) Figure 3, population level

      Show adsorbed virions physically attached to bacteria, rather than removing them completely from the image, as currently, the implication is that at a high metabolic state, there are fewer virions total, not fewer virions remaining in solution because more are adsorbed. You could go as far as to add a third "after centrifuging" row, showing the adsorbed phages stuck in the pellet and the unadsorbed phages remaining in solution.

      Thank you for this suggestion. Figure 3 has been updated to depict adsorbed virions attached to bacterial cells, clarifying that the decrease represents adsorption rather than loss of total particles. This change improves the accuracy and interpretability of the schematic.

      (4) Methods and Materials

      (a) Figure 5

      The step "estimate cell numbers from OD" appears to follow incubating plates overnight. If the cells you are counting come from the pellet produced by centrifuging 3 steps prior, you could add a fork into the black line connecting the steps, with one branch corresponding to the supernatant and phages, and the other to the pellet and cells?

      Thank you for pointing this out. The order in the figure has been corrected: cell numbers are estimated from OD before overnight incubation. This resolves the confusion without the need for branching in the workflow diagram.

      (a) Line 332

      You allow as much time as possible for adsorption without the possibility of lysis. Did you determine the lysis times / latent periods of these phages through one-step-growth-curves, or use published results, in which case please cite? Having obtained the lysis time by either method, what fraction of the lysis time did you allow for adsorption? Also, please add supplementary tables with lysis times used for the different phages.

      We thank the reviewer for this comment. We used published latent-period values as guides and verified compatibility with our own system when selecting incubation times. We have clarified this in the text and added the relevant citations. We did not use a common fixed fraction of the lysis time for all phages; instead, incubation times were chosen to allow sufficient time for adsorption but not for completion of the first lytic cycle. For λ, productive lytic development was blocked in the host background used, as in Brown et al., PNAS 2022, http://dx.doi.org/10.1073/pnas.2106005119. For ϕ80 and T5, we used published latent-period values as guides and verified their compatibility with our own system (De Paepe and Taddei, PLoS Biology 2006, http://dx.doi.org/10.1371/journal.pbio.0040193). M13 is a chronic filamentous phage and therefore does not have a standard lytic latent period; in our host–phage combination, it required more than 1 h before phage release. For T6, we relied primarily on the kinetics observed in our own system, since adsorption was unusually slow for this phage–host pair under our assay conditions. Although literature reports describe shorter T6 latent periods under specific assay conditions (Foster and Johnson, Journal of General Physiology 1951, http://dx.doi.org/10.1085/jgp.34.5.529), this is consistent with published work showing that adsorption and infection kinetics can vary substantially with host background, surface structure, and experimental conditions (Heller and Braun, Journal of Bacteriology 1979, http://dx.doi.org/10.1128/jb.139.1.32-38.1979; Storms et al., Biochemical Engineering Journal 2012, http://dx.doi.org/10.1016/j.bej.2012.02.010).

      (5) Supplementary

      Figure S1

      This data is useful in understanding the main body of the paper, and I think this should form part of a main figure (possibly with the individual experimental data points superimposed over the bars). This could come before or as part of Figure 1?

      We thank the reviewer for this suggestion. We have explored including these data directly in the main figure but found that doing so substantially reduced the readability of the figure, as the underlying table is visually dense. For this reason, we chose to summarize the results in Figure 1 and present the detailed data separately in Figure S1 of the Supplementary Material, along with the Ratio analysis, which more effectively conveys the trends without overloading the main figure.

      Reviewer #2 (Recommendations for the authors):

      Minor comments:

      (1) L16-18: This sentence could be made more accessible as 'error correction' is not an intuitive term in the phage field.

      We have updated the overall theory section including the terminology. Instead of error correction, we now refer to it as a discrimination process.

      (2) L96-98: Does this potentially indicate a trade-off where evolution for stronger binding cannot evolve at the same time as responsiveness to metabolic activity?

      We agree that this sentence made a stronger evolutionary claim than our data support. Since we only tested four laboratory phages, we cannot conclude that there is an evolutionary trade-off between stronger binding and responsiveness to host metabolic activity. We have therefore removed this sentence to avoid making an unsupported evolutionary interpretation.

      (3) L102: What does 'post-cellular' mean?

      Postcellular supernatant is simply the liquid that remains after cells have been removed. During centrifugation, the cells pellet at the bottom, and the liquid above (which can contain viruses) is the postcellular supernatant.

      (4) L105-107: Worth splitting into two sentences as it is a bit unclear if ratios are built between permissible and resistant hosts or between buffers or both.

      Thank you for the suggestion. We have rewritten this section into two sentences to clarify how the ratios are constructed, and we hope the revised wording improves readability.

      (5) L110-122: Figures S1 and S2 could be referenced here.

      References to Figures S1 and S2 have now been added in this section.

      (6) L137: As P(0) is the viral concentration in buffer, I am assuming that the phage lysate has been diluted in buffer and phages have been added to cultures from the same dilution tube to guarantee equal starting numbers, but I couldn't find this in the methods.

      This clarification has been added to the Methods and Media section of the Supplementary Information.

      (7) L243: It would be worth defining what 'hyperdiffusion' means.

      We have added a brief definition of “hyperdiffusion”.

      (8) L253-256: I do not entirely follow this explanation.

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #3. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (9) L284: Why is k_inj necessarily more sensitive to the metabolic state than k_off? Could membrane changes under stress increase k_off?

      We thank the reviewer for this important point. We agree that the model would work either by k_off or k_inj being dependent on the host metabolic state, and that our original wording was therefore too restrictive. The data do not distinguish between these possibilities; they only constrain the ratio k_off/k_inj. In the revised text, we therefore formulate the argument in terms of this ratio: reduced commitment in inactive cells can arise through an increase in k_off, a decrease in k_inj, or both, as long as k_off/k_inj becomes larger in the inactive case. What matters is therefore not which individual rate changes, but that the balance between leaving and committing shifts in a way that disfavors commitment to inactive cells. This also underlies the trade-off now discussed in the manuscript: efficient commitment to active hosts requires a small k_off/k_inj, whereas strong discrimination against inactive hosts requires this ratio to become much larger in the inactive case. We have revised the Discussion accordingly to bring it in line with what the Results actually support. Based on the comments from all reviewers, we have also revised the terminology throughout the manuscript: instead of error correction, we now refer to this as a discrimination process, and we replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (10) Figure 1: There seems to be more variation between replicates in phage Lambda than in other phages. Is this caused by receptor number heterogeneity in the population?

      Unfortunately we do not have a way to compare receptor number heterogeneity across the different phage receptors in our experiments. We therefore cannot conclude that the larger variation observed for phage λ is caused by receptor number heterogeneity in the population.

      (11) Figure S1: There seems to be a significant difference between phage Lambda viability in the two buffers - do the authors have an idea where this comes from?

      There is no difference in λ viability between the two buffers. The apparent difference in the figure is due to sampling variability.

      (12) Figure S3: Last sentence of the legend probably shouldn't say 'upper'.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      Reviewer #3 (Recommendations for the authors):

      (1) The text reads as incomplete in some places. Can the authors please provide clarifications on the following points?

      (1.1) Lines 235-256: How do the authors draw a conclusion that "a phage can detect host metabolic status" from a study that used purified LamB receptors (i.e., no live cells with any metabolism) extracted from two different bacterial species (i.e., not a difference in metabolic states)?

      We thank the referee for pointing out this lack of clarity. This was also raised by Reviewer #2. The point we intended to convey is that λ behaves differently toward E. coli LamB depending on whether it is on a living cell or isolated in buffer, but makes no such distinction for Shigella LamB, binding it in both contexts. More specifically, previous work showed that wild-type E. coli extracts could only inactivate λ in the presence of added solvents, whereas control extracts prepared similarly from Shigella did not require added solvent for λ inactivation. This observation is consistent with E. coli LamB requiring a specific state to irreversibly bind λ. We therefore meant to suggest that the capacity for metabolic-state sensing is not simply a function of phage identity, but also depends on receptor-specific properties that differ between the two bacterial species.

      We have rephrased it as follows: Notably, wild-type λ is inactivated by E. coli K-12 extracts only when solvents are added, whereas Shigella extracts inactivate λ without this requirement (Randall-Hazelbauer and Schwartz, J. Bacteriol. 1973; Schwartz, J. Mol. Biol. 1975; Schwartz and Le Minor, J. Virol. 1975). This suggests that E. coli LamB requires a specific state for irreversible binding, a conditionality absent in Shigella LamB, indicating that the capacity for metabolic-state sensing may depend on receptor-specific properties.

      (1.2) Line 270, in the abstract, and in the caption of Figure 4: The authors described the model using terms such as "an error-correction mechanism" or "standard error correction", but there is little explanation. Can the authors clarify what kind of "error" is discussed here, and how it is "corrected"? In the "standard error correction" model, what determines which method of correction is "standard"? If "error correction" is a standard term in phage-bacterial interaction modeling, please provide references.

      We agree with the reviewer that our use of the term error correction was not appropriate in this context. The proper term is discrimination process rather than error correction. We have now corrected this terminology throughout the manuscript and clarified the underlying logic in the relevant sections.

      (1.3) Line 301: The authors speculated that phage T5 is "better suited to ecological niches", but I am not sure how that is consistent with their data showing T5 is more rampant, that they infect both energy-competent and energy-depleted cells, not just depleted cells. Why "niches", and why are T5 better suited to environments "where energy-limited cells dominate", not just any environment?

      We agree that this point was not stated clearly enough. What we intended to convey is that T5 would be at a net disadvantage in a niche containing a mixture of energy-competent and energy-deficient hosts. We have updated the main text accordingly.

      (1.4) Line 303, and related to point 6.3. above: Phage λ can also infect and replicate in "starved bacterial cells" (shown in Kourilsky 1974 and Geng et al. 2024, both of which were cited in this manuscript). How do the authors reconcile these reports with the discussion point in line 303, and their data that only phage T5, but not λ, shows insensitivity to the host metabolic state?

      Our data do not imply that phage λ is unable to infect starved bacteria. As shown in Kourilsky (1974) and Geng et al. (2024), λ can indeed infect and replicate in nutrient-limited cells. Our results specifically indicate that λ infection under starvation proceeds with a reduced adsorption rate, while T5 maintains the same adsorption rate even when the host is starved. Thus, our conclusion is that T5 is insensitive to the host metabolic state at the level of adsorption, whereas λ is not. We acknowledge that the wording in line 303 may have unintentionally led to confusion, and we have revised this part of the text to avoid that.

      (2) The following comments relate to the text and figures in the manuscript. There are many places in the manuscript that could use fine proofreading and copy-editing for clarity and consistency. For example:

      (2.1) If I understand it correctly, the equation in between lines 109 and 110 should be clarified using terms such as "Free viral particles after mixing with bacteria in Arsenate and Azide" and "Free viral particles in bacteria-free buffer with Arsenate and Azide". As it stands, it is not clear which terms correspond to conditions where bacteria are present.

      The equation has been updated to explicitly indicate which terms refer to mixtures containing bacteria and which refer to bacteria-free controls, so that the correspondence between conditions is now clear.

      (2.2) Equations in between line 276 and 283, and elsewhere: Some concentration terms are enclosed in brackets ("[BP]"), while most are not.

      This notation has been clarified. We now use “[PB]” specifically to denote the transient phage–bacterium complex, distinguishing it from the product P⋅B. All other concentration terms are written without brackets for consistency.

      (2.3) Figure 4 and in equations: "BP" or "PB"?

      The notation has been made consistent throughout; we now use “PB” exclusively to denote the phage–bacterium complex.

      (2.4) Line 284 and line 286: The "inj" in "k_inj" is sometimes italicized, sometimes not.

      The notation has been standardized so that k_inj is now formatted consistently throughout the manuscript, without italicizing “inj.” Also we have replaced k_inj by k_com to reflect that our assay resolves irreversible phage commitment rather than DNA injection specifically.

      (2.5) Figure 5: Was the step "Estimate cell numbers from OD" really performed on the next day after the experiment (i.e., >12 hours after infection and phage plating), not immediately after cell washing?

      Thank you for pointing this out. The figure has been updated to reflect the correct order of steps: cell numbers are estimated from OD immediately after washing, followed by overnight incubation of the plates.

      (2.6) Figure S1: As it stands now, the x-axis of each panel can be read either as "Permissive, Resistant bacteria, Buffer" (missing "bacteria" for the first pair of bars), or "Permissive (bacteria), Resistant (bacteria), Buffer (bacteria)" (extra "bacteria" for the last pair of bars).

      The intended interpretation is the second one (permissive bacteria, resistant bacteria, buffer).

      (2.7) Figure S3: The panel letters "A" and "B" are missing in the figure. Also, it is not clear why the legend for the five phages and the legend for the measurement times are not combined.

      Following the suggestions from all of the reviewers we have removed Figure S3 as it created more confusion than clarity.

      (2.8) Strain table in the Methods and Materials: Please write genotypes with italicization, and consistently indicate mutations and deletions with the minus sign superscript or the Δ prefix. Also, for the S3222 strain: Is it really the entire Mal regulon mutated ("Mal-"), or just lamB-? In Brown et al. 2022, it was only the latter.

      Genotypes have been reformatted with consistent notation. For S3222, the correct designation is Mal-, as in the SI of Brown et al. 2022. In this case, Mal- is intended as a phenotypic designation rather than a specific genotype, and we have therefore formatted it accordingly, i.e. neither italicized nor written in lower case.

    1. Author response:

      The following is the authors’ response to the original reviews.

      In revising the manuscript, we have focused on three main priorities raised during review: (1) improving precision around evidential claims, particularly concerning vector maintenance and P2A-mediated protein separation; (2) substantially improving figure quality, accessibility, and legend clarity; and (3) correcting inconsistencies and expanding methodological detail where requested.

      This study was intended as a foundational genetic toolkit and methodological framework for Blastocystis ST7-B, establishing practical workflows for DNA delivery, endogenous regulatory-element benchmarking, antibiotic-selected recovery, clonal propagation, and reporter-based analysis in a genetically challenging anaerobic microbial eukaryote. The central evidence presented is therefore functional in nature: reproducible transgene delivery, selectable recovery and propagation of colony-derived transgenic lines, and detectable reporter expression using multiple anaerobic-compatible reporter systems.

      We agree with the reviewers that several additional experiments, including Western blot analysis of P2A-containing constructs, outward-facing PCR, plasmid rescue assays, and selection-withdrawal experiments, would further strengthen the mechanistic interpretation of the system and help distinguish episomal persistence from genomic integration. We have therefore revised the manuscript throughout to clearly separate what is directly demonstrated from what remains a plausible working interpretation or important future direction.

      Importantly, the revised manuscript no longer presents episomal maintenance or complete P2A-mediated protein separation as demonstrated conclusions. Instead, these are now discussed explicitly as unresolved mechanistic questions requiring future molecular analysis. Nevertheless, the central methodological conclusion remains unchanged: stable selectable transgene expression, recovery of colony-derived transgenic lines, and reporter-positive Blastocystis ST7-B transformants can now be reproducibly obtained.

      Reviewer #1 (Public review):

      Summary:

      This paper presents a toolkit for the transformation of Blastocystis. The authors have screened a number of selectable agents, promoters and reporter genes and present their findings. This resource will be of immense use to those in the Blastocystis field, as well as those seeking to establish transformation tools in other species where such tools do not yet exist. Establishing new transformation tools is extremely challenging, and the authors have done an excellent job.

      Strengths:

      The authors have carried out a systematic screen of promoters, reporter genes and selectable agents. They have screened numerous for each, and all the data is presented. It is good to see when things did not work as well as when things did, so this data set is extremely useful indeed.

      Weaknesses:

      The findings are reported by reporter gene assay (microscopy). No evidence is given using genetics. The authors claim that the DNA is maintained episomally. However, could it be possible that there is integration? No PCRS/RT-PCRs are shown (although it can safely be assumed that the DNA/RNA is present where the transformation was successful), nor are any Western blots. These would have been useful to show that the P2A ribosomal skipping had occurred, and that proteins were expressed individually rather than as a polyprotein.

      We thank the reviewer for the positive assessment of the manuscript and for recognising both the technical difficulty and broader utility of establishing genetic tools in Blastocystis and other experimentally challenging microbial eukaryotes. We also appreciate the reviewer’s identification of the main evidential limitations in the original manuscript, particularly regarding vector maintenance and P2A-mediated protein separation.

      First, regarding the question of vector topology and the interpretation of episomal maintenance.

      We agree that the original manuscript presented episomal persistence too strongly relative to the evidence currently available. We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working model rather than a directly demonstrated conclusion.

      The transfection system used here was adapted from Li et al. (2019), including use of the pXS2-P<sub>Legumain</sub>-derived plasmid framework. Importantly, the construct used in the present study does not contain the original Trypanosoma brucei tubulin-targeting region associated with homologous integration in the original pXS2 system. Complete plasmid sequencing confirmed that the constructs function here as heterologous expression plasmids carrying Blastocystis ST7-B regulatory elements and transgenes. While this does not demonstrate episomal persistence, it also means that genomic integration cannot be inferred from the historical pXS2 vector architecture alone.

      We further note that comparative genomic analyses by Gentekaki et al. (2017) suggest that Blastocystis lacks components of the canonical non-homologous end-joining (NHEJ) machinery, implying that homologous recombination is likely to represent the principal route for double-stranded DNA repair. Because the constructs used here did not contain Blastocystis homology arms, there is currently no obvious mechanism favouring targeted homologous integration. Nevertheless, we fully agree that genomic integration cannot presently be excluded.

      To reflect this appropriately, the revised manuscript now explicitly separates the demonstrated functional outcomes from unresolved mechanistic questions concerning vector maintenance. We also identify several future approaches that would help distinguish episomal persistence from genomic integration, including outward-facing PCR, plasmid rescue followed by full plasmid sequencing, Southern blotting, FISH, selection-withdrawal experiments, and long-read sequencing approaches.

      We have revised the manuscript throughout to remove statements implying demonstrated episomal maintenance and now present episomal persistence only as a plausible working interpretation.

      In the Methods section under Cloning, the following text has been added:

      Lines 202–206: “The constructs used in this study were derived from the pXS2-P<sub>Legumain</sub> vector described by Li et al. (2019), which adapted a heterologous expression-vector backbone for transient plasmid-based expression in Blastocystis ST7-B. Here, the same molecular backbone was used as a plasmid scaffold carrying Blastocystis-derived regulatory elements and transgenes.”

      In the Discussion, the following text has been added/edited:

      Lines 665–673: “The molecular maintenance state of the introduced constructs remains unresolved: episomal maintenance is a plausible working model, but genomic integration cannot be formally excluded. The constructs used here lack Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components (Gentekaki et al., 2017), making targeted integration by standard repair routes unlikely but not impossible. Direct assays such as outward-facing PCR, plasmid rescue followed by full plasmid sequencing, FISH, or selection-withdrawal experiments will be required to distinguish episomal persistence from integration.”

      Second, regarding P2A-mediated protein separation.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. We have therefore revised the manuscript to avoid implying that complete protein-level separation was directly demonstrated.

      The revised manuscript now states only what is directly supported by the current data: that P2A-containing bicistronic constructs supported antibiotic-selected recovery of transgenic lines together with detectable downstream reporter expression. The microscopy data therefore support functional downstream reporter expression, but do not by themselves exclude residual uncleaved fusion products.

      We selected P2A because it is a compact and well-characterised peptide with high reported separation efficiency across multiple eukaryotic systems, including microbial eukaryotes. However, we agree that P2A performance can be context-dependent, and we now explicitly identify biochemical validation of P2A cleavage efficiency as an important future direction.

      Importantly, these revisions do not alter the central methodological conclusion of the study, namely that selectable transgene expression, propagation of reporter-positive lines, and recovery of colony-derived Blastocystis ST7-B transformants can now be reproducibly achieved.

      Text inserted in the Results:

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.”

      Lines 403–404: “However, protein-level separation was not directly tested, and the extent of any residual uncleaved fusion product remains unresolved.”

      Text inserted in the Discussion:

      Lines 619–629: “P2A was selected because it is a well-characterised peptide with high reported separation efficiency in human cell lines, zebrafish embryos, and mice (Kim et al., 2011). It also has precedent across microbial eukaryotes, including the protest Dictyostelium discoideum (Zhu et al., 2023), the fungi Aspergillus niger (Schuetze and Meyer, 2017) and Ustilago maydis (Müntjes et al., 2020), and the apicomplexan parasites Toxoplasma gondii (Markus et al., 2019) and Plasmodium falciparum (Dans et al., 2024). However, P2A performance is context-dependent, and the evidence presented here is functional rather than biochemical. P2A-containing constructs support antibiotic-selected recovery and downstream reporter expression in Blastocystis ST7-B, but ribosomal skipping efficiency and any residual uncleaved product will require direct protein-level validation.”

      Reviewer #1 (Recommendations for the authors):

      (1) Please could you show a Western blot to confirm if P2A has worked? It could be that the proteins are being expressed as a polyprotein.

      We agree that Western blotting would provide the most direct biochemical assessment of P2A-mediated ribosomal skipping efficiency in Blastocystis ST7-B and would help determine the extent of any residual uncleaved fusion product. This is an important point, and we have revised the manuscript accordingly to avoid implying that complete protein-level separation was directly demonstrated.

      The current study was designed as a first-generation functional genetic toolkit for Blastocystis ST7-B, focused primarily on establishing reproducible workflows for selectable transgene expression, reporter recovery, and propagation of transgenic lines in this experimentally challenging anaerobic microbial eukaryote. The toolkit is therefore validated here through functional outcomes, including antibiotic-selected survival, stable propagation through extended passaging (>15 passages) and cryopreservation, and detectable reporter fluorescence above wild-type autofluorescence.

      P2A was selected because it is a compact and well-characterised peptide with high reported ribosomal skipping efficiency across multiple eukaryotic systems, including microbial eukaryotes, as discussed above. Nevertheless, we fully agree that direct biochemical validation would strengthen the mechanistic interpretation of the bicistronic system in Blastocystis ST7-B. We therefore now explicitly identify Western blot analysis, ideally using epitope-tagged upstream and downstream products, as an important future direction for quantitative assessment of P2A cleavage efficiency and any residual uncleaved fusion products.

      Relevant manuscript revisions are described above under the general response to Reviewer 1.

      (2) Something has gone wrong with figure formatting. Figure 2 is nearly illegible and I cannot read the text in section A. Sections B, C, and D have lost their labels and are fuzzy and surrounded by black. A similar issue affects Figure 3. Everything is just black with a few cells. It is illegible when printed.

      We thank the reviewer for highlighting these presentation issues and agree that the submitted figure quality significantly impaired readability and interpretation. The problems appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, contrast, and panel labelling in the review PDF.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved typography, panel separation, colour scaling, and legend clarity throughout. In response to additional reviewer suggestions, individual data points have now been added to Figures 2B and 2C to improve transparency and interpretability of the underlying data distributions.

      Figures 2 and 3 have been replaced with fully revised high-resolution versions with improved panel labelling, accessibility, typography, and figure legends.

      (3) The data from Figure 2B would be better placed in Table 1 with a column for robust/moderate/intermediate/weak/very weak. This would be much easier for the reader.

      We thank the reviewer for this helpful suggestion. We believe the comment refers to the promoter activity data shown in Figure 2A rather than the voltage optimisation data in Figure 2B. To improve readability and accessibility of these data, we have revised Figure 2A extensively to make the promoter activity tiers more legible and easier to interpret directly from the heat map and accompanying box plots.

      We considered incorporating simplified activity classifications into Table 1. However, activity patterns were construct-specific rather than simply locus-specific. In several cases, multiple promoter fragments derived from the same locus produced substantially different reporter outputs, and activity did not scale monotonically with promoter fragment length. We therefore felt that assigning a single categorical activity label at the locus level would oversimplify the dataset and reduce the construct-level resolution that is central to the toolkit value of the study.

      Instead, we addressed the reviewer’s concern by substantially improving the presentation and readability of Figure 2A, allowing readers to identify robust, moderate, intermediate, weak, and very weak expression constructs more directly while preserving the underlying construct-specific information.

      Figure 2A has been revised to improve clarity, accessibility, and legibility of the promoter activity tiers, allowing construct-level expression classes to be interpreted more directly from the heat map and accompanying boxplots.

      (4) How do you know if the constructs are maintained as episomes? Have you done an outward-facing PCR?

      We agree that direct molecular evidence distinguishing episomal persistence from genomic integration is currently lacking, and we appreciate the reviewer highlighting this important limitation. We have therefore revised the manuscript throughout to avoid presenting episomal maintenance as a demonstrated conclusion and now describe it only as a plausible working interpretation based on the current evidence and vector design.

      We have not performed outward-facing PCR in the present study. As discussed in the general response above, we now explicitly identify outward-facing PCR, plasmid rescue followed by full plasmid sequencing, selection-withdrawal assays, FISH, and long-read sequencing approaches as important future directions for resolving the molecular maintenance state of the constructs.

      The revised manuscript now clearly separates the demonstrated functional outcomes, including selectable transgene expression, recovery of colony-derived transgenic lines, and stable reporter-positive propagation under selection, from the unresolved mechanistic question of vector topology.

      This issue has been addressed throughout the revised manuscript, including in the Methods and Discussion sections, where episomal maintenance is now presented as a plausible but unconfirmed interpretation rather than a demonstrated conclusion.

      Minor Comments

      Line 66: is this one to two billion individuals with Blastocystis, or one to two billion Blastocystis cells per gut?

      The intended meaning was colonised individuals globally. We agree that the original phrasing was ambiguous and have corrected it for clarity.

      Lines 66–67 revised to: “…microorganisms in the human gut, and is estimated to colonise approximately one to two billion people globally (Scanlan and Stensvold, 2013).”

      Line 148: Supplier of IMDM?

      The supplier information was already present in the original manuscript as IMDM L0191 (Biowest).

      No additional manuscript change required.

      Line 157: Who annotated the dataset, the 2017 paper or the present study?

      The dataset annotation derives from Armengaud et al. (2017). We agree that the original wording was unclear and have revised this section substantially to improve clarity regarding the rationale and workflow used for promoter and terminator candidate selection.

      “The relevant Methods section has been extensively revised for clarity and expanded detail” (Lines 156–189).

      Line 166: Who predicted the 3′ UTR, the 2017 paper?

      This information derives from the NCBI annotation associated with the Blastocystis ST7-B genome based on Denoeud et al. (2011). This has now been clarified in the Methods section.

      Clarified in revised Methods section.

      Line 237: How long did it take in days?

      Approximately 2 days.

      Line 270 revised to: “…turned yellow without drug treatment, usually within 2 days post-transfection.”

      Line 325: Typo, missing gap between Figure and 1A.

      Corrected in revised manuscript.

      Reviewer #2 (Public review):

      This manuscript presents a substantial technical advance for the genetic manipulation of Blastocystis by establishing an integrated workflow for stable episomal transgenesis, antibiotic selection, clonal recovery, and reporter-based imaging in the ST7-B subtype. The study is particularly valuable because it combines multiple previously fragmented approaches into a coherent and practically applicable toolkit, including endogenous regulatory elements, optimized electroporation conditions, selectable markers, and anaerobic compatible fluorescent reporters. This methodological work greatly expands the molecular toolbox and future studies focused on both basic and infection biology can now build on the ability to express and localize proteins in fixed as well as live cells.

      The microscopy data are convincing and clearly demonstrate functional reporter expression and successful recovery of stable transgenic lines. Nevertheless, because this is primarily a methodological paper, the study would be further strengthened by the inclusion of Western blot validation of reporter expression and bicistronic constructs. In particular, biochemical analysis of the P2A-containing constructs would help assess the efficiency of ribosomal skipping and exclude the possible presence of uncleaved fusion proteins, thereby providing stronger support for the interpretation of the imaging data and the functionality of the expression system.

      We thank the reviewer for this thoughtful and positive assessment of the manuscript and for recognising the value of integrating previously fragmented approaches into a coherent and practically usable genetic toolkit for Blastocystis ST7-B. We particularly appreciate the reviewer’s recognition that the system expands the currently available molecular toolbox for both cell biological and infection-related studies in this experimentally challenging anaerobic microbial eukaryote.

      We also appreciate the reviewer’s comments regarding biochemical validation of the P2A-containing bicistronic constructs. We agree that Western blot analysis would strengthen the mechanistic interpretation of the reporter system by directly assessing ribosomal skipping efficiency and the possible presence of residual uncleaved fusion products. In response, we have revised the manuscript throughout to ensure that the conclusions remain appropriately evidence-based and do not imply that complete protein-level separation was directly demonstrated.

      The revised manuscript now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, stable propagation of reporter-positive lines, and detectable downstream reporter expression, from unresolved mechanistic questions concerning P2A cleavage efficiency and vector maintenance state. We now also identify biochemical validation of P2A-mediated protein separation as an important future direction for further refinement of the system.

      Relevant manuscript revisions addressing these points are described above under the response to Reviewer 1.

      Reviewer #2 (Recommendations for the authors):

      The quality of images could be better. The figures lacked resolution — possibly a conversion artefact.

      We agree that the figure quality in the submitted review PDF significantly reduced readability and visual interpretation. The issues appear to have arisen primarily during manuscript compilation and export, particularly affecting image resolution, typography, panel labelling, and contrast rendering.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions. We have also improved panel separation, typography, colour scaling, contrast settings, and figure legends to improve accessibility and interpretability both on screen and in print. In addition, the export workflow and file formatting have been updated to improve compatibility with journal production requirements and reduce the likelihood of compression-related rendering artefacts during manuscript compilation.

      Figures 2 and 3 have been replaced with revised high-resolution versions with improved typography, panel labelling, contrast settings, and accessibility.

      Reviewer #3 (Public review):

      Summary:

      The primary objective of this study was to establish a practical and functional framework for the propagation of stable transgenic cell lines of Blastocystis, a common animal gut microeukaryote. Although the work focused on Blastocystis ST7-B, a subtype with relatively low prevalence in humans, this choice is justified by its association with more frequent negative health effects. Beyond their relevance to the medical field, the methodological advances described here have the potential to also expand cell biology studies of this anaerobic organism, including its unusual mitochondria and redox metabolism.

      Strengths:

      Prior to this work, genetic tools for Blastocystis were very limited, relying on a single strong promoter-terminator combination. The authors successfully expanded the available promoter set across a range of expression strengths by testing two dozen variants in luciferase-based assays. Critically, they developed an integrated workflow from a modular transgenic construct design, to an expanded inventory of molecular components (promoters, reporters), optimized DNA delivery, stepwise antibiotic resistance-mediated clonal selection and propagation, and to reporter validation. The evaluation of several anaerobiosis-compatible labeling strategies for live (and fixed) cell optical imaging will be particularly useful, with the SNAP-tag system appearing especially promising for Blastocystis.

      Weaknesses:

      The presented data generally provide solid support for the conclusions that the work reached, but clarification of reasoning and several inconsistencies, as well as amendments to the visual presentation of the data, would be highly beneficial, as detailed below.

      (1) Episomal persistence of the construct:

      The manuscript repeatedly assumes, including in its title, that constructs persist in Blastocystis in their episomal form, but no direct evidence is provided. Although this interpretation is plausible, it should be identified more clearly as provisional. Nuclear genomic integration (e.g., via NHEJ) remains a possible explanation unless supporting evidence or rationale is provided to exclude it. Testing whether the phenotype persists without drug-mediated selection in the generated transgenic cell lines would help strengthen the case for episomal maintenance.

      We thank the reviewer for this important point and agree that the original manuscript presented episomal persistence too strongly relative to the currently available evidence. In particular, we agree that the title and several sections of the manuscript implied a level of mechanistic certainty that was not directly demonstrated.

      We have therefore revised the manuscript throughout to clarify that episomal maintenance should presently be regarded as a plausible working interpretation rather than a demonstrated conclusion. The revised text now explicitly distinguishes the demonstrated functional outcomes, including selectable transgene expression, recovery and propagation of colony-derived transgenic lines, and stable reporter-positive maintenance under selection, from the unresolved mechanistic question of vector topology.

      As discussed in our response to Reviewer 1, the constructs used here do not contain Blastocystis homology arms, and comparative genomic analyses suggest that Blastocystis lacks canonical non-homologous end-joining components, making targeted integration by standard repair routes less strongly supported mechanistically, although genomic integration cannot presently be excluded.

      We agree that selection-withdrawal experiments would provide useful additional evidence regarding construct persistence and have now explicitly identified such assays, together with outward-facing PCR, plasmid rescue, FISH, and long-read sequencing approaches, as important future directions for resolving the molecular maintenance state of the transgenes.

      The manuscript has been revised throughout to remove wording implying demonstrated episomal maintenance. Episomal persistence is now discussed only as a plausible working interpretation pending direct molecular validation.

      (2) Promoters and terminators:

      (2.1) There is a discrepancy between the claimed number of loci (14), from which promoters used to drive luciferase expression were derived, and those detailed as having been actually generated in Table 1 (11). This inconsistency should be corrected or explained, as it creates uncertainty around the accuracy of the dataset.

      We thank the reviewer for this careful reading and for identifying this inconsistency. We agree that the distinction between candidate loci and successfully generated promoter constructs was not sufficiently clear in the original manuscript and could create uncertainty regarding the dataset.

      The original candidate set comprised 14 loci selected for promoter and terminator discovery. However, only 11 loci yielded successfully cloned and experimentally tested promoter constructs. The remaining three loci were retained in Table 1 for completeness and transparency, as repeated cloning attempts were unsuccessful despite two independent efforts.

      We have revised the manuscript to make this distinction explicit and to clarify that the reported NanoLuc benchmarking experiments were ultimately performed using constructs derived from 11 successfully cloned loci.

      Lines 361–364: “To expand the available regulatory parts, we screened 23 NanoLuc reporter constructs containing putative endogenous promoter–terminator pairs from 11 of 14 candidate loci; three loci could not be cloned after two independent attempts and are indicated in Table 1.”

      (2.2) Based on the presented evidence, constructs benchmarked in bioluminescence assays differed only in their promoter composition. Although terminator selection is mentioned in the Methods section, no additional details are provided; for instance, Table 1 and Figure 2 only list 23 promoters in total. Figure 2A likewise shows only promoter-dependent variation. If the terminator was held constant (LeguP1?), this should be stated explicitly. The authors may then consider revising the wording of having tested “23 promoter-terminator pairs” to better reflect that only promoters varied.

      We thank the reviewer for the opportunity to clarify this point. We agree that the original presentation may have created the impression that promoter and terminator regions were independently varied and benchmarked, whereas the experimental design was primarily focused on construct-level comparison of endogenous regulatory modules.

      As described in the Methods, each construct contained a candidate endogenous upstream promoter region together with the corresponding endogenous downstream terminator region derived from the same locus. For consistency and to keep the cloning and screening strategy experimentally tractable, a fixed 500 bp downstream terminator fragment was used for each locus rather than systematically varying terminator length or independently testing terminator activity.

      We therefore retain the description “endogenous promoter–terminator pairs,” since each construct contains both endogenous upstream and downstream regulatory regions from the same genomic locus. However, we agree that the assay was not designed to independently dissect promoter versus terminator contributions to reporter output. We have revised the manuscript accordingly to make this distinction explicit and avoid ambiguity regarding the scope of the benchmarking analysis.

      Lines 365–368: “Each construct paired a candidate upstream promoter region with the corresponding downstream terminator region from the same locus, defined here as the native 500 bp sequence immediately downstream of the stop codon. Where multiple promoter lengths were tested for the same locus, the terminator fragment was kept constant (Table 1; Figure 1A).”

      This design allowed construct-level benchmarking of paired promoter–terminator modules but did not test promoter strength or terminator activity independently.

      (2.3) Promoter benchmarking was done with a plasmid lacking a selection marker, so it is unclear how the maintenance of the luciferase construct was ensured. Without selection, the observed reporter intensity could reflect differential or stochastic plasmid retention rather than promoter strength alone. The luminescence assay was performed 16-18 hours after transfection, but the rationale for this particular timeframe should be explained. In this context, the authors should explicitly state whether the experiments shown in Fig.2A represent biological triplicates or technical triplicates from a single transfection.

      We thank the reviewer for these important methodological points. We agree that the original manuscript did not sufficiently clarify the transient nature of the NanoLuc benchmarking assay or the rationale underlying the assay design and timing.

      The promoter benchmarking assay was designed as an early transient-expression screen adapted from the NanoLuc-based workflow of Li et al. (2019), with modifications, rather than as a stable-maintenance assay. No selectable marker was included because the objective was to compare relative early reporter output across constructs shortly after DNA delivery, before prolonged culture effects became dominant.

      The 16–18 h post-electroporation time point was selected based on the NanoLuc expression kinetics reported by Li et al. (2019) and empirical optimisation during assay development. This window allowed robust transient reporter detection while limiting confounding effects arising from prolonged plasmid loss, differential outgrowth, variable recovery, or later culture-level changes.

      We agree that, in the absence of selection, the observed NanoLuc signal cannot be interpreted as an absolute measure of promoter strength independent of DNA uptake efficiency, early plasmid retention, or post-transfection recovery dynamics. We have therefore revised the manuscript to clarify that Figure 2A reports relative transient reporter output under standardized early post-transfection conditions rather than isolated promoter activity alone.

      We now also explicitly state that the data shown in Figure 2A derive from three independent electroporation experiments per construct, each assayed in technical duplicate.

      Lines 241–248: “Promoter–terminator activity was assessed 16–18 h after electroporation using a transient NanoLuc assay adapted from Li et al. (2019), with modifications. This early time point was selected to capture reporter output within the transient-expression window after DNA delivery, before prolonged plasmid loss, differential outgrowth, or culture-level changes could dominate the readout. Because the constructs did not contain a selectable marker, the measured NanoLuc signal reflects early transient reporter output rather than promoter strength independent of DNA uptake, early plasmid retention, or post-transfection recovery.”

      Additional clarification added to Figure 2 legend stating that measurements derive from three independent electroporation experiments, each assayed in technical duplicate.

      (3) Figure 2:

      (3.1) Several aspects of the current design may lead to ambiguity for the reader. The boxplots are colour-coded, but it is unclear whether the colours carry meaning or are purely decorative. Because the data are already spatially separated into bins, additional random colouring is redundant and may suggest distinctions that are not intended. In addition, part A of Figure 2 is split into two panels, with the scale for the left panel shown in the right panel and some of the boxplot colours falling in the range of the scale, but not in line with their counterparts in the left panel. Because the colour use is not consistent, it is difficult to tell whether the same scale should be applied to both panels or how it should be interpreted.

      (3.2) The left panel of part A uses a diverging blue-white-red colour scheme, which is most appropriate when the midpoint represents a meaningful central value such as zero. Because the values shown in this graph are only positive, a non-diverging 2-colour scale or a colour palette such as 'viridis' would make the plot easier to interpret.

      (3.3) A black background should be avoided: 'B' and 'C' labels are invisible, and it draws attention to a distracting design feature rather than the data themselves.

      We thank the reviewer for these detailed comments regarding figure design and visual interpretation. We agree that the original presentation of Figure 2 introduced unnecessary visual ambiguity through inconsistent colour usage, the use of a diverging colour scale for strictly positive values, and poor readability associated with the dark background and low-resolution export.

      In response, Figure 2 has been extensively redesigned to improve clarity, accessibility, and interpretability. The previous blue–white–red diverging heatmap has been replaced with a sequential colour palette appropriate for positive-only expression data. Boxplot colouring has also been simplified and harmonised with the heatmap scheme to avoid implying unsupported categorical distinctions. In addition, panel organisation, typography, scaling, and legend structure have all been revised to improve readability and reduce ambiguity regarding interpretation of the plotted values.

      We also agree that the black background distracted from the data presentation and impaired visibility of panel labels and image boundaries. The revised figures therefore use white backgrounds together with clearer panel separation and improved label visibility throughout.

      Figure 2 has been completely reformatted using a sequential colour scale in panel A, simplified and harmonised boxplot colouring, larger typography, improved panel separation, revised legends, and white backgrounds throughout. Corrected high-resolution source figures have been provided.

      (4) Figure 3:

      (4.1) Individual snapshots should be separated more clearly, either by using a white background or by adding visible borders to make the overall composition clearer. As currently displayed, some boundaries between fluorescent channels resemble image artifacts rather than intentional panel divisions.

      We thank the reviewer for this helpful comment regarding figure composition and panel separation. We agree that the original presentation made it difficult to distinguish intentional panel boundaries from imaging artefacts, particularly in the low-resolution review PDF generated during manuscript compilation.

      To improve clarity, Figure 3 has been reformatted using white backgrounds, clearer panel spacing, and more explicit separation between individual snapshots and imaging channels. High-resolution source images have also been provided to ensure that fluorescence patterns, image boundaries, and panel organisation remain clearly interpretable both on screen and in print.

      Figure 3 has been reformatted with improved panel separation, white backgrounds, clearer image boundaries, and revised high-resolution source figures.

      (4.2) In parts B-D, the legend should explain more clearly what each image shows, and the figure itself would benefit from annotations. There seem to be three sub-panels in each 'condition' of part B (as well as C and D): while the middle and rightmost panel can be easily inferred to represent the fluorescent protein and bright-field image, what the leftmost panels represent is not specified. If DAPI was used to dye DNA, an explanation why mostly multiple labelled regions are visible should be provided.

      We thank the reviewer for these helpful suggestions regarding figure annotation and legend clarity. We agree that the original presentation did not sufficiently explain the composition of the imaging panels, particularly under the low-resolution conditions of the review PDF.

      To improve interpretability, the revised Figure 3 now includes clearer panel organisation, improved annotations, and expanded figure legends explicitly identifying the individual imaging channels and staining conditions shown in each subpanel. The leftmost panels in parts B–D are now more clearly identified in both the figure and legend, together with the corresponding fluorescence or staining conditions used in each experiment.

      As mentioned in the Methods sections we used Hoechst 33342 to visualise DNA; but we agree that the Hoechst 33342-labelled structures required additional clarification. The revised legend section now explains that multiple Hoechst 33342-positive regions are commonly observed because Blastocystis cells can contain multiple nuclei depending on cell stage and subtype-specific morphology.

      In addition, high-resolution source images have been provided to ensure that fluorescent signals, panel boundaries, and imaging features remain clearly interpretable both on screen and in print.

      Figure 3 legends and annotations have been revised to clarify imaging channels, staining conditions, and panel organisation. The figure caption was also edited to include: “DNA was visualised using Hoechst 33342. Most cells contained two nuclei, and smaller Hoechst 33342-positive signals consistent with mitochondrial DNA were also observed in some instances.”

      (4.3) Cell morphology and appearance differ markedly between UnaG/smURFP and SNAP-tag images, which should be explained. A microscope issue is mentioned in the main text, but if that was the cause, the authors should consider replacing the images, as the current distortions complicate interpretation.

      We thank the reviewer for this important observation and agree that the apparent morphological differences between the UnaG/smURFP and SNAP-tag panels required additional clarification.

      The images shown for the different reporter systems were acquired under different imaging conditions and microscope configurations following an instrument-related issue during part of the imaging workflow, as noted in the Methods section. As a result, direct visual comparison of cell morphology between reporter systems is not appropriate. The primary purpose of these panels is instead to demonstrate reporter detectability, live-cell labelling capability, and the characteristic fluorescence patterns obtained with the different anaerobiosis-compatible reporter systems.

      In particular, the SNAP-tag panels were included to demonstrate successful live-cell labelling without permeabilisation together with the expected increase in fluorescence signal at higher substrate concentrations, rather than to support quantitative comparison of cell morphology across imaging conditions.

      We considered replacing the affected images. However, equivalent replacement datasets acquired under directly comparable conditions are not currently available. We have therefore retained the original images but revised the figure legend to clarify the intended interpretation and limitations of these panels explicitly.

      Figure 3 legend revised to include:

      “Because images for the different reporter systems were acquired under different imaging conditions, they are presented to demonstrate reporter detectability and labelling pattern and should not be used for quantitative comparison of cell morphology across reporter systems.”

      Reviewer #3 (Recommendations for the authors):

      The reader may find the current order confusing starting with construct design before testing which drug to use for selection. The narrative would work better if it started with antibiotic selection as the first logical step for generating stable cell lines.

      We thank the reviewer for this thoughtful suggestion regarding narrative structure and agree that multiple organisational strategies are possible for presenting a methodological workflow of this type.

      We considered reorganising the Results section to begin with antibiotic selection and drug sensitivity profiling. However, we ultimately retained the overall structure because the manuscript is organised as a toolkit-development framework rather than as a strictly chronological experimental protocol. The Results therefore begin with regulatory-element discovery and construct design, which form the conceptual and experimental foundation of the toolkit, before progressing to DNA delivery optimisation, drug sensitivity profiling, clonal recovery, and reporter validation.

      We felt that this structure most clearly reflects the dependency relationships within the system: regulatory elements are required before constructs can be assembled, constructs are required before electroporation conditions can be evaluated, and selectable constructs are required before stable selection and clonal recovery can be meaningfully assessed.

      (2) The text states that the screen 'focused on the 1,000 most abundant proteins to establish a preliminary library capable of supporting varying levels of transcription.' Since the genome has ~6,000 protein-coding genes, the top 1,000 cover the most abundant proteins — not a wide expression range.

      We thank the reviewer for this important clarification. We agree that the original wording could incorrectly imply that the screen was intended to sample broadly across the full transcriptional range of the Blastocystis genome. This was not the case, and we have revised the manuscript accordingly.

      Our strategy was instead designed to enrich for candidate loci with a higher prior likelihood of supporting detectable transgene expression. Because no genome-wide promoter map, transcription start site dataset, or experimentally validated regulatory annotation was available for Blastocystis ST7-B at the inception of this work, we used the abundance-ranked Blastocystis ST4-WR1 proteomic dataset of Armengaud et al. (2017) as a practical starting point for candidate discovery.

      Importantly, the Blastocystis ST4-WR1 proteome is highly skewed, with 193 proteins contributing approximately 50% of the detected proteome and the 13 most abundant proteins contributing approximately 10% (Armengaud et al., 2017). We therefore selected the top 1,000 proteins not as a representation of the genome-wide expression range, but as a proteomics-guided enrichment strategy to identify loci more likely to contain active endogenous regulatory regions suitable for initial toolkit development.

      We have revised the relevant Methods section substantially to clarify both the rationale and the workflow used for candidate selection, homolog identification, and promoter/terminator definition.

      The Methods section (Lines 156–189) has been extensively revised to clarify the rationale underlying candidate regulatory-element selection. The revised text now explicitly states that the strategy was designed to enrich for likely active loci for toolkit development rather than to systematically survey the full range of promoter strengths across the Blastocystis genome.

      Additional methodological detail has also been added regarding:

      Use of the Armengaud et al. (2017) proteomic and proteogenomic datasets,

      Homolog identification in Blastocystis ST7-B,

      Locus selection criteria,

      Promoter boundary definition,

      And operational definition of candidate terminator regions.

      (3) The Methods contain an inconsistency: cells were left in 0.5 mL, then 1 mL was added, but then only 0.5 mL is apparently used for transfection. What happened to the 1 mL?

      We thank the reviewer for identifying this ambiguity in the transfection workflow description. The apparent inconsistency arose because the protocol description moved from bulk cell resuspension to preparation of individual electroporation reactions without explicitly stating how the intermediate suspension was used.

      After washing, approximately 0.5 mL of cytomix buffer remained above the pellet, and 1 mL of complete cytomix buffer was then added to generate an approximately 1.5 mL cell suspension. Cells were counted from this pooled suspension, after which the volume corresponding to 5 × 10<sup>7</sup> cells was transferred into each individual electroporation reaction. Following addition of DNA, each electroporation reaction was adjusted to a final volume of 500 µL with complete cytomix buffer. The remaining cell suspension was retained for additional transfections or control reactions.

      We agree that the original wording could be misinterpreted and have revised the Methods section to clarify the sequential handling steps more explicitly.

      Lines 225-229 revised to read: “The resulting approximately 1.5 mL pooled cell suspension was used for total viable cell counting using a hemacytometer.”

      “After counting, the volume corresponding to 5 x 10<sup>7</sup> cells was transferred to each electroporation reaction and combined with 25 µg of plasmid DNA. The total electroporation volume was adjusted to 500 µL with complete cytomix buffer.”

      (4) Figures 2 and 3 are too low-resolution for the font size used and for clearly viewing the microscopy images.

      We thank the reviewer for highlighting these readability issues. As noted in our responses above regarding Figures 2 and 3, the low-resolution appearance primarily resulted from manuscript compilation and PDF export artefacts affecting typography, image rendering, and panel clarity in the review version.

      To address this, Figures 2 and 3 have been completely reformatted and replaced with revised high-resolution versions featuring improved typography, panel labelling, contrast, accessibility, and image clarity for both on-screen viewing and print reproduction.

      Revised high-resolution versions of Figures 2 and 3 have been provided as described above. No additional manuscript changes were required beyond the figure revisions already outlined.

      (5) Figure 4 is confusing because the left and right panels appear inconsistent, with much higher concentrations required for growth inhibition in the culture-based assay than the resazurin assay indicated. The rationale for the resazurin assay should be explained, and the complete growth inhibition (CGI) concentration should be highlighted in the right panel.

      We thank the reviewer for highlighting this potential source of confusion. We agree that the distinction between the two assay endpoints was not sufficiently emphasised in the original figure presentation and legend.

      The apparent discrepancy arises because the two assays measure different biological endpoints under different assay conditions. The resazurin assay was used to estimate IC<sub>50</sub> values, corresponding to the concentration at which metabolic activity was reduced by approximately 50% under the assay conditions. In contrast, the small-culture assay was designed to determine complete growth inhibition (CGI), defined operationally as the concentration at which no detectable culture outgrowth occurred after incubation, using phenol red acidification as a culture-level readout.

      Because these assays measure partial metabolic inhibition versus complete suppression of detectable culture outgrowth, the corresponding concentration ranges are not expected to coincide directly. The higher concentrations observed in the right-hand panels therefore reflect the more stringent endpoint associated with complete growth inhibition rather than inconsistency between the assays.

      We agree that this distinction should have been explained more clearly in the original manuscript. We have therefore substantially revised the Figure 4 legend to clarify the rationale underlying both assays, explicitly distinguish IC<sub>50</sub> and CGI endpoints, and explain how the CGI values were used to guide subsequent antibiotic selection conditions for Blastocystis ST7-B transformants. The CGI transition range has also been made more visually explicit in the revised figure presentation.

      Figure 4 caption revised to: “Antibiotic potency and selection-window determination in Blastocystis ST7-B. Dose–response curves for puromycin, trimethoprim, and WR99210 were estimated from a resazurin-based viability assay (n = 3 independent replicates per drug per concentration). Points show mean ± SD, and the insets list the estimated IC50 values with R<sup>2</sup>-values > 0.75 for all fitted curves. The IC<sub>50</sub> estimates represent the drug concentrations that reduced resazurin-based metabolic activity by 50% under the assay conditions.”

      Right panels: “small-culture complete growth inhibition assay using 1 × 10<sup>7</sup> WT Blastocystis ST7-B cells per culture, assayed in triplicate across a wide range of concentrations. Cultures were incubated for 2 days, and outgrowth was assessed using phenol red acidification of the medium as a culture-level readout, with yellow indicating growth and red indicating no detectable growth. The yellow-to-red transition was used to estimate the concentration required for complete growth inhibition and to guide the subsequent antibiotic selection strategy for Blastocystis ST7-B transformants.”

      “IC<sub>50</sub> and CGI represent distinct assay endpoints: the former measures partial reduction in metabolic activity, whereas the latter identifies the concentration at which no detectable culture outgrowth occurs under the small-culture assay conditions.”

      (6) In Figure 3B, the unexpected UnaG fluorescence pattern could be due to protein sequestration because the protein is mildly toxic to the cell. This should be discussed in addition to the reasons already provided.

      We thank the reviewer for this thoughtful suggestion and agree that protein sequestration or reporter-associated cellular stress represent plausible alternative interpretations of the observed UnaG fluorescence pattern.

      We considered the possibility of UnaG-associated toxicity during interpretation of these data. However, under the conditions tested, we did not observe clear evidence of a substantial toxic effect: UnaG-expressing Blastocystis ST7-B cells could be recovered as stable lines, maintained under antibiotic selection, and propagated through continued culture. We therefore felt that direct attribution of the observed fluorescence pattern to reporter toxicity would currently remain speculative.

      At present, we consider the biochemical properties of the UnaG system itself to provide a more parsimonious explanation for the observed localisation pattern. In particular, unconjugated bilirubin is highly hydrophobic and would be expected to partition preferentially into lipid-rich cellular environments. This interpretation is consistent with the lipid-rich peripheral and intracellular structures previously reported in Blastocystis ST7-B (Liao et al., 2023).

      We have therefore revised the Discussion to acknowledge that the observed UnaG fluorescence pattern may reflect a combination of reporter-specific biochemical behaviour, bilirubin partitioning, local intracellular environment, or possible sequestration phenomena. At the same time, we avoid assigning toxicity as a demonstrated mechanism in the absence of direct measurements of cell fitness, reporter abundance, or bilirubin distribution. Such experiments would be required to evaluate this possibility rigorously.

      Lines 642-648: “Consistent with this, lipid-rich peripheral and intracellular structures have been reported in Blastocystis ST7-B, potentially providing favourable microenvironments for BR partitioning and contributing to the punctate UnaG fluorescence pattern (Liao et al., 2023). An alternative possibility is that the observed signal pattern reflects reporter sequestration or reporter-associated cellular stress. However, because UnaG-expressing lines were recovered, maintained under selection, and propagated through continued culture, toxicity remains a possible but untested explanation rather than a demonstrated mechanism.”

      Minor Comments

      Figure 2: Parts B and C should also show individual datapoints for better reader assessment.

      We agree that inclusion of individual data points improves transparency and interpretability of the underlying data distributions.

      Individual data points have now been overlaid on the boxplots in Figures 2B and 2C.

      Figure 3A: Separate channels (fluorescence, bright-field, merge) should be shown rather than only the merge. The current overlay is difficult to interpret, especially for colour-blind readers.

      We appreciate the reviewer’s concern regarding accessibility and interpretability. We considered separating the fluorescence, bright-field, and merged channels for Figure 3A. However, this panel was intended primarily as an overview demonstrating reporter detectability within the bicistronic construct context, while the detailed fluorescence distribution is explored more extensively in the subsequent UnaG panels. We therefore retained the merged presentation for Figure 3A. Importantly, the image is not dependent on red–green discrimination, as it combines a greyscale bright-field background with a high-contrast green/cyan fluorescence signal that remains distinguishable through brightness and contrast differences. In addition, colour-blind-friendly lookup tables (LUTs) were used throughout the revised figure set.

      To further improve accessibility, the original red annotation arrow has been replaced with a colour-blind-friendly annotation colour.

      Briefly define system components (P2A, UnaG, smURFP, SNAP-tag) and add an abbreviation list.

      We agree that brief contextual definitions improve accessibility for readers less familiar with these reporter systems. Rather than adding a separate abbreviation list, we have added short explanatory descriptions at the points where these components are first introduced in the manuscript.

      Lines 394–396: “The P2A peptide is expected to promote ribosomal skipping during translation, allowing two separate polypeptides to be produced from a single open reading frame.” Line 515–516: “UnaG, a bilirubin-binding fluorescent protein originally isolated from the muscle of the Japanese eel (Kumagai et al., 2013)…” Line 527: “smURFP (small ultra-red fluorescent protein)…”

      Abstract: “among the most prevalent microbial eukaryote” should be “eukaryotes”.

      Corrected in revised manuscript.

      Conclusion (2nd sentence): unclear what “endogenous regulatory part discovery” means.

      We agree that this phrase required clarification. The intended meaning was the identification and benchmarking of native Blastocystis ST7-B promoter and terminator elements for construct design and toolkit development. We have clarified this directly in the revised Conclusion section.

      Lines 682–683 revised to: “By bringing endogenous regulatory part discovery, namely the identification of native promoter and terminator elements, …”

      Author contributions: “critical advise” should be “advice”.

      Corrected in revised manuscript.

      Again, we thank the reviewers for their careful evaluation, constructive criticism, and thoughtful feedback on the manuscript. The review process has substantially strengthened the manuscript by helping us clarify the distinction between what is directly demonstrated experimentally and what remains mechanistically unresolved.

      The central methodological conclusions of the study remain unchanged: the toolkit enables selectable transgene expression, recovery of colony-derived lines, and propagation of reporter-positive transgenic Blastocystis ST7-B lines, extending genetic accessibility in this organism substantially beyond the previous transient transfection framework.

      At the same time, the revised manuscript now more explicitly acknowledges important unresolved mechanistic questions, including vector topology, P2A-mediated protein separation efficiency, and persistence in the absence of selection. These are now discussed transparently together with the future experimental approaches that will be required to address them directly.

      We believe the revised manuscript now presents a clearer, more rigorous, and more accessible description of a practical genetic toolkit for Blastocystis ST7-B and hope that the revisions and clarifications satisfactorily address the reviewers’ concerns.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Vasquez-Correa and colleagues describes the expression pattern of the ocelli (simple eye) gene regulatory network in ants. They correlate the expression pattern of these genes with the presence and absence of ocelli in different classes and species of ants. The presence of ocelli is a polyphenic trait in ants - understanding the molecular and developmental underpinnings of polyphenic traits is of significant interest to evolutionary biologists, developmental biologists, and ecologists. The authors propose that the presence of the latent expression of the ocellar network in classes of ants that do not display ocelli in the adults may underlie the re-evolution of ocelli within the ant lineage.

      Strengths:

      The strengths of the manuscript are that it is well written, the images are of the highest quality, and the data support the conclusions of the authors.

      We thank Reviewer 1 for their positive comments.

      Weaknesses:

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants. A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are, in fact, responsible for ocelli development in the ant.

      The reproductive caste in ants is typically composed of both winged males and winged queens. We agree with Reviewer 1 that the queen caste, which develop 3 fully functional ocelli, is an important point of comparison in our study to the wingless minor workers and soldiers. Unfortunately, however, laboratory colonies rarely produce reproductive queens, and in the field, queen production in colonies of C. floridanus occurs within a narrow seasonal window, making the collection of queen larvae particularly challenging for developmental work. In contrast, the winged males, which also develop 3 functional ocelli like the queens for help during mating flights, can be readily generated in the lab throughout the year. Therefore, we use males as a proxy for characterizing ocelli development and GRN in queens and the winged reproductive caste as a whole. Given the deeply conserved gene regulatory networks underlying this trait across insects, we believe this is a reasonable assumption.

      We also agree with Reviewer 1 that using RNAi to knock down genes in the ocelli GRN would improve the study. For completeness of the scientific record, we would like reviewers and readers to know that we actually did, in fact, invest significant effort trying to knock down otd-1 (ortholog of the Drosophila otd gene), which functions as key upstream regulator of ocellar development. In Drosophila, RNAi knockdown of otd disrupts the development of all three ocelli as well as fine morphological features on the anterior of the head. In C. floridanus, otd -1 is expressed in the head capsule and brain (see Author response image 1 in this response). Injection of dsRNA of otd-1into whole soldier-destined larvae, significantly reduced otd -1 expression in the brain relative to its control, while in the head capsule, otd -1 expression remained largely unchanged relative to its control (see Author response image 1 in this response). This indicates that in the same individual, the injected otd -1 dsRNA was able to penetrate and significantly reduce otd -1 expression in the brain, but, was unable to penetrate the head capsule, where otd -1 expression remained largely unchanged. No ocellar phenotypes could be observed in pupae or adults. Therefore, for technical (not biological) reasons, we were unable to knockdown genes in the ocelli GRN in the head capsule. We hope to solve this technical problem in the coming years to add a mechanistic explanation for the latent expression and maintenance of the ocelli GRN in workers that completely lack ocelli as adults.

      Reviewer #2 (Public review):

      Summary:

      The manuscript titled "Latent gene network expression underlies partial re-evolution of a polyphenic trait in the worker caste of ants" by Vasquez-Correa et al. aimed to study genetic mechanisms underlying developmental plasticity, especially binary polyphenism in queen vs worker ant castes. This is an interesting question regarding the extent to which phenotypic traits were altered, lost or regained, and how molecular pathways (upstream vs. downstream) can facilitate this process.

      In ants, reproductive castes (queens and males) develop wings as well as 3 ocelli for mating flights and other activities, while worker castes are wingless, and in some species, they have either no or a reduced number of ocelli. The phylogenetic analysis showed that in the Camponotini ant clade, the one-ocellus phenotype revolved in three species independently. The authors analyzed the conserved developmental pathways between Drosophila (well-established) and ants using HCR (a high-quality in situ hybridization technique). They found that although upstream genes for the development of ocelli (otd and hh) showed similar expression between castes, downstream genes (toy, eya, and so) had reduced or no expression in workers of C. floridanus, and this differential expression may lead to partial or complete loss of ocelli. Consistently, workers develop rudimentary tissues, suggesting that they initiate the ocellus developmental process but somehow stop it before adulthood.

      Strengths:

      Evo-devo approaches to reveal conserved molecular pathways of ocellus development. High-quality HCR provided convincing evidence of the expression of key genes in ocelli, eyes and antenna throughout larval development.

      Using HCR, the authors showed differential expression of downstream genes in males vs. soldiers vs. minor workers of C. floridanus, which might explain phenotypic differences between castes.

      We thank Reviewer 2 for their positive comments.

      Weaknesses:

      Although the molecular pathway is conserved, the mechanism underlying the lack of ocelli in workers remains unclear. In C. floridanus, it could be explained by the evidence of no expression of certain developmental genes, but in other species, e.g. Polyrachis rastellata, is their expression intact, or reduced? There is no control male.

      In addition, HCR in species with partial re-evolution (if their genomes have been sequenced) would be useful to understand the mechanism. For example, there might be differential spatial expression between medial and lateral ocelli.

      We agree with Reviewer 3 that investigating the mechanisms underlying the lack of specific ocelli in these and other species is the next step for this research. Here, our main focus was instead on trying to explain the mechanisms underlying partial reversion of ocelli through the persistence of ocelli GRN expression in adult workers lacking ocelli. We therefore focused on the latent expression of the ocelli GRN in Polyrachis rastellata, a species that completely lack ocelli in adult workers, and how it may have facilitated the partial reversion of a single ocellus in its congener Polyrachis bihamata. Therefore, although we did not reveal specific interruption points in the ocelli GRN in Polyrachis rastellata, our results showing that this species expresses three genes of the ocelli GRN, offers sufficient evidence that this network is conserved and likely facilitated the partial reversion to a single ocellus in P. bihamata.

      We also agree with Reviewer 3 regarding the male control in P. rastellata and obtaining the species in our study that have undergone partial re-evolution. Unfortunately, these ants occur in Southeast Asia and are very difficult to collect. For males in P. rastellata, our colony died before we could try to induce male development. However, given the deep conservation of the network in the males of a genus within the same subfamily (Camponotini), we feel it is reasonable to assume that the network would also be conserved in the males of P. rastellata, especially since the genes we sampled are conserved in workers that do not develop ocelli as adults. As am sure the Reviewer may know that this is a continual challenge of working with emerging models in evodevo.

      Reviewer #3 (Public review):

      Summary:

      This paper examines the loss and re-evolution of specific organs during the evolution of ants. The authors show that these organs, the ocelli, disappear and are re-evolved in different ant species and in different ant castes within these species. The authors show that this is linked to to a conserved GRN discovered in Drosophila, that appears to underlie the development of the ocelli, and demonstrate that this GRN appears to remain active in the developing heads of ants that have no ocelli- implying that it is the evolutionary latency of this GRN that allows loss and subsequent evolution.

      Strengths:

      This manuscript has outstanding imaging of a very difficult developing organ, and the key data, fluorescence in situ hybridisation, is done well and clearly shows what the authors wish to demonstrate. The methods are well described and underpin the whole work.

      The authors convincing demonstatrate that gene expression patterns imply the conservation of the ocellus gene regulatory network from Drosophila to ants. They further show that this network is present even in ants that don't produce an adult ocellus, but do show that in those species, loss of a developing nascent ocellus (which they identify) occurs at the same time as an interruption in the expression of the key genes in the GRN. All of this data is beautifully presented and explained.

      We thank Reviewer 3 for their positive comments.

      Weaknesses:

      There is one key weakness in that there are no functional students that indicate that the GRN actually does make the ocellus, though the expression patterns are convincing. This applies to loss of the ocellus as well. It would be nice to see that transient loss of the ocelli GRN might lead to loss of ocelli in ant species that have them. These are very difficult things to achieve, as the key genes have earlier developmental roles, such that CRISPR knockouts would not be interpretable, and transient RNAi in the head capsules of developing pupal ants would be challenging.

      We agree with Reviewer 3 that functional experiments in species where workers both have ocelli present and absent is a key next step in this research. Please see our response to Reviewer 1 on our failed attempts to achieve this. We are therefore grateful to Reviewer 3 for acknowledging the challenges in trying to establish RNAi and CRISPR in the head capsules of developing workers in these ants. Also, please see our response to Reviewer 2 on the difficulty of finding and collecting these ants, which occur mainly in Southeast Asia.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      One improvement that could be made is to include imaginal discs of the queen ants as well as scanning electron images of the ocelli of the queen ant to match the pupal stage images of the worker and soldier ants.

      A second improvement is to attempt a gene knockdown using RNAi or similar methods to ensure that the genes that are being studied are in fact responsible for ocelli development in the ant.

      Please see our response to Reviewer 1 above.

      Reviewer #2 (Recommendations for the authors):

      For the questions below, if there is no experimental evidence, consider addressing them in the Discussion.

      Do sizes of ocelli different between castes? For example, even workers have 1-3 ocelli, their sizes are smaller than those of males/queens, especially in workers with 1 ocellus. If so, might it be continuous (not binary) changes in downstream gene expression that control ocellus size, with no ocellus below threshold? Does this favor the hypothesis of threshold but not switch?

      We thank Reviewer 3 for highlighting an important point about the size and development of ocelli. Observations suggest that ocelli tend to be larger in queens and males than in workers and soldiers in species with ocelli. However, we lack quantitative data to test this conclusively. We now include a sentence on the Discussion stating that an important avenue of future work should investigate whether threshold or switch mechanisms influencing the presence/absence, as well as size, of ocelli between queens and workers.

      For the species whose workers have a single ocellus, are there variations, e.g. spanning from 0, 1 to 2? If 2, always one medial plus one of the two laterals? If always one, it would be a good control for staining to see up- vs down-regulation of downstream gene expression within the same individual.

      We agree with Reviewer 2 that this is a fascinating approach to our question. We have not observed natural wild-type variation in the number of developing ocelli in the same-sized individuals in the worker caste. However, in a distantly related leaf-cutting ant species (Atta cephalotes) belonging to different subfamily (the Myrmicinae) individuals with different head-to-body scaling within the same colony can vary in the number of ocelli. For example, soldiers of Atta cephalotes include individuals developing one, two, or three ocelli. These configurations can appear as only the median ocellus, only the two lateral ocelli, or even the median plus a single lateral ocellus. Interestingly, these correlations vary with changes in the size and head-to body scaling, suggesting that each ocellus can undergo different degrees of development, with one or more remaining vestigial or completely absent. On the other hand, workers in other species consistently develop a single ocellus, like in workers of Polyrachis bihamata, with no correlation to size or head-to-body scaling. These cases highlight how evolutionarily labile this trait is among workers of different ant species, which supports our proposal that the underlying gene regulatory network remains latent, thereby facilitating the emergence of novel trait combinations. We therefore agree on the importance of comparing the developmental mechanisms underlying these patterns temporally across larval stages and between individuals within a colony. We have now incorporated 2 sentences into the discussion, stating that this will be an important avenue for future work.

      Is there any function of a single ocellus in workers, or just a consequence of incomplete down-regulation of gene expression?

      Thank you again for highlighting these important points that help us to elaborate on the discussion of our study. The functional role of ocelli in species that develop these structures remains largely understudied. However, for some species particularly within the Formicinae clade the function of the three ocelli in workers has been investigated, revealing that they serve as a celestial compass that facilitates navigation. We reference these findings in our Introduction and Discussion to illustrate that the presence of three ocelli in workers can represent an adaptive trait. In contrast, the functional significance of a single ocellus or of partially developed ocelli remains an important question. This knowledge gap presents a promising avenue for future research to understand the adaptive value of reduced, partially suppressed ocellar development. We have now added a sentence in the discussion stating this.

      In previous studies, JH treatment can increase the number of ocelli in workers, consistent with its role in promoting reproductive development. In the ocellus developmental pathway, what causes the reduction of downstream gene expression in C. floridanus? Does JH directly regulate their expression?

      We thank Reviewer 2 for proposing yet another interesting question for future investigation, which we have added to the Discussion.

      The only current evidence available in C. floridanus is a recent study (MacMillan et al. 2025), in which minor workers and soldiers were treated with JH at different developmental stages. Unfortunately, no evidence of ocelli induction was observed in JH-treated individuals, suggesting that the mechanisms of ocelli development in C. floridanus might be highly canalized, especially in species that exhibit worker polymorphism (inter-individual variation in size and head-to-body scaling within the worker cate). However, more studies are required to understand why in Monomorium pharonis (no worker polymorphism) ocelli development can be readily induced by JH, while in another C. floridanus (with worker polymorphism) it appears quite difficult.

      "In D. melanogaster, the head develops from the eye-antenna disc" This statement is not correct. The brain does not belong to the eye-antennal disc.

      We thank Reviewer 2 for catching the misspelling. We have changed the name to eye-antenna disc in the sentence.

      Reviewer #3 (Recommendations for the authors):

      It is hard to see the developing ocelli in Figure 7 - could the authors increase the contrast to make them more visible?

      We have made the suggested changes to Figure 7 in the main article, and it has indeed improved the figure.

      Author response image 1.

      RNAi knockdowns in developing soldiers of Camponotus floridanus show a reduction of otd -1 expression in the brain, but no effect on otd -1 expression in the eye-antenna disc. A. HCR revealing otd -1 expression in the brain B. qPCR of otd -1 expression after RNAi knockdown shows significantly reduced otd -1 expression in the brain, C. HCR revealing otd -1 expression in the eye-antenna disc D. qPCR of otd -1 expression after RNAi knockdown shows no significant affect on otd -1 expression in the eye-antenna disc.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      Cisplatin, a platinum-based chemotherapeutic agent, induces intra- and interstrand crosslinks, thereby blocking DNA replication and transcription and triggering apoptosis. The authors aim to demonstrate that DNA polymerase κ (Polκ), traditionally seen as a translesion synthesis (TLS) polymerase, able to synthesize DNA through DNA lesions, plays a non-catalytic, structural role in stabilizing replication forks and protecting cells from cisplatin-induced cytotoxicity. A key finding of this work is the identification of two novel molecular axes: PCNA-Polκ-Polδ, which facilitates efficient DNA replication; PCNA-Polκ-USP18, which stabilizes DNA damage response proteins. These findings provide actionable therapeutic targets for overcoming head and neck squamous cell carcinoma chemoresistance, a cancer with rising incidence and limited treatment options.

      Strengths:

      The study relies on a robust experimental design, including Polk allegedly CRISPR-Cas9 knockout, siRNA knockdown, and rescue experiments with wild-type, catalytically dead, and PCNA interaction-deficient Polκ variants, supporting a non-catalytic role of Polκ. The work also reports a strong implication of Polk in cisplatin resistance, the identification of USP18 as a possible Polk partner and the consequences of Polk depletion on post-translational stabilisation of DNA damage response proteins.

      Thank you so much for appreciating our efforts to demonstrate role of Polκ mediated axes in cisplatin resistance in head and neck cancer cells.

      Weaknesses:

      The findings reported in this manuscript cannot be generalized to all cisplatin resistance mechanisms, as cells may develop multiple adaptive strategies to survive chemotherapy. Polκ's role varies across cancer types. For example, it is downregulated in stomach and colorectal cancers but upregulated in HNSCC, lung, and ovarian cancers. Thus, its use as a biomarker or drug target may be context-dependent.

      We completely agree with you, and the presented data only support Polκ's role in HNSCC as demonstrated in both acute cisplatin exposure as well as the cisplatin-resistant HNSCC models. Other cell and cancer types may adopt different strategies for cisplatin resistance.

      Acute cisplatin exposure is sufficient to trigger Polκ upregulation to levels similar to those in resistant cells. However, it remains unclear how long this upregulation persists and to what extent it contributes to survival. Further, the sensitivity of cisplatin-naïve H357 or SCC9 cells (H357-S and SCC9-S) to Polκ knockdown has not been addressed. This is a critical question, as acute cisplatin exposure induces Polκ expression to levels similar to those in resistant cells. This could argue against a direct role for Polκ in mediating resistance and instead suggest indirect mechanisms (like Polκ-dependent mutations during adaptation).

      Since H357-S and SCC9-S cells are highly sensitive to cisplatin, knocking down of Polκ unlikely will alter the phenotype, as other TLS DNA polymerases like Polκ and Polκ play critical role in such lesion bypass. Since no other DNA polymerase was upregulated in these cells upon cisplatin exposure and in the cisplatin-resistant cells, it was intriguing to demonstrate a direct role of Polκ in chemoresistance and that has been proven in this study. Since the catalytic activity of Polκ is not required to induce chemoresistant in these cells, we strongly believe that Polκ-dependent mutagenesis play minimal or no role in adapting cells to tolerate cisplatin. Nevertheless, we will knock down Polκ in these cells and determine cisplatin sensitivity

      The experimental design and results aimed at demonstrating the existence of a PCNA-Polκ-USP18 axis (Figure 9A) do not fully support the conclusion that these proteins form a stable complex. This set of experiments also lacks essential controls, such as the immunoprecipitated bait and the amount of immunoglobulins precipitated in all conditions. This also applies to the colocalization experiments in cells shown in Figure 9B. Images are poor and lack quantification. Further, Polk is seen mainly cytoplasmic in the upper panel, while it is nuclear in the lower panel. Discrepancies in Polk subcellular localization are also evident in the Supplementary data.

      We appreciate the Reviewer's critical and insightful comment. In our view, the interaction between Polκ and USP18 is very specific as USP2 and IgG alone do not pull down Polκ. Similarly, we also show that both Polκ and USP18 interact with PCNA. We agree with the reviewer that the existence of a stable complex of PCNA-Polκ-USP18 has not been fully demonstrated in the current version. We will perform additional experiments to strengthen our finding: a) Co-IP experiments with Polκ PIP mutants (wild-type vs. mutant) should be performed to determine whether USP18 loses its ability to bind PCNA in the absence of Polκ-PCNA interaction. b) Mapping the domain in Polκ that is involved in USP18 binding and their Co-IP experiment. Additionally, high resolution co-localisation images including quantified data will be provided.

      USP18 is known to deubiquitinate ISG15-modified proteins (not just ubiquitin). The study does not rule out ISGylation as a contributing mechanism.

      We find the point raised by the reviewer is very intriguing, however, as it will require a significant amount of time and effort to demonstrate ISGylation of DDR proteins and deISGylation by UPS18, and the insight that we may gain is unlikely to add to the central theme of this paper, we will expand this in our subsequent related study. Thank you for the suggestion.

      The experimental design involving analysis of DNA synthesis dynamics at a single-molecule level is not appropriate. Over interpretation of the data in several parts of the manuscript and lack of rigor in performing the experiments. Inappropriate consideration and absence of discussion of previously published literature directly related to the subject studied in this manuscript. Discrepancy with a previous report regarding the role of Polκ in Chk1 phosphorylation (Tonzi et al., eLife 2018). Synergic effect of T2AA inhibitor and Cisplatin have been already described in « naive » cancer cells (Inoue et al, 2014).

      Thank you very much for the suggestions. We will take care of the portions and modify as suggested. The necessary reference will be added as appropriate.

      Another critical point is that the proliferation rate of Polk-depleted cells is slower than that of wild-type cells. Hence, the colony formation assay shown in Figure 2B can be misleading, since the observed differences can be interpreted only as a proliferation problem.

      Thank you for pointing this out and we will modify the portion for better clarity.

      Reviewer #2 (Public review):

      Summary:

      Building on earlier studies, the authors report a role for pol kappa in mediated cisplatin resistance. Their data on dispensability of pol kappa catalytic activity for cisplatin resistance is consistent with previous reports. They further demonstrate that the PIP box of pol kappa is critical for cisplatin response. Based on these observations, the study concludes that targeting pol kappa and PCNA interaction can be a viable approach to overcome cisplatin resistance.

      Strengths:

      Indications that interaction between Pol kappa PIP box and PCNA can be targeted to overcome cisplatin resistance.

      Thank you for appreciating our finding that the PIP box of Polκ is critical for cisplatin response

      Weaknesses:

      (1) The study has used a model of cisplatin resistance and found that the phenotype is specifically reliant on upregulation of Pol kappa. They also observe that in this model of cisplatin resistance, there is rapid degradation of multiple repair proteins, including ATM, ATR, HR and NHEJ proteins upon knocking out Pol kappa. However, it is unclear how the resistant model was derived. Also, since the data and almost all experiments in this manuscript were performed with a single model of cisplatin resistance, the conclusions should be taken with caution.

      We are extremely sorry for the lack of clarity. Please note that two cisplatin-resistant models (H357 and SSC9) have been used and the results were very consistent in both cells. Fig. 1C clearly demonstrates about the generation of these resistant models and the original reference has been already cited.

      (2) There are also inconsistencies in findings. Increased G2 arrest and no change in origin firing are being observed despite a significant reduction in Chk1 protein levels.

      Thank you for pointing this out. In our view, the increased G2 arrest is due to more fork stalling or collapsed than the new origin firing. Also, in our assay we observed less than 10% of new origin fired DNA fibres, and that could be the reason of no significant change in new origin firing among various cells.

      Reviewer #3 (Public review):

      This manuscript investigates the role of PolK in cisplatin repair. While in general it is considered that polK is not involved in the repair of cisplatin-induced DNA damage, the authors show that in a very specific scenario, namely cisplatin-resistant head and neck cancer cells, loss of PolK causes cisplatin sensitization, implying a role in cisplatin repair by polK in these cells. It is also implied that these cells acquire cisplatin resistance by overexpressing polK, but this is not really investigated. The authors then go on to show that DNA replication in the presence of cisplatin is affected by the loss of polK in these cells and also identify USP18 as a potential polK interactor in these cells with a similar phenotype. They claim that polK and USP18 form a pathway that allows cisplatin tolerance in these cisplatin-resistant head and neck cancer cells. The findings are interesting and useful to the field; however, the manuscript, in its current form, has several issues. Most importantly, the mechanism of USP18 has not been investigated. In addition, the manuscript does not flow fluidly, and instead, various experiments are put together without a clear logic. Some of the claims are not substantiated by the data shown.

      Thank you very much for finding our study interesting and the pending concerns will be addressed as suggested.

      (1) The experiments in Figure 1 using a few cell lines from various types of cancers are not enough to conclude that polK expression is specifically induced by cisplatin in some types of cancers but not others. Since the focus of this study is head and neck cancer, the authors should show the expression of PolK after cisplatin treatment in more head and neck cancer cell lines, and not just the two investigated.

      In this study, we have explored eight different cell types (breast, brain, liver, head and neck, pancreatic, prostrate, lungs, and kidney) to check the expression of Polκ upon cisplatin exposure, and HNSCC cells only showed Polκ up-regulation. Therefore, we went ahead for further demonstration of the role of Polκ in cisplatin resistance in OSCC using four different cell models (H357-S, H357-R, SSC9-S, and SSC9-R). By adding more cell lines to study will unlikely change the central theme of the paper. Yes, by acquiring and analysing clinical samples from the cisplatin responder and non-responders would have strengthen our finding.

      (2) It is unclear to me why the authors include H357-S in their experiments. If the idea is that these cells acquire resistance because they overexpress polK, then the authors should investigate this by exogenously overexpressing PolK in H357-S cells and test if these cells are cisplatin resistant.

      It’s an interesting point and we will check whether overexpression of Polκ in H357-S cells could induce resistance to cisplatin and alters IC<sub>50</sub>. Thank you for the suggestion.

      (3) In addition, the authors should create the polK knockout in H357-S cells as well and include it as a control in their experiments.

      We appreciate your suggestion. As suggested by Reviewer #1 also, we will check the phenotype of Polκ knockdown H357-S cells.

      (4) Page 6, line 28: the comet assay does not measure DNA degradation, but rather DNA breaks.

      Thank you for the suggestion, we will modify the text accordingly.

      (5) Figure 4B: How does the overexpression of PolK mutants compare to endogenous PolK expression? It is important to assess if this expression is similar or of much higher magnitude.

      Please note that GFP-Polκ has been overexpressed in H357 Polκ knockout cells to nullify the effect of endogenous Polκ, otherwise we will not be able to test the role of various Polκ mutants.

      (6) Page 9, line 22: "For such a function, the catalytic domain of PolK becomes dispensable, whereas its interaction with PCNA is sufficient to drive efficient replication". I do not understand what data the authors used to make this claim. The interaction and colocalization studies should be performed with the PIP mutant. Similarly, this mutant should be used in the HU DNA fiber assays.

      We are extremely sorry for the lack of clarity. The inference has been derived from two sets of experiments as shown in Fig. 4C and Fig. 4D (and is with HU).

      (7) It is unclear how USP18 acts. What are its substrates? Chk1/2, BRCA1, BRCA2? This needs to be investigated. The impact of PolK on this activity needs to be assessed as well (is PolK needed for USP18-mediated de-ubiquitination of these DSBR proteins?). As it stands, the manuscript does not address the mechanism of USP18 in DNA repair, which is billed as the main finding of the paper.

      It has already been demonstrated in Fig. 9C where by knocking down USP18, the DDR proteins like Chk1, Chk2, CtIP, and Artemis can be recovered for ubiquitin-mediated proteasomal degradation. The same results are also obtained when its interacting partner Polκ is deleted. In our view, the presented results have sufficiently demonstrated the role of Polκ-Usp18 in the repair of cisplatin adducts through DDR proteins.

      (8) Do PolK and USP18 interact directly? Experiments using recombinant proteins would be useful to address this.

      We appreciate your suggestion. Since the Usp18 protein is not readily available, we will not be able to show; however, we believe the interaction is direct, and we will be able to map the binding site in Polκ.

    1. Author response:

      eLife Assessment

      This is a potentially important study comparing LTP mechanisms between primates and rodents. The experimental methods have some possible confounds, and the power (replicates) and design of the statistical methods could be strengthened, hence the support for the central claims of species differences is currently incomplete.

      We thank the Editor and the Reviewers for taking the time to carefully review our manuscript and for providing constructive comments and suggestions, as well as the opportunity to revise our work.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This is an important paper examining LTP induced by theta-burst stimulation in hippocampal slices from macaques and rats. While both species show theta-burst-late-LTP, only the non-human primate theta-burst-late-LTP showed synaptic tagging and capture that converts early-LTP into late-LTP in an independent synaptic pathway.

      Strengths:

      Synaptic tagging is a fundamental feature of repeated 100 Hz-tetanus-induced LTP, whereas theta-burst induction is arguably more physiologically relevant. Thus, synaptic tagging during theta-burst may differ in the two species, a distinction that may prove important in the mechanisms underlying the cognitive differences between the species.

      Weaknesses:

      Bursts repeated at the frequency (~5 Hz) of the endogenous theta rhythm induce strong LTP, primarily because this frequency disables feed-forward inhibition and allows sufficient postsynaptic depolarization to activate voltage-sensitive NMDA receptors. Therefore, the species differences may be due to differences in inhibition, rather than in molecular mechanisms of maintenance. One way to assess the relative strengths of this early induction mechanism in rats and macaques is to examine the "depolarization envelope" during the sequential bursts, which may be determined from the recordings already obtained. (Larson and Munkácsy, Theta-burst LTP, Brain Res 2015 Sep 24:1621:38-50. doi: 10.1016/j.brainres.2014.10.034)

      Another issue is that the PKMzeta-antisense oligodeoxynucleotides block the synthesis of the kinase. However, Mei F, Nagappan G, Ke Y, Sacktor TC, Lu B (2011), BDNF Facilitates L-LTP Maintenance in the Absence of Protein Synthesis through PKMzeta. PLoS ONE 6(6):e21568, provided evidence that BDNF and theta-burst stimulation can act to increase PKMzeta by a protein synthesis-independent mechanism, presumably through decreased degradation. Therefore, the absence of an effect of the PKMzeta-antisense does not exclude the possibility that persistently increased PKMzeta is the mechanism of theta-burst-late-LTP maintenance in mice or macaques. This issue is worth discussing.

      We sincerely thank the reviewer for the positive evaluation of our study and for highlighting the significance of examining synaptic tagging and capture following theta-burst stimulation (TBS) in rodents and non-human primates.

      We agree that TBS is a physiologically relevant induction paradigm and that differences in inhibitory circuit dynamics may also contribute to the species-specific effects observed in our study. As highlighted by Larson and Munkácsy (2015), repeated bursts delivered at theta frequency (~5 Hz) can transiently suppress feed-forward inhibition through GABAB receptor-mediated mechanisms, thereby enhancing postsynaptic depolarization and facilitating NMDA receptor activation. We therefore agree that species differences in inhibitory regulation and burst-evoked depolarization may contribute to the distinct expression of synaptic tagging and capture observed between rats and non-human primates.

      We further agree that analysis of the “depolarization envelope” during sequential bursts may provide additional insight into the relative strengths of early induction mechanisms. We will therefore perform these analyses using the existing recordings and compare the depolarization envelope between rodents and NHPs in the revised manuscript. Following the reviewer’s suggestion, we will expand the Discussion section to acknowledge the potential contribution of inhibitory circuit dynamics and depolarization envelope differences during sequential bursts.

      Importantly, however, we believe that differences in downstream molecular maintenance mechanisms also contribute to these species-specific effects. In support of this, our molecular analyses revealed enhanced recruitment of plasticity-related proteins and transcriptional pathways in NHP hippocampus following TBS, including increased expression of BDNF and PKCζ. These findings suggest that both induction-related network properties and downstream molecular stabilization mechanisms may collectively contribute to the enhanced associative plasticity observed in NHPs.

      We also thank the reviewer for the important point regarding PKMζ antisense experiments and the study by Mei et al. (2011). We agree that the absence of an effect of PKMζ antisense oligodeoxynucleotides does not necessarily exclude a role for persistently elevated PKMζ in the maintenance of theta-burst late-LTP. As demonstrated by Mei et al., BDNF together with theta-burst stimulation can maintain late-LTP in the absence of protein synthesis, potentially through stabilization of PKMζ protein levels by reducing degradation rather than through de novo synthesis. However, these findings are not directly comparable to our study, since our experiments involved theta-burst stimulation alone without exogenous BDNF application. Interestingly, our results suggest species-specific differences in the interaction between BDNF and PKMζ signaling pathways. In rats, TrkB/Fc-mediated blockade of BDNF impaired TBS-LTP maintenance, whereas PKMζ inhibition alone had no significant effect. In contrast, in NHP hippocampal slices, inhibition of either BDNF signaling or PKMζ alone failed to abolish late-LTP, whereas simultaneous inhibition of both pathways disrupted LTP maintenance.

      These findings suggest that endogenous BDNF signaling and PKMζ may operate through partially redundant or compensatory mechanisms, particularly in the primate hippocampus. Therefore, although our findings indicate that de novo PKMζ synthesis may not be strictly required under the present experimental conditions, we cannot fully exclude the possibility that protein synthesis-independent stabilization or maintenance of PKMζ contributes to theta-burst late-LTP maintenance in rodents or NHPs. We will now clarify this point in the revised Discussion section.

      Reviewer #2 (Public review):

      Summary:

      This study compares theta-burst stimulation (TBS)-induced synaptic plasticity in hippocampal CA1 slices from rats and non-human primates (Macaca fascicularis). The authors report that while TBS induces persistent LTP in both species, only primate hippocampal slices exhibit synaptic tagging and capture (STC) under these conditions. They further show increased BDNF and PKMζ expression following TBS in primates and propose that a redundant BDNF/PKMζ signaling architecture supports persistent plasticity in primates, whereas rodent TBS-LTP depends primarily on BDNF. The work aims to identify species-specific specializations in associative plasticity with implications for translational neuroscience.

      Strengths:

      The topic is potentially important because direct comparisons of hippocampal plasticity mechanisms between rodents and primates are rare.

      Weaknesses:

      (1) Limited biological replication in the primate experiments

      The manuscript's strongest claims rely on data obtained from 36 slices from 7 monkeys, qPCR analyses with n=3 biological replicates, and Western blot analyses with n=3 biological replicates. The effective sample size for species-level conclusions is therefore not large. The manuscript frequently treats slices as independent observations while drawing conclusions about species differences. This is particularly problematic for electrophysiological experiments because multiple slices appear to originate from the same animals. The statistical unit should be the animal, not the slice, unless nested analyses are performed.

      The authors should (1) report the number of animals contributing to each experiment, (2) provide animal-level analyses, (3) use mixed-effects or hierarchical models where appropriate, and (4) clarify whether multiple slices from the same monkey contributed to the same experimental condition. Without these analyses, the evidence for species-specific mechanisms remains weaker than presented.

      We thank the reviewer for this important and thoughtful comment regarding statistical interpretation and biological replication. We agree that, particularly for electrophysiological experiments where multiple slices may originate from the same animal, the effective sample size for species-level conclusions should be considered at the animal level rather than solely at the slice level.

      In the revised manuscript, we will clearly indicate the number of biological replicates (animals) together with the number of slices contributing to each electrophysiological experiment, as well as the biological replicates used for qPCR and Western blot analyses. We will also clarify whether multiple slices from the same NHP/rat contributed to the same experimental condition. These details will be incorporated into the figures and figure legends wherever appropriate.

      In addition, we will perform animal-level analyses by averaging slice responses within each animal prior to statistical comparison and, where appropriate, apply hierarchical or mixed-effects statistical models to account for the nested structure of slices within animals.

      We acknowledge that the number of non-human primates (NHPs) available for this study was inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with primate electrophysiology and tissue collection. Consequently, achieving sample sizes comparable to rodent studies is often not feasible in NHP research. Nevertheless, to further strengthen the biological robustness of the findings, we are currently in the process of obtaining additional NHP brain samples and plan to repeat key experiments in an additional 3-4 animals. We believe these revisions and additional experiments will substantially strengthen the statistical rigor and overall interpretation of the study.

      (2) The central STC conclusion requires stronger controls

      The most important result is that TBS supports STC in primates but not rats (Figures 1F-G). However, several alternative explanations are not excluded. For example, only a single interval (30 min) between TBS and WTET is examined. Classical STC studies characterize tag duration, PRP availability window, and temporal asymmetry. The current work does not determine whether primates exhibit longer tag persistence, increased PRP synthesis, altered capture efficiency, or merely a shifted temporal window. A temporal series (e.g., {plus minus}15, {plus minus}30, {plus minus}60, {plus minus}90 min) would substantially strengthen the mechanistic interpretation.

      We thank the reviewer for this insightful comment regarding the mechanistic interpretation of the STC findings. In the present study, we selected the 30 min interval based on well-established classical STC paradigms in rodents, where this interval reliably falls within the effective tagging and capture window. Using this experimentally validated interval allowed us to directly compare whether TBS is sufficient to support STC in primates versus rats under equivalent experimental conditions. Accordingly, the primary objective of this study was to determine whether TBS-induced STC varies across species, rather than to comprehensively define the temporal dynamics of the tagging window.

      We agree, however, that the current experiments do not distinguish whether the primate-specific effect reflects prolonged tag persistence, enhanced plasticity-related protein (PRP) synthesis, altered capture efficiency, or a shifted temporal window. Addressing these possibilities would indeed require systematic temporal interval analyses (e.g., ±15, ±30, ±60, and ±90 min), which represent important future directions. Such experiments are particularly challenging in non-human primates because the availability of primate tissue and experimental resources for large-scale electrophysiological studies remains limited and is currently beyond our experimental capacity due to substantial ethical, logistical, financial, and technical constraints.

      Nevertheless, we fully agree with the reviewer that these experiments are important for advancing the mechanistic interpretation of the findings. Similar temporal analyses have recently proven informative in our rodent studies (Chong YS, Ang SR, Sajikumar S. Commun Biol. 2025;8:553). Importantly, we are currently in the process of obtaining additional non-human primate samples and plan to extend the present work by examining an additional 60 min temporal interval to further characterize the temporal properties of synaptic tagging and capture in non-human primates.

      (3) Species differences may reflect tissue quality or preparation differences

      The manuscript compares 5-7 week-old rats with 5-7 year-old monkeys. These are very different developmental stages. Moreover, euthanasia methods, extraction procedures, and post-mortem handling are different. These factors can affect BDNF expression, protein synthesis, LTP magnitude, and transcriptional responses. The authors should discuss these caveats more explicitly.

      We thank the reviewer for raising this important and insightful point. We agree that differences in developmental stage between the experimental groups represent an important consideration when interpreting potential species-dependent effects. In the present study, rat experiments were performed in 5-7 week-old animals, whereas non-human primate (NHP) tissues were obtained from 5-7-year-old monkeys. This difference largely reflects the practical, ethical, and logistical constraints associated with NHP research and tissue availability. We acknowledge that these ages are not developmentally equivalent and that maturation state may influence BDNF signaling, protein synthesis capacity, synaptic plasticity thresholds, and transcriptional responses relevant to late-LTP and STC mechanisms.

      We also recognize that differences in euthanasia procedures, tissue extraction, slice preparation, and postmortem handling between rodent and primate tissues may influence tissue physiology and electrophysiological properties. Although extensive care was taken to optimize tissue viability and maintain stable recordings within each species, these variables cannot be completely excluded as contributing factors to the observed differences.

      Accordingly, we will revise the Discussion section to more explicitly acknowledge these limitations and clarify that our findings support potential species-dependent differences under the present experimental conditions, rather than definitive intrinsic species-specific mechanisms. Nevertheless, despite the inherent challenges associated with NHP electrophysiological studies, we believe that the present findings provide an important initial framework for understanding the translational relevance of synaptic tagging and capture mechanisms across species.

      (4) Statistical reporting is incomplete

      Many comparisons report exactly Wilcoxon p = 0.0313 and U-test p = 0.0022, across numerous experiments. This suggests very small sample sizes and discrete nonparametric distributions. The manuscript should report exact n values for each comparison, effect sizes, and confidence intervals.

      Second, many genes and proteins are tested. No correction for multiple testing is described. The authors should state whether corrections were applied, and if not, justify this choice.

      We thank the reviewer for this important comment regarding statistical reporting and interpretation. We agree that the repeated occurrence of identical exact p-values in several nonparametric analyses reflects the relatively small sample sizes and the discrete nature of the statistical distributions. This issue is particularly relevant for the NHP experiments, where biological replication is inherently limited because of the substantial ethical, logistical, financial, and technical challenges associated with obtaining and processing primate tissue.

      In the revised manuscript, we will provide exact n values for all comparisons, including the number of biological replicates (animals) and slices where applicable. We will also include additional statistical details, including effect sizes and confidence intervals where appropriate, to improve transparency and facilitate interpretation of the reported findings. Furthermore, we are currently in the process of obtaining additional NHP samples and will attempt to include more biological replicates in the revised version to further strengthen the robustness of the analyses.

      We also agree that the issue of multiple testing should be addressed more explicitly, particularly because multiple genes and proteins were examined. In the revised manuscript, we will clearly state the statistical correction methods applied for multiple comparisons where appropriate. For analyses in which corrections were not applied, we will provide justification, noting that several experiments were based on hypothesis-driven candidate targets rather than exploratory large-scale screening analyses. These statistical considerations will be clarified in the Methods and Results sections.

      (5) Interpretation and significance

      The study addresses an important and understudied question: whether associative synaptic plasticity mechanisms differ between rodents and primates. The finding that TBS can support STC in the primate hippocampus is potentially novel and impactful. However, the mechanistic evidence remains incomplete, the molecular analyses are underpowered, and several key controls are missing. At present, the data support the conclusion that under the specific experimental conditions tested, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than in rat slices.

      The stronger claims regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and redundant BDNF/PKMζ architecture require additional experimental support.

      We thank the reviewer for this thoughtful and balanced assessment of our work. We agree that the present data primarily support the conclusion that, under the specific experimental conditions examined, TBS-induced plasticity in primate hippocampal slices exhibits greater associative persistence than that observed in rat slices. We also agree that broader interpretations regarding evolutionary specialization, fundamentally distinct plasticity rules, altered STC thresholds, and potentially redundant BDNF/PKMζ-related mechanisms require additional mechanistic investigation and experimental validation.

      Accordingly, we will moderate these interpretations throughout the revised manuscript and clearly state that these conclusions remain preliminary. We will further emphasize that additional experiments, including increased biological replication, expanded temporal analyses, and further mechanistic investigations, will be necessary to more conclusively define the basis of the observed species-dependent differences. Within our current experimental capacity, we are actively working to obtain additional non-human primate samples and plan to incorporate additional biological replicates and key follow-up experiments in the revised version to further strengthen the robustness of the findings.

      At the same time, we believe the present study provides an important initial contribution to an understudied area by directly examining synaptic tagging and capture mechanisms in the primate hippocampus. Given the limited availability of non-human primate electrophysiological data in the field, these findings may offer a valuable framework for future studies investigating the translational and evolutionary relevance of associative synaptic plasticity mechanisms across species.

      Reviewer #3 (Public review):

      Summary:

      In this manuscript, the authors have undertaken an investigation of differences between two mammalian species, the brown rat and the crab-eating macaque, in the mechanisms supporting a well-established model of long-term Hebbian synaptic plasticity, Schaffer collateral to CA1 Long-term potentiation (LTP) in the hippocampus. LTP has been long-studied and deeply characterised due to its potential importance in modeling a strong candidate process for the central mechanism of learning and memory. LTP was first discovered in lagomorphs (rabbits), but has since been much more widely studied in rodents (mostly rats and mice), and there has been some complementary work revealing LTP in non-human primates and even in humans, revealing largely overlapping canonical mechanisms of induction, expression, and maintenance. More specifically, this study puts a particular focus on the fascinating associative features of this form of lasting synapse-specific modification, in which a synaptic input can be stimulated with a relatively weak induction protocol that will not produce lasting plasticity on its own, but can undergo lasting LTP if paired with stronger stimulation on a separate synaptic input to the same neuron. This associativity mechanism is particularly attractive within the Hebbian synaptic plasticity framework as it provides a candidate mechanism for associative forms of learning in which stimulus-stimulus, stimulus-reward, stimulus-punishment, or action-outcome associations are formed. A particularly attractive feature of this associative LTP is that there can also be a substantial time-lag between the strong stimulation of one pathway and the weaker stimulation of the other synaptic input, which only undergoes lasting LTP by hijacking the proteins synthesized as a result of strong stimulation elsewhere. This observation has led to the famous tagging and capture hypothesis as an explanation of how such synapse-specific change can be achieved on both stimulated inputs but not on other synaptic inputs, given the potential requirement for cell-wide protein synthesis. This theory, for which there is very strong experimental evidence, posits that a protein tag is left at synapses that have been stimulated with sufficient vigor in recent history, serving as a key mechanism to ensure that those weakly stimulated synapses will undergo change when a larger-scale LTP event occurs due to stronger stimulation elsewhere within a relevant time window. Again, this idea is attractive as it can explain how we might form associations between events that occur slightly separated in time. The manuscript goes on to show that an induction protocol that is particularly physiologically relevant, theta burst stimulation, produces this tag and capture associative effect in ex vivo slices of Macaque hippocampus, much more readily than in side-by-side ex vivo slices of rat hippocampus. Moreover, the manuscript delves into the importance of well-characterised LTP maintenance mechanisms, including PKMzeta and BDNF, which are key factors that ensure that altered synaptic change is maintained for long periods of time despite substantial molecular turnover in the neuron. The observation in this manuscript is that a degree of redundancy for these mechanisms exists in the primate species but not the rodent species, as both mechanisms need to be inhibited to return LTP to baseline in the Macaque, but only one needs to be inhibited to have that effect in the rat. A major emphasis of this study is that there may be a step-wise difference in associative learning mechanisms between rodents and primates that may contribute to their differing cognitive capacities, although I believe a lot more evidence would be required to reach that conclusion.

      Strengths:

      The strengths of this study are that it is technically very proficient and is from a laboratory that has a long history of seminal work on synaptic tagging and capture. The cross-species comparison, particularly involving non-human primates, is also very hard to achieve, and a major strength here is the side-by-side comparison of slices from rat and monkeys. Further strengths of the study are the use of a number of experimental strategies, including both observation and intervention, to demonstrate differential involvement of LTP maintenance mechanisms. A final major strength is conceptual, as it is undoubtedly useful not only to identify shared mechanisms of plasticity between commonly used model organisms and either humans or much more closely related species such as old world monkeys, but also to reveal differences that have the potential to contribute to differences in memory/cognition.

      Weaknesses:

      The findings of this study are a very useful building block for understanding how generalisable mechanisms of LTP are. However, arriving at really substantial conclusions from these findings is challenging, as there are a number of variables that are unaccounted for in this study that may explain the differences that have been observed between rats and monkeys. One example of a potential confound to these interpretations is that rats are nocturnal/crepuscular animals, and macaques are diurnal animals. Thus, to undertake a like-for-like comparison, it would be necessary for the rats to be on a reversed light-dark cycle to ensure that the wake cycle of the rat (dark) is being compared with the wake cycle of the monkey (light). It is possible that the authors have done this, but it is not mentioned in the methods section. The reason this is important is that there is a substantial body of work indicating that different mechanisms are at play in hippocampal LTP during wake and sleep. Transcripts and proteins related to synaptic function are dramatically differentially regulated during sleep-wake cycles, and phosphorylation states of key proteins involved in plasticity are also altered. Moreover, synaptic tagging and capture are specifically disrupted by sleep deprivation. Perhaps the authors have already considered this factor and appropriately reversed the light-dark cycle of their rat subjects, in which case a clarification in the manuscript would be useful. Nevertheless, I have used this as an example because there is a variety of potential confounds that may explain the difference between SC-CA1 TBS LTP in rats and monkeys, e.g., circadian rhythms, degree of enrichment, natural light vs indoor lighting, diet, degree of inbreeding, strain, etc. Thus, to make strong conclusions about the potential for differences in plasticity rules/mechanisms and how those may contribute to differences in cognition, I think it would be necessary to compare a wider variety of species, including a good representation of each order (e.g., nocturnal rats and diurnal squirrels, new and old world primates) and not just a single exemplar. I understand, of course, that this is really pushing the boundaries of practicality, but I see no other way to make a strong conclusion or to generalise to mechanisms or properties of plasticity in rodent’s vs primates. Thus, while I believe the manuscript presents really admirable work, I am not sure the findings are at all easy to interpret.

      We thank the reviewer for this thoughtful and insightful comment, as well as for the encouraging appreciation of our long-duration plasticity recordings and associative plasticity experiments, which are both technically demanding and time-intensive. We fully agree that interpretation of cross-species differences in synaptic plasticity requires careful consideration of multiple biological and environmental variables, including circadian state, enrichment conditions, strain differences, diet, lighting conditions, and species-specific behavioral ecology.

      Regarding the specific concern related to circadian phase and sleep-wake state, the reviewer raises an important point. Rats are nocturnal animals, whereas macaques are diurnal, and hippocampal plasticity mechanisms are known to be influenced by circadian rhythms and sleep-dependent regulation of synaptic proteins and signaling pathways. Previous studies have demonstrated modulation of LTP, synaptic tagging and capture and protein synthesis in rats across normal sleep-wake cycles. We therefore agree that these factors may influence plasticity outcomes and should be carefully considered in comparative studies.

      Studies have further shown that theta frequency is highly sensitive to sleep-related manipulations. Specifically, theta frequency decreases immediately after sleep, remains elevated during sleep deprivation, and rapidly declines following recovery sleep. In aged animals, these effects appear comparatively attenuated, suggesting reduced sleep-dependent modulation of theta dynamics with aging. Therefore, disruption of normal circadian or sleep-wake patterns may significantly alter theta activity and associated plasticity mechanisms within a species and may not accurately reflect physiological baseline states (Utku Kaya et al., 2026).

      In our experiments, recordings from rats and macaques were performed during their respective active phases under standardized laboratory housing conditions, and we will further clarify these details in the revised Methods section. Nevertheless, we acknowledge that circadian state and related physiological variables cannot be completely excluded as contributing factors to the observed differences between species.

      More broadly, we agree with the reviewer that the present study does not permit definitive conclusions regarding universal “rodent versus primate” rules of synaptic plasticity. Our intention was not to propose a generalized dichotomy between rodents and primates, but rather to report that, under the experimental conditions used here, SC-CA1 TBS-LTP and associated synaptic tagging mechanisms differed between rats and macaques. We agree that broader evolutionary or cognitive interpretations would require systematic comparative analyses across multiple species, including both nocturnal and diurnal rodents as well as diverse primate species. Such studies would provide a stronger framework for distinguishing conserved versus species-specific mechanisms of plasticity.

      At the same time, we believe the present findings remain important because they provide one of the first direct experimental comparisons of SC-CA1 TBS-LTP-associated plasticity mechanisms between rodents and non-human primates under controlled ex vivo conditions. Although the interpretation should be done cautiously, the observed differences raise the possibility that certain metaplastic or protein synthesis-dependent mechanisms may not be fully conserved across species. Accordingly, we will revise the Discussion section to better emphasize the exploratory and comparative nature of the study, while explicitly acknowledging the limitations and potential confounding factors highlighted by the reviewer.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      This article describes a very ambitious metascience project aimed at testing the reproducibility of a corpus of publications conducted in Brazil. The strength of the approach lies in its systematic, multicenter replication design. The authors focus on three commonly used experimental paradigms in biology: the MTT assay, RT-PCR, and the elevated plus maze.

      The effort is commendable and reveals a rather low rate of reproducibility, in line with findings from fields considered less reproducible in the life sciences, such as cancer biology.

      Strengths:

      The study is supported by a substantial dataset, incorporating multiple independent replication attempts and the use of stringent, well-defined protocols, which strengthens confidence in the overall conclusions.

      We thank the reviewer for the comments.

      Weaknesses:

      (1) Being neither an expert in metascience nor in statistics, I cannot fully judge the methodological aspects of the article or its extensive supplementary material. I will therefore focus my comments on readability. I found the manuscript difficult to digest. The authors should improve readability if they wish to reach a broad audience of experimental biologists. In particular, they should simplify the description of protocols and highlight the key findings more clearly, using accessible language. See specific points below

      We can try to simplify the description of protocols at specific points for example, by providing an overarching description of the study design in the beginning of the Methods, rather than citing our previous eLife paper (Amaral et al., 2019), as suggested below. The methods are indeed quite extensive, but the this may be inevitable in a large-scale project such as this and we note that Reviewer #2 thought that part of the supplementary material should be incorporated back in the main text, which is a suggestion in the opposite direction. It may thus be hard to strike a balance between readability and comprehensibility that can address both reviewers’ opinions.

      (2) The article appears to oscillate between:

      (i) a description of the approach and the inherent challenges of such a multicenter replication program

      (ii) an estimation of reproducibility.

      These could potentially form two separate articles: one aimed at a broad audience emphasizing key results, and another focused on methodological aspects for a more specific metascience audience. The Results section currently contains redundancies and is difficult to follow for non-experts in statistics. I also find it challenging to extract the main findings.

      There is a bit of redundancy between tables and text, but this was intentional to make both of them self-explanatory. We also think stating the results in the text can allow us to make each of the replication criteria clearer, a concern that was also mentioned by the reviewer.

      As for requiring particular expertise in statistics for understanding, we mostly disagree. The main results (Tables 1 and 2, Figure 2) are expressed as percentages, and the only statistical concepts needed for interpreting these results are understanding prediction and confidence intervals. For this, we could provide a bit more guidance on their interpretation in the Methods section. Beyond that, most of the secondary results (e.g. Figure 3 and Figure 4) involve linear correlations, which is about as simple as statistical analysis gets.

      Of the results presented in the main manuscript, only Table 3 contains anything beyond percentages and correlations. We do agree that the meaning of each ratio in this table could be more clearly described, but there are essentially no expert-level statistics involved in their calculations.

      Other than that, the main statistical issues are the ideal way to aggregate the results from different replications for which we use different strategies for robustness purposes. However, all of these results are already in the supplementary material, so we don’t feel they interfere to much with the readability of the main manuscript.

      A possible improvement would be to include an initial section clearly describing the protocol (replication of a single experiment, across several labs, for three types of assays), followed by a concise presentation of the main results regarding reproducibility in Brazilian science with subsections.

      This is indeed a good idea, and we plan to include an initial overarching description of the project in the Methods section of the revised manuscript.

      Methodological details could be moved either to a Supplementary Information or to a more specific article, while being summarized in the Discussion.

      Again, this is the opposite of what was suggested by Reviewer #2, so we would rather keep the Methods section more or less at its current level of detail.

      (3) This study evaluates the reproducibility of a single experiment from each article, taken out of its broader context. While this provides an estimate of reproducibility, it does not directly contribute to resolving uncertainties within a specific field. This may represent a limitation compared to other reproducibility projects that attempt to replicate multiple key claims within a given study (e.g., in cancer biology or Drosophila immunity). I found that a weakness is that it does play a role in cleaning a field of wrong statements.

      The reviewer is correct in his interpretation. Evaluating the main findings of articles or cleaning a field of wrong statements was never a goal of our study (and we were clear about this from the start). Our aim with the project was metascientific (i.e. evaluate the reproducibility of biomedical experiments with a set of common methods) rather than driven by a particular interest in the findings themselves. This is reflected by our choice of selecting experiments from a random sample of articles from multiple fields, rather than filtering by area of interest or importance. It also underlies our choice to evaluate experiments rather than claims, as this was more statistically tractable and potentially more objective as a meta-research goal.

      To be clear, we don’t feel this approach is inherently better or worse than evaluating claims in the literature, as in the Drosophila immunity article case (i.e. Westlake et al., 2026), which is also an important goal. They are merely approaches that answer different questions. Ultimately, we probably made our choice based on (a) our expertise/interest in meta-research rather than in the fields the replications stemmed from and (b) an attempt to engage Brazilian researchers in the project in a way that was non-confrontational and minimized backlash from their peers. We feel this was valuable for many of the lessons learned, although it also meant learning less about the research findings in question.

      Even though this was not a goal of the study, there is some knowledge obtained about the findings that is indeed largely absent from the current manuscript. We do not feel the current format allows for much discussion of 45 different findings, but we do have plans to address these in future articles (as outlined in our response to point 5). In the meantime, qualitative descriptions of each experiment can be found at https://osf.io/w5z9a. This is already mentioned in the Methods but could be reiterated in the results as well.

      (4) The observation that external observers can predict which experiments are likely to be reproducible is interesting and should be more clearly emphasized.

      We did not go too deep into that finding because we are publishing a separate article focused on the prediction project, which should look into factors that correlate with prediction accuracy, both at the level of predictors (e.g. research field, career level) and of individual predictions (e.g. information taken into account for each answer). We also feel that, given the multiplicity of predictors in the prediction analyses, these findings are a bit tentative, as the strongest predictors may be subject to effect size inflation from the “winner’s curse” effect (as outlined by Reviewer #2). We can try to emphasize it a little more in the discussion (although it already merits a whole paragraph on pages 23-24), but we feel we would be able to discuss it more critically in a follow-up article.

      (5) The manuscript frequently refers to future publications. It would be helpful to clarify what is included in the present article versus what is deferred to subsequent papers.

      Indeed, some of our results did not fit this overarching analysis and were left for future publications. One of them is already available as a preprint, while the others are currently in preparation. Specifically, other results from the project should be spread about across five different articles.

      (a) A narrative article focused on challenges and lessons learned with the project, already published as a preprint at https://osf.io/preprints/metaarxiv/8y3tg_v1 (Amaral et al., 2026).

      (b) An article analyzing the prediction survey and markets results in detail (following the pre-analysis plan detailed in https://osf.io/6av7k/files/pjhgd and adding some exploratory analyses on prediction rationales).

      (c) Three articles describing the results of specific experiments with each experimental method (MTT, PCR, elevated plus maze) along with a discussion of aspects inherent to the method that seem to influence reproducibility.

      We can add this information more explicitly to the Methods section, including the links to the papers that have already been published at the time the manuscript is revised.

      Reviewer #2 (Public review):

      Summary:

      This is an important contribution to science, not only because large-scale replication studies remain rare despite their value, but also because this one focuses on research that was underrepresented in previous large-scale efforts. The findings reveal concerningly low replicability in this field, pointing to a problem that warrants immediate attention. Particularly noteworthy is the study's sampling strategy: by randomly selecting experiments from a wide range of publications based on methods, rather than filtering by research area, importance, or citation counts, the authors have produced results that are potentially more representative of the broader literature than those of previous large-scale replication projects in this and other fields. Overall, this is a fantastic contribution that I will be recommending and using in all my open science talks, and from which I have learned a great deal. Congratulations to the team!

      Thanks!

      Strengths:

      A study of this scale inevitably requires an enormous amount of work and methodological care, and this one is clearly both robust and thoughtfully designed. I want to particularly acknowledge the considerable efforts the authors have made to ensure the robustness of their findings. The use of multiple approaches to estimate replicability, combined with a substantial battery of sensitivity analyses, including a multiverse approach on top of everything else, clearly reflects the authors' genuine commitment to understanding their results and the limits of their conclusions. The transparency and sharing of all protocols, materials, and challenges and limitations encountered is also outstanding.

      We once more thank the reviewer for the compliments.

      Weaknesses:

      There were several instances during my reading of the methodology where I felt the authors relied too heavily on the external supplementary materials, at the expense of basic detail in the main manuscript. I appreciate how overwhelming it can feel to integrate more into an already substantial paper, but without some minimum integration, the reading experience and overall comprehension are too often compromised, at times posing more questions than answers. And it is unrealistic to expect most readers to engage with the extensive supplementary materials provided. Please see the comments below for specific suggestions.

      We do acknowledge that the article currently includes a lot of supplementary material. This includes both supplementary figures/tables relating to the paper and many supplementary methods files (mostly hosted at the Open Science Framework). However, we also note that this is already a rather long paper as it stands and that Reviewer #1 has made the opposite suggestion of simplifying it. Thus, it may be hard to strike a balance that will suit all preferences, and we feel that maybe our attempt has landed somewhere in the middle of both reviewers’ ideal versions of the paper.

      Additionally, I found the discussion rather underdeveloped. There is relatively little engagement with the broader literature, not only with replicability studies from other fields, but more generally with relevant meta-research work on publication bias, blinding, risk of bias, citation practices, etc. Some of the most novel and interesting findings in the paper also receive less attention than they deserve, and the discussion at times reads as a repetition of the results section rather than a critical engagement with them. I would encourage the authors to engage more deeply here, as the study clearly has much more to say. Doing so would further highlight why this study is important for the answers it provides and the questions it can spur. Again, please see the comments below for specific suggestions.

      We can try to engage with some of the above-mentioned literature in more depth in particular replication studies from other fields (some of which have appeared after our preprint (e.g. Tyner et al., 2026) and with the risk of bias and transparency literature (e.g. Serghiou et al., 2021). That said, we note once more that the article (and the Discussion section) are already quite long, and that analyzing each of these articles in depth is likely to be unfeasible.

      Specific suggestions:

      Page 1, abstract: "while t values for replications were positively correlated with researcher predictions about replicability, and negatively correlated with the rate of publications by the original article's last author" - I need to address the question: why t values and not effect sizes, p values, or something else? Update after reading the study: although the authors used others, they seem to place more emphasis on t values, which is not well explained. Without a clear explanation, it just left me wonder why, given that effect sizes would, in principle, be more information.

      Our original plan was to use p values as a predictor (see protocol at https://osf.io/9rnuj), but we later realized this was inadequate as it did not account for effect direction (i.e. significant effects in the opposite direction as the original may yield low p values, but this should not count as replication success). We thus switched to t values to be able to assign positive and negative signs depending on effect size direction. We note that, as we are using non-parametric Spearman coefficients (in which the module of t correlates negatively with the p value), the two approaches are effectively equivalent when original and replication effects have the same direction. This change was accounted for and justified in our list of protocol deviations at https://osf.io/9hj7t.

      Effect size (in relative terms) is already being used in the second predictor in the analysis (i.e. effect size decrease), as our idea was to use one significance-based predictor and one effect size-based predictor, to match what was done for the replication rates). We feel that using relative effects (e.g. response ratios) by themselves may not be as adequate, as for experimental methods with large coefficients of variation and/or low sample sizes (especially PCR ones), one can find large relative effects that are nevertheless far from statistical significance. This also makes relative effects not very commensurable between methods.

      We do believe there is a fair argument, however, to use standardized effect sizes as an alternative to t values (i.e. difference measured in standard errors of the mean) to measure significance/evidence strength. As some replications ended up underpowered, low t values may sometimes be due to insufficient statistical power/low sample size rather than replication failures. Using standardized effect sizes is not devoid of pitfalls (e.g. they can be quite variable when sample size is low), but it is worth doing as a robustness analysis.

      That said, there are a few statistical issues to be decided on how to calculate this (e.g. whether studies should be meta-analyzed using standardized mean differences rather than relative ones for this purpose, or whether an analog of the standardized effect size should be calculated for the log ratio of means). We would have to look more carefully into the multiple possibilities to decide on the best approach (and we do accept suggestions!).

      In the meantime, we note that running the prediction analysis using only experiments with ≥80% power yields a slightly higher correlation of t scores with researcher predictions (ρ = 0.49, p = 0.005), so we do not think that these underpowered experiments affect the trend too much. If anything, they could be masking a higher correlation between researcher predictions and replicability.

      Page 2, paragraph 2: "reproducibility (defined here as reaching the same results when analyzing a set of data)" - In my opinion, this definition is vague enough that it encompasses not only reproducibility (same data, same methods) but also robustness (same data, different methods), and I would therefore recommend providing a more precise definition. The same applies to replicability (different data, same methods), since the definition used does not highlight the importance of using the same methods, and thus also encompasses generalisability (different data, different methods). Explicitly clarifying these distinctions is particularly important as the field grows and the terms become increasingly mixed up and confusing.

      We agree that we should make the description more precise (e.g. “reaching the same results when analyzing a set of data in the same way” for reproducibility and “finding similar results with new data collected under similar conditions” for replicability). We will update these definitions in the revised manuscript.

      Page 2, paragraph 3: "All of these issues raise concerns about the replicability of published results - something that has not been evaluated systematically in the country" - I would suggest providing more information about why those factors may lead to expected lower replicability, ideally with a couple of sentences supported by references. As it stands, less experienced readers may not follow the argumentation and may consider it speculative.

      We would argue that the reader would be correct in this case: the argument is a bit speculative. It does go in the direction of what is generally accepted within the field (i.e. that publication pressure can lead to lower reproducibility for a range of factors), but we’re not sure this connection has been demonstrated empirically, except for indirect evidence (such as the lower reproducibility in papers stemming from top institutions and “trophy journals” in, the higher frequency of positive results in US states with more researchers in Fanelli, 2010, or the higher number of problematic images for highly productive researchers in some countries in Fanelli et al., 2022. We could cite this evidence in the introduction and make the speculated connection more explicit, perhaps adding modeling work as well (e.g. Ioannidis, 2005; Smaldino & McElreath, 2016) to explain why this could be the case. But essentially, our opinion is that the connection remains a speculation.

      Page 3, paragraph 2: "We then opened a public call for Brazilian labs that could replicate experiments using these methods and models, advertised by email, social media and lectures in conferences and institutions, to which 73 labs initially responded" - Since recruiting is an important component of this study, I would recommend providing additional details so the reader can better assess how comprehensive and unbiased the recruitment process was. AND Page 5, paragraph 2: Please provide more information about this open call: how was it advertised, where, and when? This is needed so that the reader can assess its comprehensiveness and potential biases. Even the link provided is not specific enough to understand the process, as it only states: "Calls were open to participants > 18 years old with current or previous experience in experimental research in any field and were advertised via e-mails, lectures and social media."

      We can offer a more detailed description of the recruitment process (e.g. number and distribution of lectures, social media strategy used, etc.), although we would rather do this in a supplementary document so as not to make the Methods section even lengthier. We note, however, that we never aimed to recruit a “representative sample” of labs from the country: we were busy enough trying to get enough labs for the project to happen, and aware that the call would be inevitably biased by our own communication capabilities and personal networks.

      That said, the response rates for different regions of Brazil do generally match the distribution of research labs and graduate programs within the country (with some distortions likely caused by our personal networks, such as the large number of labs in Rio de Janeiro state), and seem to indicate a rather wide dissemination of the call. One way to visualize this would be to present the distribution of corresponding articles from the original studies selected for the replication (or even from the whole sample of articles obtained for experimental selection) along with the distribution of labs at different stages of the project in Figure S3, which generally show similar patterns. This would actually lend support to our statement that “the population of labs that performed replications was largely similar to the one that produced the original results” in the discussion.

      Page 3, paragraph 2: "Based on the expertise of respondents and a feasibility analysis by the coordinating team, we selected 3 outcome assessment methods for replication" - Since this choice determined what was ultimately studied and who could participate, I would like to see more information to understand it: was it based on the most common expertise among respondents? How was feasibility defined and estimated?

      We tried to find the combination of methods that would maximize the number of labs that would be included in the project. This is explicitly stated in our Methods Selection document at https://osf.io/qxdjt, but could be stated more explicitly in the paper as well.

      Page 3, paragraph 3: How was the manual screening performed? Was it done by one or more people? Was there double-screening to ensure reliability of the screening protocol? Did the authors use a specific decision tree or tool? How were conflicts between observers resolved? Were any other validation steps taken to ensure reliability? The same comments apply to the data extraction (who, how many, validation, protocol, etc.).

      We initially used single screening by three different reviewers (see https://osf.io/6av7k/files/u5zdq for criteria), as we were merely looking for a sample of experiments; thus, comprehensive inclusion of all eligible studies was not a priority. After this initial screening step, inclusions were confirmed in a consensus meeting with the three reviewers involved.

      Data extraction was also done by a single individual, but the resulting data led to a protocol that was later checked by two reviewers who had access to the paper and were explicitly oriented to judge whether the protocol consisted in a valid replication. Thus, discrepancies between what was in the paper and what was included in the protocol could potentially be flagged at these stages (as they were in many cases). We do note, however, that this is likely not as effective to prevent errors as having data extracted independently, as reviewers may overlook mistakes more easily when comparing two documents rather than extracting data anew. We did find that some errors in extraction slipped by, such as an MTT experiment where treatment concentration was inadvertently changed from mM to μM in a particular protocol step; this was picked up and corrected by 2 out of the 3 labs, but not by the third one, leading the latter replication to be invalidated.

      Page 3, paragraph 3: As a non-expert, I would need more context about the expected average cost of experiments in this field; otherwise, I cannot assess how representative this sample is or whether potential biases may exist (e.g., cheaper experiments perhaps being expected to be less replicable than more expensive ones). Could expected costs also have affected the reduction in geographical coverage eventually observed in this study (Figure S3)?

      As stated in the manuscript, we initially capped experiments at a predicted cost of R$ 5.000 (around USD 1336 at that time), considering reagent cost alone (as equipment and labor was provided by labs), as mentioned in the manuscript. Exclusion rates for that reason were 12/74 (16%) for MTT experiments, 36/132 (27%) for PCR ones and 4/40 (10%) for EPM ones. This is stated at

      This turned out to be an underestimation in many cases, especially as it did not account for pilot experiments, need for repetition, etc; thus, many experiments ended up costing considerably more than that ceiling. As we had included a contingency fund for those cases which we expected would occur , we avoided removing experiments from the sample for this reason as much as possible. Nevertheless, one elevated plus maze experiment ended up not being replicated for cost reasons, as the necessary rat strain was provided by a single facility in the country, meaning that a large number of rats would have to be acquired and transported to all labs at a cost that we were not able to cover.

      As these costs were covered by the coordinating team, we do not feel that this is likely to underlie the reduction in geographical coverage. Other reasons related to lab structure could have led to labs in less well-resourced regions to leave the project, but they probably has nothing to do with the experiments selected.

      That said, the cost cap does mean that the selection of experiments is not completely representative of the literature, but is enriched in relatively cheap and simple experiments which were able to perform (which was our next step for selecting the final sample of experiments. Exclusion rates due to lack of lab expertise and/or infrastructure to perform the experiment were 21/56 (37%) for MTT experiments, 67/89 (75%) for PCR ones and 7/34 (21%) for EPM experiments.

      We will try adding some of this information to the flowchart in Figure 1, as we agree it provides more context on the representativeness of the selected experiments.

      Page 6, paragraph 2: "(on a scale of 1 to 5)" - Could you clarify whether 1 means no deviations and 5 means everything deviated? Is that how it was phrased to participants? Was there a threshold used by the coordinating team to decide how many deviations were acceptable? (I would briefly clarify all scales mentioned below to allow easier interpretation throughout.)

      The scale ranged from 1 (No relevant differences) to 5 (Very relevant differences that prevent considering the study as a direct replication). This scale was used for both the lab and the validation committee scores, and is described at https://osf.io/xgth2 (debriefing protocol) and https://osf.io/e3fjg (validation protocol).

      For the validation committee, we did use a threshold (any score of 4 or a sum of scores of 10 or more among 3 evaluators) to decide what had to be discussed to decide on inclusion, as mentioned on Page 7 of the Methods. For the labs, we used no threshold labs answered the protocol deviation question as a scale, but the decision of whether to consider the study a valid replication or not was not tied to this score.

      We can make both of these points (meaning of the scale and connection to lab’s decision to consider the replication valid) clearer in the Methods section.

      Page 6, paragraph 4: How were long-text answers (e.g., justifications) reviewed? Was this done manually by one or more members of the coordinating team, or using any text interpretation tool? What steps were taken to ensure the interpretation of these answers was as objective as possible?

      For the initial analysis of justifications, one reviewer read all answers and flagged those that seemed to concern reproducibility of the methods (e.g. “we replicated the protocol exactly as planned”) rather than results reproducibility (e.g. “effects went in the opposite direction”). We then revised these answers among the whole coordinating team to decide whether we should contact the lab asking them to revise them. We can add this information to the Methods section.

      For classifications of the justification into categories (i.e. Table S7), justifications were classified by two independent reviewers based on categories created after an initial inspection of the data, and discrepancies were resolved by consensus. We can add this information to the table legend.

      Page 8, paragraph 1: "If issues were found, the lab and coordinating team reviewed them via email until the sources of errors were identified and corrected (see https://osf.io/58vsx for details)." - Could you please provide information about how often these disagreements arose and briefly explain their causes? I am struggling to understand why these discrepancies occurred and how frequently. Without more detail, the error rate presented in the next paragraph is a little concerning.

      After we extracted data from the lab spreadsheets and summarized the results by code, labs received the results by e-mail and were asked to fill in a form on whether the results were in agreement with what they had found (see details at https://osf.io/nfr6y). Discrepancies in results at least 1 experiment were noted by 36% of the 53 (out of 56) labs that responded. Many of these stemmed from the coordinating team misunderstanding issues such as group identity or experimental unit identification in the spreadsheet. Others had to do with different ways to perform calculations (e.g. relative gene expression or % time spent in open arms). In some cases, simple errors in data transcription or typos caused the discrepancy.

      We were also surprised (and concerned) by the number of experiments in which we later found data errors that were not detected by this process (e.g. 18% of total). Our best understanding of this is that not every lab checked the results with the necessary care, as some errors were quite obvious, as in experiments in which sample size was different, or in which group labels were reversed. Ultimately, agreeing with a form that says “did you find any discrepancies?” may have been performed as a box-ticking exercise with little attention, and was probably not the ideal way to check data which led us to start reviewing results in live meetings afterwards. This is discussed in more detail in our challenges article (Amaral et al., 2026)

      Page 8, paragraph 4: Please provide the version of any package or software used throughout, and make sure to cite R appropriately (R Core Team XXX).

      R 4.5.1 was used for the analysis. We can add this information (which was present in the data repository in the R session info.txt file) and provide the R reference in the manuscript as well.

      In addition, did the authors calculate the log ratio of means (ROM/lnRR) using escalc()? If so, please report this.

      If not, I would recommend doing so, as escalc() implements recommended small-sample adjustments that produce slightly different values compared to a simple manual calculation of log(mean1/mean2).

      Yes, we did use the escalc() function for this calculation (for both the replications and the original effect sizes). We can mention this in the manuscript.

      Page 10, paragraph 1: "Coefficients of variation from the original study were compared to the mean coefficient of variation of its replications using Wilcoxon's signed rank test" - I wonder how these CVs were calculated - whether simply as SD/mean or using escalc() from the R package metafor, which includes a correction for small-sample size. This may affect the fairness of the comparison, particularly since CVs from original studies are expected to be slightly overestimated given their smaller sample sizes relative to the replications.

      We calculated the coefficients of variation as the pooled SD divided by the mean of both group means. The reviewer is correct about the possibility of small-sample effects in this case (which we were not aware of). We will thus look into the possibility of implementing this via the escalc () function in the analysis of the revised manuscript.

      We also acknowledge that this could be a source of bias in the comparisons between original and replication CVs (albeit likely a minor one). That said, we note that sample sizes are not always larger in the replication for some experiments with large original effects, power calculations sometimes yielded lower sample sizes in the individual replication, albeit infrequently. On average, though, replication sample sizes were indeed larger.

      I also have concerns about using the mean CV of all replications and comparing it to a single CV value, as this ignores the uncertainty around that mean.

      This is indeed the case; that said, the CV of the original effect also has random error relative to the true population CV and in that case, there is no way to estimate the uncertainty, as we have a single measure of that parameter. So there is probably no way around ignoring uncertainty in this case.

      We also note that we are looking for evidence of systematic CV inflation across all experiments (rather than for a statistically robust comparison between the CVs of any individual replication). For the sake of measuring this systematic inflation, the use of multiple experiments does allow us to estimate variability at the experiment level which should incorporate the lower-level variability between individual replications if this is not included in the model. Thus, we do not feel that our procedure introduced a systematic bias in the analysis at the experiment-level (although one could argue that it may lead to less precision).

      An additional check could involve calculating the log coefficient of variation ratio (lnCVR; Nakagawa et al. 2015, Methods in Ecology and Evolution; implemented in escalc()) between the original CV and each replication CV, and running a random-effects (or multilevel) meta-analysis that accounts for shared-control non-independence. I believe this would provide a more robust approach, as it does not ignore the uncertainty around the mean CV of the replications - uncertainty that, if neglected, is expected to increase the likelihood of false positive findings. This concern would also apply to the subsequent analysis on absolute means.

      We thank the reviewer for this suggestion, which indeed seems like an option in this case. We will look into this possibility, although we cannot guarantee at the moment that we will implement it, as we were not previously familiar with the method and will have to study it in more detail.

      Page 10, paragraph 2: The change in geographical distribution shown in Figure S3 appears rather striking, with western states disappearing step by step. Should the reader be concerned about the eventual geographical representability of the sample?

      Yes, but there are likely different reasons for that. Labs leaving after being included may have been due to those in less privileged regions of Brazil (e.g. the northern and western regions of Brazil, generally speaking) having more difficulty in persisting in the project. That said, most of the “disappearance” happens between registration and inclusion which usually has to do with the labs not working with the methods that were ultimately included in the project. We also note that most of the states that lose representation were those that had a single lab to begin with, which may make the visual pattern more striking than the actual trend (as states in the South/Southeast also lose labs, but don’t disappear from the map).

      We note again that we never planned to achieve geographical representativeness when recruiting the labs on the contrary, we were aiming to maximize the number of available labs to run the project. That said, we do agree that for the sake of examining whether the population of labs is similar to the one that generated the original experiments (a claim that we do make in the discussion), this representativeness is important to assess. Once more, to allow the reader to evaluate this, we plan to add an additional map to Figure S3 to describe the Brazilian states where the original experiments came from (based on corresponding author affiliations) in which a similar bias towards the South and Southeast Region can be observed.

      Page 15, Figure 3A: I wonder whether adding 95% CIs calculated from the sampling variance of each ratio would improve interpretation and help readers appreciate the real differences between the dots (i.e., means) - along the lines of a forest plot.

      We agree that this would be useful information, and can experiment with the possibility, but our feeling is that the figure will likely become too noisy in cases where the 95% CIs overlap (which are quite frequent). If this is indeed the case, an option to allow the reader to examine this would be better to add an explicit link to the forest plots for each individual experiment (https://osf.io/sx9gv) in the figure legend.

      Page 17, section "Predictors of replication success": It is unclear to me how the decision was made about which results from Figure 4 to present in the text. Intuitively, given that correlations were calculated for both t values and lnRR (and other metrics), I would have expected that whenever a result is highlighted in the text, the authors also report how it changes depending on the metric used - for example, the interesting result regarding the 5-year number of publications, whose correlation is notably lower when using lnRR (−0.31 vs. −0.18). Presenting this nuance in the text would reduce the risk of inadvertently giving the impression of cherry-picking.

      We selected the highest correlation values for each continuous outcome (t score and lnRR) and presented these separately in the text. This is a systematic way to perform the selection, but is obviously subject to the “winner’s curse” effect. We agree that adding both metrics for each predictor would be a fair way to keep this in perspective for the reader, but we would have to think about how to do this without sounding too confusing (as results for the two main outcomes are quite different).

      We do note, however, that the outcomes are indeed different and are expected to vary independently in some cases. For the correlation with replication probability predictions, for example, the effects in opposite directions would likely be expected, as larger original effect sizes will likely lead to larger probabilities to be assigned, but also to a higher possibility of effect size decrease. This low correlation between outcomes is probably something that should be pointed out and discussed in the revised manuscript.

      Page 23, paragraph 1: (this comment should have come during the first % reported, but only in the discussion I realized how important this would be for comparing estimates) I wonder whether the authors should calculate 95% confidence intervals for all their percentages (and those of Errington et al.) using the Wilson method via the function binom.confint() in R, which handles extreme proportions (0% or 100%) more gracefully. This would ensure that uncertainty around these percentages is not neglected and would aid interpretation when comparisons are made.

      We had given this some thought when writing the manuscript – but ultimately opted not to include confidence intervals for our replication percentages and to use the replication rates as descriptive measures only (as done in other replication studies such as (Errington et al., 2021).

      Even though we aimed for our sample of original experiments to be as systematic as possible, it is ultimately constrained by many factors (the choice of methods, the particular expertise of the labs, etc.) thus, adding confidence intervals represents the uncertainty around the replication rate of a very specific population of experiments, which is not directly comparable to those included in other replication efforts in any case.

      We will reconsider whether we should include confidence intervals for replication rates: although doing this for every replication rate in Table 1 and Table 2 may end up being too much information, it could probably be done at least for the replication rates of the main analysis in the text. We note that calculating confidence intervals for percentages is straightforward, requiring only the numbers that are in the table thus, any reader that wants to estimate uncertainty for those rates should be able to do it easily.

      We will also point out the uncertainty around the percentages mentioned in the discussion when comparing our replication rates with those of other studies, which we agree is an important issue to touch on.

      In addition, in the next sentence, the authors are comparing correlation coefficients, at least verbally, these could in principle be transformed into Pearson's r and assigned 95% confidence intervals following meta-analytic workflows, which would better allow us to assess whether these correlations are meaningfully larger or smaller, and help avoid potentially misleading arguments.

      Both correlations in that case are non-parametric (e.g. Spearman’s ρ), so they cannot be directly transformed into Pearson’s r without making assumptions about the distribution (which we would probably avoid doing given the very marked outlier in our own). We can calculate a non-parametric confidence interval for our own correlation coefficient by resampling, but we will have to investigate whether this can be done using the available data from (Errington et al., 2021) (which is probably the case if effect sizes for all experiments have been shared).

      Page 24, paragraph 2: The following result is really interesting and I would love for the authors to expand on it a little. There must be other meta-research studies that, despite not studying replicability directly, have explored a similar predictor: "Other features of the original article were generally uncorrelated with replication outcome, although large rates of publications by the last author were associated with lower replicability, suggesting that incentivizing publication volume may be counterproductive for the reliability of results."

      It is indeed interesting, and seems to confirm an intuition that has long been present in the reproducibility field, but actually has little evidence to support it: if anything, there is evidence in the opposite direction in psychology (Youyou et al., 2023), although they looked at cumulative publication number, while we used number of publications in a fixed interval.

      We can expand a bit further on that finding: that said, we do note that the correlation is relatively weak and has a p value of 0.04. Thus, given the multiplicity of predictors would not be that unlikely to occur by chance, even though it seems intuitive. Thus, even though the relationship seems intuitive, we think it should be considered tentative at best and would refrain from discussing it in too much detail.

      Page 25, paragraph 1: I believe the authors could explore if there is evidence for "incorrect labeling of error bars (Cumming et al., 2007; Vaux, 2004)" by plotting log(SD) vs log(mean) across all original studies, and exploring if large outliers (i.e., points largely deviating from the positive regression) exist. That should provide some insights into whether some values reported as SD in the original studies were indeed SE, which I am assuming is what the authors of the study are referring to when they say "incorrect labelling of error bars" here.

      Yes, that is what we mean by “incorrect labeling of error bars” (as can be grasped from the cited references).

      We can perform this regression, which seems relatively straightforward to do. That said, we note that another likely cause for outliers at least for cell line studies would be the use of different (and eventually inadequate) experimental units (e.g. having error bars that represent technical replicates of the same measurement rather than truly independent experiments). We suspect that this may have an even greater effect in terms of causing error bars not to express the same thing and the regression will not help in differentiating the two causes.

      We should also note that different types of experiments may be expected to have very different SDs, so the regression is likely to have a lot of error associated with it. In particular, it’s probably worth doing separate regressions for each method, to account for the likely difference in CVs between animal and cell line experiments, for example. This could also help tease apart the two causes above, as the experimental unit problem mentioned above will likely only be observed for cell experiments.

      Code: I could not engage with the data and code, but I would like to highlight that the organisation and clarity of the GitHub repository is of high quality.

      Thanks!

      Reviewer #3 (Public review):

      Summary:

      The authors conducted a large-scale replication effort of lab-based biomedical experiments with an emphasis on the country of origin and who conducted the replication experiments. The authors aimed to understand this context in both the outcomes produced, but also in the approach. Finally, the authors aimed to conduct multi-lab replications to provide richer data from the replications. Overall, the authors find replication rates that are like other large-scale replication efforts in the biomedical space. The authors provide rich detail into the three experimental techniques that were the focus of this effort, potential moderators of replication success, and challenges in conducting replications and coordinating a large-scale crowd-sourced effort.

      Strengths:

      The paper is outstanding in being transparent and calibrated in how the results are presented. While the authors were challenged by mundane aspects (e.g., difficulty with logistics), unexpected aspects (e.g., COVID pandemic), and very insightful aspects unique to conducting replications (e.g., experimental issues). The authors also provide variation in how they present the results, including confirmatory, multiverse, and exploratory analysis. A unique strength for this study is the rich in-depth insights about the process and interpretation of conducting replications, including predicting replication success in the lab-based biomedical space.

      We thank the reviewer for the compliments. Again, a more extensive list of insights can be found in our challenges article (Amaral et al., 2026), which we will cite in the revised version.

      Weaknesses:

      The study has weaknesses that the authors acknowledge in their discussion, such as lower number of replications than originally planned that limited the intended effort to compare multiple experiments with multiple attempts against a single original experiment. Another weakness is the limited discussion connecting these findings to the Brazilian research ecosystem.

      We acknowledge the missing replications as a weakness, and we hope we have made that point clear in the discussion.

      Concerning the Brazilian research ecosystem, we could try to explore this in more detail in the introduction. In particular, we believe that a better understanding of the Brazilian academic system, including its regional disparities and the general composition of its workforce (which is largely composed of undergraduate and graduate students), can be useful in interpreting some of the findings.

      We can try to provide a bit more context at the end of the introduction (perhaps between the last 2 paragraphs, which would also address a point made by Reviewer #1), and also in different points of the discussion including those comparing replication rates with other studies or discussing infrastructural difficulties, some of which may be specific to the Brazilian context (such as difficulties in acquiring specific reagents or licenses). Still, we reiterate that, due to the lack of studies with comparable samples in other regions, we cannot tease apart the factors that are specific to Brazil from those affecting lab biology as a whole from the data alone.

      References:

      Amaral OB, Neves K, Wasilewska-Sampaio AP, Carneiro CF. 2019. The Brazilian Reproducibility Initiative. eLife 8:e41602. DOI: https://doi.org/10.7554/eLife.41602

      Amaral OB, Valério B, Carneiro CFD, Mota GPS, Neves K, Abreu M, Tan PB. 2026. Challenges for building up confirmatory science in lab biology: lessons learned from the Brazilian Reproducibility Initiative. MetaArXiv, DOI: https://doi.org/10.31222/osf.io/8y3tg_v1

      Errington TM, Mathur M, Soderberg CK, Denis A, Perfito N, Iorns E, Nosek BA. 2021. Investigating the replicability of preclinical cancer biology. eLife 10:e71601. DOI: https://doi.org/10.7554/eLife.71601

      Fanelli D. 2010. Do pressures to publish increase scientists’ bias? An empirical support from US states data. PLoS One 5:e10271. DOI: https://doi.org/10.1371/journal.pone.0010271

      Fanelli D, Schleicher M, Fang FC, Casadevall A, Bik EM. 2022. Do individual and institutional predictors of misconduct vary by country? Results of a matched-control analysis of problematic image duplications. PLoS One 17:e0255334. DOI: https://doi.org/10.1371/journal.pone.0255334

      Ioannidis jpa. 2005. why Most Published Research Findings Are False. PLoS Medicine 2. DOI: https://doi.org/10.1371/journal.pmed.0020124

      Serghiou S, Contopoulos-Ioannidis DG, Boyack KW, Riedel N, Wallach JD, Ioannidis JPA. 2021. Assessment of transparency indicators across the biomedical literature: How open is open? PLOS Biology 19:e3001107. DOI: https://doi.org/10.1371/journal.pbio.3001107

      Smaldino PE, McElreath R. 2016. The natural selection of bad science. R Soc Open Sci 3:160384. DOI: https://doi.org/10.1098/rsos.160384, PMID: 27703703

      Tyner AH, Abatayo AL, Daley M, Field S, Fox N, Haber NA, Hahn KM, Struhl MK, Mawhinney B, Miske O, Silverstein P, Soderberg CK, Stankov T, Abbasi A, Aberson CL, Aczel B, Adamkovič M, Albayrak N, Allen PJ, Andreychik M, Awtrey E, Axxe E, Azevedo F, Bader MD, Bago B, Bailey J, Bakker M, Banik G, Banks GC, Baskin E, Batruch A, Beatteay A, Behr SM, Berente N, Berry Z, Białkowski J, Bodroža B, Boeschoten L, Bognar M, Bokhove C, Bonfiglio D, Bouwman R, Brady TF, Braithwaite SR, Briceño Jiménez G, Brick C, Bricka T, Briker R, Brown AN, Brown GDA, van Aert RCM, Caldwell K, Capitan S, Capitán T, Chandler J, Charles T, Chartier CR, Chawdhary R, Cheng KJ, Chopik WJ, Clark B, Colvin VE, Comer CC, Costantini G, Coupé T, Cummins J, Czernatowicz-Kukuczka A, de Leeuw J, Dobolyi D, Druckman JN, Duan J, Dujmović M, Dunleavy DJ, Durkee PK, Emery C, Esterling KM, Evans TR, Fedor A, Fernández-Castilla B, Fiala N, Field JG, Fong N, Fonseca MA, Freeman ALJ, Freese J, Geiger SJ, Geng J, Getz LM, Geven LM, Gleibs IH, Gonzales DP, Gooty J, Gourdon-Kanhukamwe A, Greculescu C, Griffin SM, Grigoryan L, Grunow M, Gunby N, Hall B, Hanel PHP, Hannon EE, Harper S, Held MJ, Hickman L, Higgins NC, Hippel S, Hoeppner S, Hong S, Hostler TJ, Inzlicht M, Izydorczak K, Jaeger B, Jankowsky K, Jarke-Neuert J, Jensen M, Jokić B, Jolles D, Jolly P, Jones AM, Juanchich M, Kačmár P, Kapoor H, Keljanovic A, Koirala S, Kołczyńska M, Kouroupaki D, Kühnen U, Landgrave M, Larson MJ, Laulié L, Lawrence ACE, Le Forestier JM, Leahy KE, Lee S, Leslie J, Lewis SC, Limnios C, Lin H, Liu A-C, Lloyd JW, Ludvig EA, Lynott D, MacDonald J, Mallik P, Mallinson DJ, Marinazzo D, Martarelli CS, Matacotta J, McBride A, McHugh C, McMillan G, Méndez E, Metzger M, Michaelides MP, Michalak J, Micheli L, Miller JK, Milyavskaya M, Molden DC, Monjaras AG, Moreau D, Morrow A, Moya C, Mudrik L, Mulder LB, Munt KA, Nandi A, Nason K, Nast C, Nave G, Nax HH, Neubauer F, Nguyen PLL, Nichols AL, Nilsonne G, O’Boyle E, Oettinghaus J, Oh J, Oshana A, Ostermann T, Ostrowski RP, Oyebanjo A, Panczak R, Patrianakos J, Pavez I, Pavlov YG, Persson S, Perugini M, Peters K, Pieters C, Ponizovskiy V, Porter ND, Prenoveau JM, Purić D, Purol MF, Puthillam A, Quinn KA, Ramljak M, Reed WR, Ritchie M, Ritzau M, Roche SP, Rodela R, Röer JP, Ropovik I, Rothschild J, Saal J, Safadi H, Samaha J, Sanchez M, Sankaran S, Santos D, Sargent AC, Sauter M, Schmidt K, Schnabel L, Schroeder AN, Schuetz SW, Schuetze BA, Schulte-Mecklenbeck M, Schütz A, Sevigny EL, Shackleton E, Shafranek RM, Shaki S, Shakya S, Sirota M, Sisco MR, Sitnikov MM, Slevc LR, Smalarz L, Smith CT, Snyder JS, Sommet N, Sonmez F, Spellman BA, Stanulewicz-Buckley N, Stock G, Street CNH, Strømland E, Sundelin T, Syed M, Szabelska A, Szaszi B, Szumowska E, Tagat A, Täuber S, Tay L, Thapa S, Thatcher J, Tsaklakidou D, Tummers L, Turkovich E, Tutor MV, Urbanska K, van ’t Veer AE, van Assen M, van de Ven N, van den Goorbergh R, Vargo EJ, Vaughn LA, Vazire S, Vermeulen JM, Vo DTH, Volkman V, Wagenmakers E-J, Wagner D, Walasek L, Walter F, Warmelink L, Wei L, Weißflog MI, Weller N, Wichman AL, Wilbiks J, Williams JR, Wolfe K, Wort F, Wright R, Wulff JN, Xue X, Yan VX, Yang Y, Yoon S, Žeželj I, Zhang Y, Ziano I, Zogmaister C, Zupan Z, Zwaan RA, Nosek BA, Errington TM. 2026. Investigating the replicability of the social and behavioural sciences. Nature 652:143–150. DOI: https://doi.org/10.1038/s41586-025-10078-y

      Westlake H, David F, Tian Y, Krakovic K, Dolgikh A, Juravlev L, Bournonville TE de, Carboni A, Melcarne C, Shan T, Wang Y, Mu Y, Kotwal A, Pirko N, Boquete JP, Schüpfer F, Rommelaere S, Poidevin M, Liu Z, Kondo S, Ratnaparkhi GS, Chakrabarti S, Liu G, Masson F, Xiaoxue L, Hanson MA, Jiang H, Cara FD, Kurant E, Lemaitre B. 2026. Reproducibility of scientific claims in Drosophila immunity: A retrospective analysis of 400 publications. eLife 15. DOI: https://doi.org/10.7554/eLife.108404.1

      Youyou W, Yang Y, Uzzi B. 2023. A discipline-wide investigation of the replicability of Psychology papers over the past two decades. Proceedings of the National Academy of Sciences 120:e2208863120. DOI: https://doi.org/10.1073/pnas.2208863120

    1. Author Response:

      We thank you for this assessment of our work and the positive assessment of the overall theoretical framework. We can fully answer the concerns, in particular regarding data quality, and will provide detailed answers in the following directions:

      “insufficient description of the data”: We will describe the data as much as possible and will share the data and the analysis code.

      “lack of included equations and code”:  We will share the mathematical equations in the supplementary material, and the full code on an online repository. The reason why our data repository (10.5281/zenodo.18480481) is not yet public is that it cannot be changed after publication. For the review process, we provide a github link to data and code here https://github.com/oliviercotto/eLife_epidR. We will ultimately share the link to the final version of the files on Zenodo.

      “definitions of antibiotic use that are not complete”: We will complete the definition of antibiotic use, which is the use of any antibiotic between 7 days and 3 months before sampling. Children who used any antibiotic 7 days before sampling were not included in the study. The type of antibiotic used is given in supplementary material S1: 93% of the antibiotics prescribed are beta-lactams (amoxicillin, amoxicillin/clavunalate, oral 3rd generation cephalosporins).

      “low sensitivity of assays for carriage”: the carriage study conducted in Sweden is used to get plausible estimates of carriage duration parameters in infants.

      • Strain definition is based mainly on randomly amplified polymorphic DNA (RAPD), not colony morphology. Strains with distinct morphology but the same RAPD profile are considered one strain. Conversely, it was checked that strains of the same timepoint with the same morphology most often had the same RAPD profile.

      • We did check that these data are not much affected by imperfect sampling: observations of a strain ‘disappearing’ from sampling then ‘reappearing’  at later timepoints are rare (14 out of 273 strains). This is why we did not correct these occurrences in the previous version of the analysis. In the revised version, we will add a description of these occurrences and correct them. This correction did not significantly alter the inferred parameters in our preliminary analyses.

      • At a broad level, the fact that E. coli clades vary in their carriage duration is very well established across multiple independent datasets; the precise value of carriage duration difference for “persistent” vs. “transient” that we inferred here (a two-fold difference, supplementary material S3) is actually relatively conservative, in the sense that other studies have detected more important differences. We will create a table summarising available evidence on colonization parameters of E. coli to show that the insights from the Swedish data are qualitatively robust.

      • Yet, we will conduct a range of sensitivity analyses to see how the inferred costs of ESBL resistance vary when varying differences in carriage durations, competition and niche differentiation.

      “technical issues with statistical prior selection and parameter identification”: We disagree there is a “technical issue” with prior selection: The fact that resistance is costly, hence that our priors are left-bounded at 0 for the cost parameters, is a prior expectation based on the observation that resistances do not go to fixation. If resistances only conferred an advantage in treatment, but zero cost, then they would quickly evolve to 100% frequency–contrary to what is observed in virtually all epidemiological studies of resistance. That said, we will relax the definition of these priors to test that the data is also compatible with a strong cost on some traits, and no cost or a “negative cost” on other traits.

      Regarding parameter identification and the specific comments on the inference of colonisation parameters: we re-inferred all colonisation parameters with direct inference assuming specific functional forms. This does not alter much the final colonisation parameters that we then use for our main inference. We will also conduct sensitivity analyses to examine how changing some of the colonisation parameters (carriage duration, competition and niche differentiation)  would alter the main inference.

      “application of non-regional ECDC surveillance data to France”: We will clarify our text, as there is a misunderstanding here: we do not use non-regional ECDC surveillance data for inference. We use ECDC data (i) for illustrative purposes, to show that trends in ESBL in France in this surveillance system are very similar to those observed in France in our focal dataset, thus showing the consistency and representativeness of our data. (ii) to give an overview of the weak and inconsistent association of ESBL with age across Europe, thus supporting the relevance of our approach even if our data concerns infants and children. We will make sure this is clarified in the updated version of the manuscript.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have carefully addressed the insightful comments provided by the reviewers which thoroughly increased our comprehension of the dynamics of centriole amplification. The manuscript has been revised accordingly and put in the context of the two papers we published since our last submission, showing that MCC differentiation is a genuine cell cycle variant. A point by point answer to all reviewer comments is provided below.

      Briefly:

      We have streamlined terminology and nomenclature in text and figures / better define experimental conditions with nocodazole

      We have tested the role of dyneins in the dynamics of centriole amplification

      We have done correlative light and electron microscopy on the early stages of centriole amplification

      We have analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors

      Collectively, this allowed us to make a clearer parallel with what occurs during centriole duplication and to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles.

      We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Public Reviews:

      Reviewer #1 (Public Review):

      The manuscript by Boudjema et al. describes the cellular events underlying centriole amplification and apical migration to allow the assembly of hundreds of motile cilia in multi-ciliated cells. For this, they use cell culture models in combination with fixed and live cell imaging using antibody staining and fluorescence from endogenously tagged centriole and deuterostome markers, respectively. The work is largely descriptive and functional analyses are restricted to treatment with the microtubule depolymerizing drug nocodazole. The imaging is state-of-the-art including confocal microscopy, live imaging with optical sectioning and high optical and temporal resolution, as well as super-resolution imaging by ultra-expansion microscopy.

      The study does a good job of providing a very detailed description of the dynamics of centrioles and deuterostomes that lead to centriole amplification and apical migration in multiciliated cells. This detailed view was missing in previous work. It also reveals the involvement of microtubules at multiple steps: the formation of a cloud of deuterostome precursors, the nuclear envelope tethering of newly formed centrioles, their separation, and their migration to the apical surface.

      It would have been useful to expand the analysis of the role of microtubules by including analyses of the requirement for specific microtubule motors, for a better understanding and additional evidence that microtubule-based transport is involved. A weak point is that there is no visualization of microtubules together with deuterosomes and centrioles at the different steps of centriole amplification and migration, to directly address how these structures may interact with and move along microtubules.

      Overall, apart from experimental aspects and since this is largely a descriptive study, the manuscript would benefit from more precise language and a better description of the complex events underlying centriole amplification and movements.

      We have streamlined terminology and nomenclature, clarified the description of the complex events, and test the role of dyneins in centriole amplification. Microtubules density in MCC does not allow to extract information from imaging. In addition, we have done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Altogether, our new data allowed to demonstrate that centriole biogenesis in the MCC cell cycle is marked by the superimposition of 2 canonical centriole cycles. We believe the manuscript will interest a broader readership since it now provides more fundamental insights on the mechanism of centriole biogenesis.

      Reviewer #2 (Public Review):

      This important work will be of interest to centriole and cilia cell biologists. It describes in detail how microtubules control multiple aspects of centriole amplification in brain multiciliated cells. This study provides a greater time-resolved and molecular proteomic mapping of the different steps involved, with or without microtubule disruption. Boudjema et al. show that microtubules are important throughout the centriole amplification process, from the early stages, where the procentrioles emerge from a pericentriolar "nest", through the growth stage where microtubules maintain the perinuclear localisation, to the detachment stage, where microtubules assist in perinuclear disengagement and apical migration. The results are generally well supported by the evidence, but the manuscript would benefit significantly from some heavy editing to introduce more niche terms, standardize abbreviations in text, and labels on figures to help bring the readers, especially non-specialists, along with them - increasing the accessibility of their work.

      We thank the reviewer for his/her enthusiasm. We have streamlined terminology and nomenclature and clarified the description of the complex events to increase the accessibility of our work. We also replied points by points to his/her specific comments.

      Reviewer #3 (Public Review):

      Summary:

      In this manuscript, Boudjerna and Balagé et al. aim to elucidate the spatial origin of centriole amplification and the mechanisms behind the formation of an apical-basal body patch in multiciliated cells (MCCs). To this end, they focused on the role of microtubules and developed new tools for spatiotemporal and high-resolution analysis of different stages of centriole amplification, including the centrosome stages, A-stage, G-stage, and MCC-stage. Among these tools, the MEF-MCC cells grown on micropatterns stands out for its versatility as it is not tissue-specific and does not require epithelial cell-to-cell contact for differentiation. Additionally, the CEN2-GFP; mRuby-DEUP1 knock-in mouse model was used to study different stages of centriole amplification in physiological brain MCCs. This model offers an advantage over the previously described CEN2-GFP model by enabling the resolution of early events in centriole amplification through the visualization of DEUP1-positive structures and their dynamics. Finally, the authors leveraged powerful imaging techniques, including super-resolution microscopy, the U-ExM, and high-resolution live cell imaging in order to detect and track centriole amplification, elongation, disengagement, and migration.

      By combining the MEF-MCC and knock-in mouse model with spatiotemporal imaging in control and nocodazole-treated cells (treated acutely or chronically), the authors define the sequence of events during centriole amplification, revealing the critical roles of microtubules for the first time. Initially, the centrosome-mediated microtubule network forms, organizing a pericentrosomal nest from which procentrioles and deuterosomes emerge. Their findings indicate the importance of microtubules in recruiting and maintaining pericentriolar material clouds that contain DEUP1, PCNT, SAS6, PLK1, PLK4, and tubulins. Following the amplification stage, the procentrioles mature, leading to cells displaying numerous MTOCs, as demonstrated by regrowth experiments. Mature centrioles then disengage from deuterosomes, attach to the nuclear envelope, and migrate to the apical surface facilitated by microtubules.

      Strengths:

      The manuscript provides new insights into the regulatory function of microtubules in centriole amplification. Addressing the role of microtubules during different stages of centriole amplification required the development of new tools to study brain MCCs, which will be useful in future studies of MCCs. A notable strength of this manuscript is the authors' thorough and quantitative analysis of highly dynamic processes in MCCs. The precision and detail in describing these dynamic events are impressive. This comprehensive analysis advances our understanding of MCC biology.

      Weaknesses:

      The role of microtubules and other molecular players during different stages of centriole amplification in brain MCCs can be further studied and strengthened using the tools developed in the manuscript. A more quantitative description of some of the analysis performed in the manuscript is required to strengthen the conclusions.

      We thank the reviewer for his/her enthusiasm. We have tested the role of dyneins in the dynamics of centriole amplification, done correlative light and electron microscopy on the early stages of centriole amplification and analyzed a new single cell RNA seq dataset comparing canonical and MCC cell cycle variants in mouse brain progenitors. We also replied points by points to the reviewer specific comments.

      Recommendations for the authors:

      As you will see, all reviewers felt that the analyses of the involvement of microtubules should be strengthened by including controls and additional experiments. Also, they agree that significant text editing would help to improve the manuscript's accessibility and readability.

      Specifically, they would suggest (1) streamline terminology and nomenclature in text and figures; (2) better define experimental conditions with nocodazole (concentrations used, effect on microtubules, effect on canonical centriole duplication); and (3), in the absence of other complementary genetic perturbation experiments, add a limitations paragraph in the discussion about conclusions drawn from nocodazole treatment alone.

      Reviewer #1 (Recommendations For The Authors):

      Main issues:

      (1) The authors use variable terminology to describe the same or similar events/structures. For example, in Figure 1 they refer to "centrosome stage" where they observe a pericentrin "cloud", which they later refer to as a "nest". In all other figures the first stage is not referred to as the "centrosome stage" but as the "cloud stage". Again, they also describe the "cloud" as a "nest" occasionally, but not always. In the cartoon, the nest is termed "centrosome cradle". The variable and inconsistent use of terms is confusing and the authors do not provide any explanation for the use of one vs. another.

      The text is now corrected. The centrosome stage corresponds to the stage preceding the beginning of centriole amplification in MCC progenitor. The pericentrosomal cloud of centriole and deuterosome elements forms later on, during the amplification A-stage. The formation of this cloud marks the beginning of A-stage, and persists up to G-stage where it dissolves. When we show that the cloud hosts the first stages of centriole biogenesis, we defined it as a “nest”. We do not use anymore the term craddle.

      (2) What prompted the authors to use the term "nest"? It gives the impression that they describe aspecific physical entity/structure (also depicted in this way in Figure 3P, with microtubules outside of this structure), but what is the evidence for this?

      The cloud is the spatial entity and the term “nest” is used to define a function of this transient compartment. We decided to keep the term “nest” as we now identified it with correlative light and electron microscopy, in addition to U-ExM, and show that the accumulation of centriole and deuterosome elements is accompanied by the formation of immature procentrioles, deprived of MT walls, as well as immature and empty deuterosomes. The scheme with MT outside the cloud/nest is misleading as we see MT organized by the mother centriole. We have now changed this.

      (3) The "nest" may simply be a dynamic accumulation of precursor particles around the centrosome, similar to what has been described for centriolar satellites. Rather than proposing a new entity, I suggest testing whether the "nest" particles may colocalize with PCM1 and thus may be related to centriolar satellites. Based on the data, the nest would simply be the centrosomal MTOC that organizes a radial microtubule array on which particles move around its center. In the absence of other evidence, I am not convinced that a new term is needed.

      We totally agree with the reviewer: the centrosome, as MTOC, concentrates centriolar and deuterosome components. This cloud is consistently dissolved when MT are depolymerized or dyneins inhibited. So, the physical entity is a “cloud”. We used the term “nest” to propose one function for this cloud which is to form deuterosomes and centrioles, before they move away for maturation. In fact, deuterosome and centriole formation are hindered when the cloud is dissolved. We have tried to edit the text all over the manuscript to make it clearer.

      (4) Role of MTs: are microtubules required or do they just facilitate some of the investigated events?

      The reason why the role of MT has not been tested yet during centriole amplification is probably because MT not only constitute the cell cytoskeleton on which molecular motors ride to transport cargos or distribute forces, they are also the core component of the structures we are studying. This is why we have tested a range of nocodazole concentrations and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (Fig. 4 Supplementary 1A-B). This may lead to an underestimation of the role of MT but we cannot study the role of MT on centriole amplification if centrioles cannot be formed.

      Does multi-ciliation in these models eventually occur normally under the concentrations and treatment conditions used here? This should be tested and discussed in the context of whether microtubules are indeed required and at what step of the entire process (amplification, migration, ciliogenesis) they may be critical.

      We did both chronic and acute treatments.

      Chronic treatments were done to test the overall efficiency of centriole amplification when MT (or dyneins) are perturbed. Chronic treatments were used to assess the role of MT (or dyneins) on the global efficiency of centriole and deuterosome formation (number of cells able to amplify, number/size/loading of deuterosomes, final number of centrioles (Fig. 4H-I, Fig. 4 Supplementary 2 B-D). In these chronic treatment, we focused on centriole amplification and not ciliation since it was the scope of this study. Also, we did not take ciliation as a readout of amplification because ciliation is relying on MT polymerization.

      Then, we also did acute treatments to test the role of MT (or dyneins) at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration; Fig. 4, 5, 7, 8 and associated supplementary figures). Since one stage is dependent on the precedent one, this enabled us to decipher the direct role of MT (or dyneins) on each single stage. We have now edited text, methods, legends and pictograms to be clear on whether acute or chronic treatment was done.

      (5) Can the authors include control (non-amplifying) progenitors in their analyses? It would be useful to know what the signal and distribution of each specific marker are before differentiation begins (before the cloud stage).

      Non amplifying progenitors are analyzed and constitute the so-called “centrosome stage”. We have now precised it and called it the “progenitor stage”.

      (6) Figure 2: Again, the terminology is confusing, since the authors describe that DEUP1 forms a "cloud" with centrin during the A stage.

      Corrections have been done as explained in point 1.

      (7) Description Figure 3: the authors introduce yet another term: "halo" A-stage. Is this the early A stage? Again, this is not explained and confusing. More systematic and consistent description is needed.

      Corrections have been done as explained in point 1. The term halos is used un the lab as it was the first term we used in our Nature paper in 2014 in reference to the halo described by Erich Nigg when they overexpressed Plk4. It was an error to use it in the manuscript.

      (8) Nocodazole treatments: the used concentrations are quite high.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely (Fig. 4 Supplementary 1A-B).

      (a) To avoid non-specific effects the authors should test what the minimal concentration is that completely depolymerizes microtubules in their cell model and perform analyses at this concentration.

      We have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A-B), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). In case it was not clear, we refer to this now several time and more clearly in the text and methods.

      (b) They should demonstrate depolymerization of microtubules by microtubule staining in the acute and chronic noc treatments and at the different noc concentrations used.

      This is, and was, in supplementary material (same, Fig. 4 supplementary 1A).

      (c) The authors should demonstrate that the used nocodazole concentrations do not impair normal centriole biogenesis during the cell cycle in these cells; if so, impaired assembly of centriole wall MTs may contribute to the observed effects in Figure 4.

      As mentioned in point 8b, we have of course tested a range of nocodazole concentrations at the beginning of the study (Fig. 4 supplementary 1A), and used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4). The ability of the cells to form centrioles during chronic treatments were always assessed using immunostainings of SAS6 and/or CEN2-GFP signals (now exemplified in Fig. 4 Supplementary 1B). We also did EM analysis on cells treated with the highest doses of nocodazole (Nocodazole 10 uM for 24h) and this showed that centrioles can form with, what seems to be MT walls, in cells totally deprived of cytoplasmic MT fibers (Fig. 4 Supplementary 3-4). However, this does not show that all the cells can, because the number of cells that can be analyzed by EM are not sufficient to conclude. Also, one cannot assess whether MT walls are properly polymerized. However, the absence of MT walls should not change the results of the Figure 4, which are based on DEUP1, SAS6 or CEN2-GFP signals for deuterosomes and centrioles. Also MT depolymerization affects the formation of deuterosomes, which should not be altered by MT wall defects as it is not affected, even when centriole formation is blocked (LoMastro et al., 2024). Last but not least, we now show that blocking dyneins, as a comparable and even greater effect, on the formation of the cloud, deuterosomes and centrioles (Fig. 4C-I and Supplementary Fig. 4), which confirms that MTOC function, rather that MT wall formation, explain the centriole biogenesis alteration shown in Figure 4.

      (9) The authors repeatedly refer to the centriole-to-centrosome conversion of amplified centrioles and how this resembles centriole-to-centrosome conversion during the cell cycle. However, they incorrectly claim that this occurs at the G2/M transition. PLK1-dependent modification occurs at this stage, but conversion and PCM recruitment only occur after mitosis (see original work by the Tsou lab, which needs to be cited here).

      We agree with the reviewer. We have now added additional data to show clearly that centriole biogenesis, which requires two cell cycles to proceed in cycling cells, is accelerated during the MCC cell cycle variant where the elongation and maturation cycles are superimposed. This is now clearly shown in Fig. 3, 5, 9 and discussed.

      (10) Figure 6H-J: the authors claim that at low noc concentration, more D-stage cells showed incomplete disengagement than in controls, but the effect is shown only for the highest 10 µM concentration. Do any eof the phenotypes in Figure 6 also occur at the lowest noc concentration (assuming it depolymerizes MTs)? Again, it is crucial to demonstrate this, to exclude unspecific effects not linked to MT depolymerization.

      An error was made on the figure (but not in the legend). In Figure 6, chronic treatments are at 1 or 5 µM. Only acute treatments were done using 10 µM. In both cases, MT are not entirely depolymerized in these experiments (Fig. 4 supplementary 1A).

      (11) Disengagement, Figure 7: The authors describe that DEUP1 signal spreads all over the cytoplasm and becomes diffuse during this process, but one cannot see a diffusive signal throughout cells in the figures.

      We pushed the contrast to make it clearer but the deuterosomes are still bright at this stage and it is difficult to have both signal clear (now in Fig. 6B). We have also changed the example in video (now video 19) to show it more clearly with DEUP1 channel alone.

      (12) Figure 7: localization of disengaged centrioles at microtubule "nodes" is not clear from the images. There are many centrioles and random colocalization may be expected simply based on the high number. Higher resolution and/or magnification and quantification would be needed.

      We have edited and now say that centrioles “colocalize” with MT which, since centrioles nucleate MT, seems normal. We agree that it could be random, but given the density of MT, and the number of centrioles, it does not seem opportune to us to quantify. We can just say that we never see centrioles is regions that are deprived of MT.

      (13) The term "diffusive" to describe slow centriole movements in Figure 8 suggests that it is not motor or force-dependent, but there is no evidence for that. Movement based on opposing forces could produce a similar result, but would not be considered diffusive.

      We agree. We have changed “diffusive” by “diffusive-like”.

      (14) The manuscript would greatly benefit from the analysis of some candidate motor activities that may drive the movement and migrations of centrioles in this system. This would support the importance of the microtubule network for the specific steps in these processes, and better define its role beyond "being required". Dynein may be a candidate or minus end-directed kinesins. Since chemical inhibitors are available, these types of experiments would be straightforward.

      We formerly tested ciliobrevin but had hard time because of the small stability of the drug. Since our submission to eLife, we tested dynapyrazol and dynarestin and found dynapyrazol very efficient in dissolving the Golgi, a good readout of dynein inhibition. We sought to test the role of dyneins, using dynapyrazol, on (i) the formation of the pericentrosomal cloud in A-stage, (ii) the oscillation of DEUP1+ structures during A-stage, (iii) the number, size, loading of deuterosome, (iv) the final number of centrioles, (v) the migration to the nuclear membrane and (vi) the final apical migration of centrioles. The results are now inserted in main and associated Fig. 4, 5, 7, 8, 9.

      (15) Discussion:

      "the role microtubules" lacks "of"

      This is now edited.

      "This lack is..." Lack of what?

      This is now edited.

      "reflexive link" - meaning of "reflexive" is not clear in this context

      We have removed it.

      In my opinion, the study does not identify a nest composed of DEUP1, PCNT, and Centrin2; it only shows that these components accumulate as particles around the centrosome, which functions as MTOC. Consequently, it seems that the "nest" does not exist when MT is depolymerized. One could consider the center of the centrosomal MT array as a nest in this context, but there is no evidence of a specific new structure as suggested by the way the term is used in the manuscript.

      This is what we want to say: the center of the MT array become a nest in this context. We do not state that there is a specific new structure. We just say that MT and dynein dependent concentration of centriole and deuterosome components exists and that this region nests the birth of centrioles and deuterosomes. Also, this compartment is restricted in time and space, which justifies to use a specific term. The MTOC exists in the progenitor cell, while this compartment, marked by DEUP1, Centrin, PCNT accumulation, appears at the beginning of amplification and grows during A-stage to be dissolved at G-stage when all the deuterosomes and centrioles have moved away.

      What is the evidence that "DEUP1 is a centrosomal protein before building deuterosome structures"? It would be good to refer to the specific experiment. Does DEUP1 localize at centrioles also in the absence of microtubules? If not, I would not consider it a centrosomal protein.

      We have removed this statement to avoid misinterpretation.

      "This reminds the centriole-to-centrosome conversion..." the sentence is missing an "of"; also, again the authors confuse the order of events during the cell cycle, where centrosome conversion occurs after completion of mitosis, not at G2/M transition.

      We have removed this statement to avoid misinterpretation. Also, see Point 9.

      "microtubule dependent nuclear migration" should be rephrased; it sounds as if the nucleus migrates.

      This has been changed

      The following discussion of disengagement being linked to association with the nuclear envelope and resembling the process in cycling cells is misleading. In cycling cells movement of centrioles along the nuclear envelope occurs at G2/M and drives centrosome separation (separation of centriole pairs) in preparation for mitosis, not centriole disengagement.

      We are now clearer. We compare centriole-loaded deuterosome organization around the nuclear membrane to the migration of new centrosomes during early prophase (Fig. 5F-H, Fig. 5 Supplementary 2G-K).

      Regarding the possibility that forces by microtubules generated by the daughter centriole drive disengagement also in cycling cells, I would argue that this is unlikely since the daughter centriole can only nucleate microtubules after disengagement has occurred (and conversion to centrosome/PCM recruitment). Once this happens, it may physically separate the disengaged centrioles, which is a different type of activity. Indeed, originally the term "disengagement" was coined to specifically describe the loss of the perpendicular engagement of daughter centrioles with their mothers (Tsou and Stearns, Nature, 2006).

      We have removed this statement to avoid misinterpretation. The perpendicular engagement is difficult to assess on deuterosomes but we do see by live imaging, that attachment changes during D-stage, before centrioles detach clearly from deuterosomes.

      "high resolutive" should be "high resolution"

      Edit done.

      "splitted" should be "split"

      Edit done.

      "Consistently, when the mitotic oscillator is dis-inhibited and cells enter pseudo-mitotic events, centrioles show clear and rapid cell-cycle like clustering" This sentence is not understandable without further explanation; what does mitotic oscillator refer to? What are pseudo-mitotic events? What is cell cycle-like clustering?

      We have removed this statement.

      Minor:

      (1) Abstract: "Centriole number must be restricted to two..." Since cells are born with two centrioles and have 4 centrioles (2 pairs) when they enter mitosis, this sentence is inaccurate.

      The sentence has changed.

      (2) Abstract: "reflexive link"; I am not sure what the term "reflexive" refers to?

      We have removed this statement to avoid misinterpretation.

      (3) Figure 1C, D: it should be described better that the larger magnification panels represent overlays of many cells and what marker they show. This is not obvious since the smaller single-cell panels always show two different markers. Also, it would be more useful to show also single cells in the magnified view. The overlay does not allow us to see if a marker forms a cloud or a single dot, which is as important as the cell-to-cell variation in distribution.

      We have clarified this in the text and the legend. The cell-to-cell variation cannot be estimated with the overlay, but the projection from several cells (number precised) allows to see that the signal is confined in a restricted region. Or not. Which is what we wanted to analyze.

      Related to the above, the authors say that pericentrin forms a cloud at the top left in panel D, but there is only one confined centrosomal dot in the single-cell panel.

      The sentence has changed.

      (4) Results, Figure 2F; video 4: The authors claim connection and disconnection of DEUP1 aggregates with centrosomal centrioles; can the authors comment on the spatial resolution including in z in this movie to support this claim? Can they exclude that the structures are in proximity of each other rather than "connected"?

      This is a single z-section of 500nm. The resolution in xy is 128nm/pixel. Given the sizes of deuterosomes and a mature centriole, and given the fact that we observed this dynamics in several cells in live, we can state that the structures are connected. This is consistent with deuterosomes frequently observed “kissing” the daughter centriole by EM in the present manuscript (Fig. 2D, Fig. 2 supplementary 3 and 4 and Fig. 4 Supplementary 3-4). One has to look carefully at the daughter centriole (marked “dc”) and span in on the serial sections to see the connected deuterosome (marked by a star): this is at very early stage and therefore it is small. We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both structures are frequently sticked to each others on tens of nanometers.

      (5) The term "dynamics" as used in the manuscript should be plural.

      It has been used plural, except when for “dynamic microtubules” and “dynamic attachment to the nucleus”, which we think is ok? We have not found any other singular uses in our manuscript.

      (6) Figure 5: what does "YL1/2 procentriole intensity" refer to in panel F? This should be the intensity of microtubule asters.

      This has been modified.

      (7) Figure 6 - supplement 1B: contrary to the claim in the text, one cannot see tight colocalization with the nuclear pore marker. This seems to be a very small subset of particles and even in those cases colocalization is not tight. Also, what is the relevance of nuclear pore colocalization?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. What we want to say is that there is a tight connection with the nuclear envelope as shown by the localization of NPC on the same z-section as centrioles. This is why we present a single z, to show that centrioles and NPC are on the same z-plane of 500nm. NPC are stained to outline the nuclear membrane. This is also clearly visible for G-stage centrioles in the XY plane. We have now added an entire z-stack on video 18.

      Reviewer #2 (Recommendations For The Authors):

      To improve accessibility of their manuscript, we would suggest making the following edits:

      (1) Define 'specialist' or 'niche' terms each time you introduce them, such as 'pericentrosomal nest', or 'flower-like structures'.

      This has been clarified.

      (2) Have a think about abbreviations, again ones that work for people outside the project- this paper uses 'PC' for 'procentriole' but for many 'PC' is 'Parental centriole' or Figure 6J talks about 'D total' or 'D partial', leaves readers confused.

      This has been clarified.

      (3) Standardize your abbreviations throughout particularly for your treatments- sometimes Noco sometimes, NOCO, or your imaging experiments sometimes Cen-GFP, sometime CEN2-GFP (Figure 7A, D vs. Figure 6) or DEUP1- mRuby, DEUP1-mRuby3 or mRuby3-DEUP1?

      We now use Nocodazole or Noco in the text and the figure respectively, CEN2-GFP and mRubyDEUP1.

      (4) About 10% of the population, including several key figures in this field, are red-green color blind. Although 4 colour fluorescence is difficult to get right for everyone, choosing palettes (especially for two colour panels) is inclusive. More so, greyscale or inverted monochrome images make it easier for everyone to visualize changes in localization, size, and intensity. Red on black small foci is particularly difficult to discern. For example, Figure 3 - more individual channels in grayscale with arrows to mc, dc, and cilia would be helpful - difficult to distinguish stainings.

      We thank the reviewer for this comment and for this recommendation of being more inclusive. We have done the changes.

      To improve the conclusions drawn, we suggest some revisions below:

      (1) Since the paper really hangs on it, a clearer description of the rationale for when, how long and how much nocodazole treatment was done is needed. The logic currently is difficult to follow seemingly random jumps 10x concentration are used. Microtubules control many aspects of cell biology and could be impacted. For example, I particularly found Figures 6D and H difficult to follow i.e. the timing for 6H seems off.

      MCC develop a very dense and stable MT network that is not comparable to cycling cells. MT are very difficult to depolymerize entirely. We have of course tested a range of nocodazole concentrations at the beginning of the study and shown the extent of MT depolymerization under each treatment. We used concentrations where MT are perturbed but not entirely depolymerized, allowing centrioles to be produced (see answer to point 4 reviewer 1). The level of perturbation of MT and consequences on centriole formation at the different timings and doses were done for each experiment and are exemplified in Fig. 4 supplementary 1A-B. This figure was already present in the first version of the manuscript but we have now edited text, methods and pictograms to clarify this.

      (2) Perhaps an extension of this point- in general how interdependent are the processes? If there is a defect at the nest stage, how much are the later defects secondary to this, or do MTs genuinely play direct roles at all stages or are these knock-on effects? How do the authors rule this out? Defects in the nest, lead to smaller and more DEUP1+ foci, with defects in concentrating procentriole factors and centrin, which lead to... For example, Figure 4B looks like centrin is reduced upon noco treatment? Does noco treatment affect Cetn2GFP levels globally? Individual channels grayscale would help visualise this better.

      See also our answer to reviewer 1 point 8c.

      The stages are indeed interdependent. This is why we did both chronic and acute treatments. Chronic treatments were done to test the overall efficiency of centriole amplification when MT are perturbed. We typically used low dose of 1µM because nocodazole remains 48h in the culture medium. Acute treatments were done to test the role of MT at each stage of amplification (A-amplification, G-growth, D-disengagement, M-migration). Most of the acute treatments were done live and nocodazole was applied after the first time point of live monitoring. We used 10µM to have a rapid effect, and because nocodazole remains only several hours in the culture medium. This allowed to monitor the stage “n”, in cells where the stage “n-1” was completed without any drug which allowed to analyze a stage without having perturbed the precedent one.

      We now also test the consequences of dynein inhibition using both acute and chronic dynapyrazole treatments. We show that except for centriole migration, dynein inhibition phenocopies MT depolymerization (centriole number, perinuclear organization and disengagement as well as deuterosome number/loading/size).

      Nocodazole chronic treatments do affect intensity of CEN2-GFP at G-stage centrioles suggesting an altered A-to-G transition. In D-stage, CEN2-GFP signal seems normal. We now mention this in the text and in the Fig. 4 Supplementary 1B.

      (3) The authors nicely show the importance of MTs in the structure of the nest from which procentrioles and DEUP1 positive structures emerge. They suggest this nest may be what supports procentriole generation in the absence of DEUP1 and parental centrioles. Firstly how does this nest look in the absence of DEUP1 and/or parental centrioles (centrinone treatment)? This may be what they are trying to show in Figure 5 Supplement 1 but it currently is very difficult to digest what it is showing relative to controls and whether this is significant in the way it is plotted.

      The nest is conserved in the DEUP1KO with or without centrosomal centrioles, as shown by accumulation of Centrin and PCNT at the center of the self-organised MT network (Mercey et al., 2019). This is in fact what motivated our study on the role of MT in centriole amplification. We have edited the legend to precise the quantification done, which is not related to this question. In this quantification, we show that the increased propensity to accumulate PCNT by centriole-loaded deuterosomes between A and G-stage is maintained in the absence of deuterosomes, indicating that centrioles themselves accumulate/recruit PCNT.

      (4) Can you do CLEM on DEUP1-Ruby and these early foci at the cloud stage to see if they are visible at the ultrastructural level, relative to procentrioles, microtubules, and other electron-dense structures?

      We thank the reviewer for this question. We have done CLEM on the pericentrosomal cloud during very early steps of centriole amplification. This showed that DEUP1 early accumulation at the centrosome corresponds to a region rich in fibro granular aggregates, suggesting that DEUP1 may be translated here, through locally concentrated centriolar sattelites, known to be involved in local translation. Then, small deuterosomes and immature centrioles are formed, within this cloud of sattelites, confirming that the pericentrosomal cloud is a nest for centriole biogenesis (Fig. 2C-D + Fig. 2 Supplementary 2-6 for control and Fig. 4 Supplementary 3-4 for nocodazole treated cells). This also shows that immature deuterosomes are not necessarily round shaped, and can be deprived of centriole loading.

      (5) Check the scale bars- see Fig 4E. Check throughout.

      Done.

      (6) Figure 3 Supplement 1 and 2 don't match the legend and are likely reversed - which one is right?

      Done.

      (7) Technical issue - I couldn't play videos 6 or 16? Check these work.

      Done.

      (8) Nomenclature mammalian proteins- mouse or human- should be all caps DEUP1, PLK4, SAS6,etc. Watch your units- space between number and unit.

      This has been done.

      (9) Many of the graphs involve three biological replicates but why not plot the mean of each of the three experiments and do stats? The number of events measured may conflate the significance. Try using Superplots.

      Here is how we proceed: we count the number of occurrence of the phenotype we monitor, and the total number of cells. We apply a X<sup>2</sup> to test whether there is a significative difference between our replicates in each condition. If not, we pool the number of occurrence of the phenotype we monitor and the total number of cells for the 3 replicates, and for each condition. Finally we apply a X<sup>2</sup> between the different conditions. This is how we usually proceed to avoid comparing a mean of percentages. This is now explained in the methods.

      Minor points:

      (1) "DEUP1 is a centrosomal protein and assembles deuterosomes in the pericentrosomal region in brain MCC". I am not sure you have evidence that DEUP1 is a centrosomal protein. You don't seem to study the relationship between centrosomes and DEUP1? Rewrite this title and tone down this claim.

      This has been modified.

      (2) Why the crossbow micropattern (versus some other shape) - seems very specific but not discussed?

      We wanted a shape where centrosome is not localized at the center of mass of the nucleus. Among the corresponding patterns, the crossbow was the one where differentiating cells had less propensity to detach.

      (3) Figure 2 - are the foci of DEUP1 at the cloud stage smaller than at A stage? How do they grow? Measure the diameter at cloud stage, just after they leave the cloud and then once they move away from centrosomal cloud and each other. If so, and they do indeed grow in size from the cloud stage to the growth stage which I think your images suggest - do you envision this happening with the gradual addition of DEUP1 rather than fusion?

      Early deuterosomes are not easy to detect by light microscopy, because of accumulation of DEUP1 in the cloud. We did CLEM on the cloud of early A-stage cells to resolve the earliest deuterosomes which are often very small (see Fig. 2D, Fig. 2 Supplementary 2-6) suggesting that they grow, either by fusion, which we never observe in our movies at later A-stage, or by accretion of DEUP1. However, by light microscopy, we can detect very early but big deuterosomes, which we see splitting later on into smaller ones. So, we cannot conclude on the mechanism that regulate deuterosome size. This is now discussed in the discussion of the manuscript.

      You say in the discussion:

      "Consistently, we never observed fusion events of DEUP1 condensates in our time-lapse experiments. More importantly, we did FRAP experiments on endogenously tagged mRuby-DEUP1 in cells at the different stages of centriole amplification, and did not find significant recovery, supporting that centrosomal DEUP1+ foci and deuterosomes are not liquid-like structures (Figure 8 Supplementary 2)." How do you prove there is no fusion of deuterosomes?

      It is always difficult to prove the absence of something, we agree! But we did tens of movies with high temporal resolution and never observed fusion events. But, as we say in the previous question, the very early deuterosomes can be very small and we do not distinguish them from the DEUP1+ cloud by live imaging. So at this stage, we cannot say. But later on, during A- or G-stage and when deuterosomes are outside the cloud to be easily observed, we very often observe deuterosomes bumping into each others and stay in close contact for minutes, but then moving away. This, for us, supports the lack of fusion properties. But the question remains open. We now explain this in the manuscript and have added an example in video 28.

      If they are getting bigger as I think your imaging suggests from cloud to growth stage, then how is this happening?

      MT depolymerisation and dynein inhibition leads to the formation of very small deuterosomes. Dynein inhibition can even lead to a block in the formation of new deuterosomes suggesting that DEUP1 concentration is a crucial parameter for condensation into deuterosomes. Deuterosome growth may happen through oligomerization of DEUP1 molecules allowed by their dyne-independent concentration. Sorokin in 1968 proposed that a supersaturation of deuterosome components may lead to their solid crystallization into deuterosomes. Deuterosome size can also be regulated by a more complex molecular cascade, involving post-translational modifications of DEUP1 or PCM, such as phosphorylations driven by the cell cycle machinery. This would be consistent with the fact that deuterosomes are very big in the absence of CCNO, a cyclin required for entering the MCC cell cycle variant. This will need further investigations.

      I'm not sure FRAP actually proves fusion doesn't happen.

      Agreed, this is not what we wanted to say, we clarified. The FRAP experiment just suggests that it is not liquid-like.

      It is technically difficult to laser ablate individual or only subsets of deuterosomes...

      This is what was done but anyway, FRAP does not firmly show that deuterosome compartments are not liquid-like as we now precise.

      (4) How do you fix your cells for expansion as you have no preservation of cytoplasmic microtubules? You are saying that there is a "nest" of MTs but beta tubulin ONLY stains the cilia and centriole - why is this? Tyrosinated tubulin on regular confocal shows strong cytoplasmic staining. See Figure 3.

      Cytoplasmic microtubules do not preserve well through the expansion process. We did try a few different fixations and pre-extraction methods but they come at a trade-off to preserving centrioles. i.e. we could either preserve cytoplasmic tubes or centrioles but not both with the same processing method.

      (5) "PCNT puncta partially overlap with centrin (Figure 3 Supplementary 2C). At this stage, PLK4, the master regulatory kinase, and SAS6, one of the first centriolar components are either absent or present as small foci within the cloud, often on the wall of the parent centrioles (Figure 3B-C)." some arrows to highlight this would be useful - difficult to see?

      We have tried to make arrows on what is now Fig. 3 Supplementary 1 G, but there is to many CENTRIN colocalizing with PCNT. We have enhanced the contrast of the merge to make it more visible.

      (6) Figure 3I legend - what are the arrows pointing at? Yellow and white on inserts? ". Around the same time as tubulin, centrin is also recruited to procentrioles (Figure 3I). This stage is probably the stage that we previously documented as A"

      However you see centrin at DEUP1 foci in D, and you don't show any eg. SAS6 or PLK4 positive DEUP1+ structures lacking centrin specifically, centrin seems to be present on all the procentrioles in Figure 3I. Did I miss it where you show centrin negative procentrioles in the cloud?

      Fig. 3I (now Fig. Supplementary 1J), yellow arrows are pointing at centrioles with non-acetylated MT while white arrows point at acetylated MT. This is now indicated in the legend.

      Regarding CENTRIN, it is present as a diffuse staining around the centrosome since the very beginning of amplification (now in Fig. 3 Supplementary 1A with different contrasts), in addition to compose the parental centrioles. This staining can therefore overlap with DEUP1 staining when DEUP1 appears (Fig. 3 Supplementary 1B, E) but not necessarily. In live we observe that CENTRIN and DEUP1 foci can move independently at early stages (Fig. 2 Supplementary 1B, video 2). This is later on, as shown now in Fig. 3 Supplementary 1J (previously Fig. 3I), that procentrioles are all strongly positive for CENTRIN.

      A new paper (Laporte et al., Cell 2024) recently showed that the recruitment of CENTRIN on duplicating procentrioles first occurs at the distal end, visible by a small dot, and then appears gradually at the level of the inner scaffold when procentriole reach 160nm, the stage where POC5 appears, which corresponds to the A-to-G transition in our MCC progenitors (Al Jord et al., 2014). One can therefore consider that the same is happening in our cells, and that, with the CENTRIN cloud, we have difficulties to detect the distal CENTRIN dot. We have changed the text to add this reference and discuss CENTRIN apparition in MCC procentrioles.

      (7) " The DEUP1 asymmetry previously described at the centrosomal daughter centriole (Al Jord etal., 2014) becomes visible in some cells during the cloud stage (Figure 3B, N; Figure 3 Supplementary 2B) and in a majority of cells" difficult to see - maybe enlarge and single channel from Figure 3F-H in the supplemental Figure 3 to emphasise this?

      We have either changed the pictures or the contrast to be more representative with the quantifications. This is visible in Fig. 3A, D, E, G; Fig3. Supplementary 1E and now using correlative light and EM in Fig. 2 Supplementary 2, 3, 4 and Fig. 4 Supplementary 3-4. One has to look carefully at the daughter centriole (marked “dc”). We have not zoomed in since previous manuscript have already described this at later stages with bigger deuterosomes. You can refer to main or supplementary figures in previous manuscripts (Al Jord 2014, Khoury Damaa 2024) where serial sections span the entire deuterosomes and daughter centrioles and show, with nanometric resolution, that both strutures are frequently sticked to each others on tens of nanometers.

      (8) Do you have videos of DEUP1 oscillations with nocodazole to show a lack of oscillations?

      We have now added videos of DEUP1 oscillations under nocodazole and dynapyrazole treatments.

      (9) "In addition, co-staining of centrioles and nuclear pore proteins show a tight colocalization(Figure 6 Supplementary 1B)." I see the colocalisation in panel 1 but less obvious with panel 2 maybe have some more zoomed in panels and some quantification of the colocalization? Is it more striking at the G stage than the D stage?

      We edit and change the phrasing as ‘colocalization with NPC’ is not the good term. There is too many centrioles and NPC, they cannot do otherwise than colocalize… What we want to say is that there is a tight connexion with the nuclear envelope. This is why we present a single z, to show that centrioles and NPC are on the same z-plane. This is also clearly visible for centrioles that are loaded on deuterosomes that are around the nuclear membrane in the XY plane. We also added a video to show an entire z-stack of this kind of staining.

      (10) "Indeed, SAS6 normally disappears from procentrioles when centrioles are docked, just beforeciliation (Al Jord et al., 2014). This suggests that centrioles were able to degrade SAS6, a process also dependent on APC/C (Strnad et al., 2007), but failed to disengage from deuterosomes." Figure 6 Supplement 1E-F - are you sure it wasn't that Sas6 wasn't loaded correctly at the earlier stage and so is reduced recruitment rather than premature disengagement of Sas6? If it is indeed premature disengagement of Sas-6 - what about CP110 - does the CP110 get loaded and is it still present in noco treated cells arrested in the D phase?

      We do not observe SAS6-negative procentrioles on deuterosomes at G-stage but only on deuterosomes in D-stage cells (cells with partly disengaged procentrioles). This is why we hypothesize that, because of the long duration of D-stage and knowing that SAS6 is finally degraded at the end of amplification (Al Jord et al., 2014), we are in the presence of cells where SAS6 has been degraded but where centrioles did not manage to disengage. This is now clarified in the text.

      (11) Can you track deuterostome splitting live? Maybe not enough spatial or time resolution?

      One has to monitor in 3D (multiple z because deuterosomes move a lot), 2 colors, high temporal resolution (dt=2-5’; to be able to track a single deuterosome), and long duration (deuterosomes are sometimes touching each other and then moving away, giving the impression that they split). This eventually leads to the bleaching of the mRuby fusion protein… We have put an example of what we think is a deuterosome splitting in Fig. 6E (former Fig. 7D). But we decided to finally monitor with low temporal resolution (dt=40’) to avoid photobleaching, and analyze numerous deuterosomes and cells to quantify the number and size of deuterosomes over time in single cells.

      (12) The MT nodes - can you segment the tyrosinated MTs and define nodes and then quantify theDEUP1 presence on them?

      Please see answer to reviewer 1 regarding this point.

      (13) Figure 8 supp 1 (E): Representative XY distribution of CEN2-GFP+ centrioles at the end of migration (Sas6 negative) in brain MCCs treated with DMSO, Nocodazole 1µM and 5µM (48h). Scale bar, 5µm Bit more detail on how you define fully migrated vs still migrating centrioles in z. You say you are using Sas-6 negativity to define fully migrated cells in the legend, yet you say noco treatment leads to premature sas-6 negativity, and yet the apical migration takes longer upon noco treatment?

      Nocodazole does not lead to premature SAS6 negativity but to a partial disengagement which lead to SAS6 negative “mature” centrioles being still connected to deuterosomes. We define complete migration when all the centrioles are on the apical side of the nucleus. We now clearly define what “apical” migration stands for in the main text and changed the pictograms in Fig. 8G to clarify this.

      (14) Figure 8H and video 18 - it isn't obviously clear to me that the noco-treated cells are "more erratic" or how you decide what counts as apically migrated successfully. How do you control for drift in z? Can you track individual centrioles as you did in untreated and define what is "erratic about their movement?

      Erratic means that the centrioles are moving away from each others, and back, in a non-predictable way, instead of migrating up and gathering. The drift in z of the whole cell is visible because there is always some centrioles, that are apically located at the beginning, that remains on the apical membrane, probably because they are already docked.

      We have indeed followed the centrioles individually in the nocodazole condition. However, in the control, the XYZ coordinates of one of the centrioles of the centrosome, which normally don’t move, are substracted to the coordinates of all the other centrioles as explained in the method section. This allows to have a subcellular reference, and to circumvent the movements of the cell, which are non-negligible at all at this timescale. In the nocodazole treated cells, the centrosomal centrioles share the erratic movements of the other centrioles and can migrate up and down, which exclude them as a reference. Since the nucleus is also moving a lot, we were left with no reference point.

      (15) Figure 8 supplement 1E can you quantify the final area of centriole patch in XY upon noco treatment?

      It was in main Fig. 8J and is now in Fig. 8 Supplementary 1F.

      (16) Figure 8J legend- MBB is never defined as an acronym.

      Thank you for pointing this.

      (17) Define what is the frequency and how is it calculated - Figure 8J.

      This is the MBB patch area in µm<sup>2</sup>

      Text edits:

      (1) "Altogether, these results suggest that, in this non-tissue-specific proxy of MCC progenitors, microtubules organize the onset of centriole amplification in the pericentrosomal region."

      Sentences have changed.

      (2) "Increasing the temporal resolution to 5-15s reveals that DEUP1+ foci observe an exhibit oscillatory dynamics to at the centrosome (Figure 2E, colored arrows, Video 3, 5/10 cells observed for 1-4min)."

      Sentences have changed.

      (3) "stage procentrioles were involved in this perinuclear migration and distribution. In fact, this dynamic is reminiscent of the centrosome migration that occurs during the G2-to-M progression in cycling cells in preparation for mitotic spindle organization. In cycling cells, this" Grammar - maybe change to "stage procentrioles were involved in this perinuclear migration and distribution. This is reminiscent of the centrosome migration that occurs during the G2-to-M".

      Sentences have changed.

      (4) "We then wondered whether these microtubule-dependent dynamics was were required for an efficient subsequent centriole disengagement during the following D-stage."

      Sentences have changed.

      (5) "Then, monitoring tens of disengagement movies, we identified a transient stage during which disengaging procentrioles redistribute isotropically in the 3 dimensions, along the nuclear membrane (Figure 6A, 4:30, Video 7) before losing its contact to migrate to the apical surface (Figure 6A, 6:30 to 14:00)."

      Sentences have changed.

      (6) Discussion: "Since pioneer electron microscopy studies on basal body production in quail oviduct MCC 35 years ago (Boisvieux-Ulrich et al., 1987, 1990; Boisvieux-Ulrich et al., 1989), this work is the first to assess the role of microtubules in the now finely described centriole amplification process. This"

      Sentences have changed.

      (7) "Using live imaging on brain MCC, we highlight the existence of a nest composed of DEUP1, PCNT and Centrin2, pre-assembled before the onset of centriole amplification onset."

      Sentences have changed.

      (8) "Recently, formation of DEUP1 pure condensates in solution as well as FRAP experiments after overexpression of DEUP1 in MCC progenitors suggested that deuterosomes where are not liquidlike structures (Yamamoto & Kitagawa, 2019). Consistently, we never observed fusion events of DEUP1."

      Sentences have changed.

      (9) "This reminds is reminiscent of the centriole-to-centrosome conversion occurring at the G2-M transition followed by the associated microtubule dependent nuclear migration of new centrosomes at mitosis onset (Agircan et al., 2014)."

      Sentences have changed.

      (10) "Following individual trajectories requires high resolutive resolution spatio-temporal live imaging while avoiding excessive light exposure which disturbs centriole migration (Boudjema et al., 2024)."

      Sentences have changed.

      (11) "Using high temporal resolution microscopy, we further identify that individual dynamics is are complex and can be splitted between divided into the baso-apical migration, where centrioles move in a processive and more..."

      Sentences have changed.

      Reviewer #3 (Recommendations For The Authors):

      (1) Growing MEF-MCCs on micropatterns has successfully mimicked the dynamics of centriole amplification in brain MCCs, allowing the authors to study the spatial origin of procentrioles. Since this is a powerful system, a more quantitative description of the system will be informative and beneficial for future studies. For example: What is the efficiency of this system? Do the cilia that form in MEF-MCCs motile?

      The system of MEF-MCCs has been described in a previous paper from the Kintner lab. It seems that growing the MEF-MCCs on micropatterns did not ameliorate the ciliation which is partial, probably due to the absence of an apico-basal polarity.

      (2) Figure 2: The analogy drawn by the authors between DEUP1 oscillatory dynamics and centriolar satellites is intriguing. In early amplifying cells within the cloud, do these DEUP1 structures co-localize with the satellite marker PCM1?

      We have added immuno stainings of PCM1 in mRuby-DEUP1 / CEN2-GFP cells in Fig. Supplementary 2E. Within the centrosomal cloud, DEUP1 colocalizes with PCM1. Interestingly, this PCM1 concentration at the centrosome is dependent, at least in part, on dyneins. Then, PCM1 can localize around the deuterosomes, but it is never colocalized with deuterosomes (not shown). This is also showed by immuno-EM in Zhao et al., 2019. Although it was shown that PCM1 is a proximity interactor of DEUP1 (called ccdc67 at that time) by Firat-Karalar et al., 2014., absence of PCM1 staining on deuterosomes does not favor the hypothesis of PCM1 and DEUP1 being part of the same entities. One could hypothesizes that DEUP1 is transcribed locally within the satellites, explaining the colocalization of the 2 proteins and the + BioID results, and then form PCM1negative deuterosomes.

      (3) The authors propose a physical link between deuterosomes and centrosomes based on their oscillatory behavior. How are the oscillatory dynamics of DEUP1 affected by nocodazole treatment or inhibition of microtubule motors (i.e ciliobrevin treatment)?

      These oscillations are inhibited by nocodazole (Fig. 4D). They are also inhibited by dynapyrazole (Fig. 4D). We never succeeded in having a nice disruption of the Golgi apparatus with ciliobrevin and therefore we did not used it.

      (4) In addition to nocodazole treatment, it would be important to determine the consequences of microtubule stabilization by taxol and inhibition of microtubule motors during critical stages of centriole amplification where microtubules are reported to play a role for the first time in this manuscript. Another interesting area of investigation will be to study the extent to which microtubule PTMs contribute to these processes.

      We now blocks dyneins during the different stages of amplification. The results are in main and associated Fig. 4, 5, 7, 8. The role of microtubule PTM, is not in the scope of this manuscript.

      (5) Describing microtubule dynamics along with Centrin/DEUP1 dynamics will be informative in assessing whether these structures associate and/or move along microtubules? Have the authors performed their imaging experiments with SIR tubulin?

      Yes, we have tried hard! But we have encountered different obstacles:

      3-color video microscopy is phototoxic,

      siRTubulin is bleaching very rapidly

      The density of microtubules in MCC makes the observation hardly informative

      (6) Figure 5: The role of PLK1 in centriole-centrosome conversion and generation of multiple MTOCs can be tested with a PLK1 inhibitor for further confirmation.

      We have also tried but inhibiting Plk1 blocks the A-to-G and G-to-D transitions so it was not possible to uncouple the role of Plk1 in stage transitions versus centriole maturation.

      (7) Figure 6: The tight co-localization of nuclear pore proteins with centrioles poses questions about the role of nuclear pore proteins or other nuclear proteins that are associated with centrioles during centriole disengagement and migration. Considering the existing literature on centrosome-nucleus attachments, can there be a way to test this question within the scope of this manuscript?

      We have tried to deplete Nup133 but it’s killing the cells. Our additional experiments now show that the nuclear migration of centrioles during G-stage is dynein dependent, reinforcing the parallel with centrosome migration in prophase. We also added results from our scRNA sequencing (Fig. 5 Supplementary 1) showing that some key players of centriole migration to the nuclear membrane are conserved in the MCC cell cycle variant, and expressed with a comparable dynamics as to the canonical cell cycle.

      (8) Figure 8: Manually tracking a subset of migrating centrioles to define their dynamics during centriole migration and docking provides valuable analysis for determining the molecular mechanism of these processes. In addition to microtubules, does actin contribute to this process? Since centrioles eventually migrate to the apical side in nocodazole-treated cells, there should be other molecular players involved in this process.

      We did block actin polymerization but we found that the different stages were affected and that it would be better to dedicate a whole manuscript on the role of actin during each stage of amplification. We discuss the migration mechanism, and the putative role of actin, in the discussion.

      (9) The legends for Supplementary Figures 1 and 2 in Figure 3 are mixed and need correction.

      Figures have been remodelled.

      (10) In Figure 3P, the term "PLK4+" is labeled in bright green, which is not clearly visible. It maybe beneficial to change the color of this label for better visibility.

      We have tried to correct this.

      (11) Figure 6F quantifies "% tethered flowers" on the nuclear membrane. When quantifying, is the3D localization of DEUP1 flowers in both DMSO- and Noc-treated cells considered? A flower may appear to be on the nucleus in 2D, but it could be detached from the membrane in a 3D view.

      The quantifications are done in 3D. However, flowers that are below or above the nucleus are not quantified since the space is confined and the resolution in z to small to see whether they are connected or not. This is now precised in the legend.

      Before the editors proceed with an updated assessment, they've requested that we pass on some of the comments that have arisen as part of the evaluation of your revised manuscript. They feel that these concerns should be addressed before we proceed with issuing a formal assessment and publishing the revised Reviewed Preprint:

      We thank the reviewers and the editors for the corrections and insighfull comments. We apologize for our delayed answer and hope our corrections in the main text and some of the figures will give them satisfaction.

      The revised manuscript is greatly improved with nice new data regarding the role of microtubules. It also has changed quite a bit including the title. The new focus is on the cell and centriole cycle variants in MCC. While this helped to focus the study, there remains an important issue related to the interpretation of the data and the proposed 2-in-1 cycle model. Before providing the final updated assessment, we ask you to address the following points (which were raised already in the first round of review): The manuscript still contains statements that are not aligned with published work and the current view in the field regarding the timing of events during canonical centriole biogenesis. These timings are in conflict with your model that 2 centriole cycles are "superposed" in the MCC cell cycle variant, as currently presented. An alternative straightforward interpretation would be that multiciliogenesis uses an accelerated centriole duplication cycle where key steps occur concomitantly or in short succession instead of being separated by mitotic divisions as in the canonical cycle.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we think we have proposed. When correcting our confusions as regard to centriole-to-centrosome conversion (as explained below) and putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (corrected Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a superposition of events that; although driven by the same molecular machinery, are normally occuring in two consecutive cell cycle. We explain ourself briefly in two paragraphs, before answering point by point to the questions of the reviewers.

      As regard to centriole-to-centrosome conversion:

      We thank the reviewer for pointing out that we used “MTOC conversion” for what is normally called “centrosome maturation”. We have removed the term “centriole-to-centrosome conversion” during the first round of revision but we now realize that “MTOC conversion” leads to the same misinterpretation as regard to the literature on centriole duplication.

      The reviewer asks us to refer to the work of the Tsou lab (Wang 2011, reference now added in the manuscript) showing that daughter centrioles are “modified” (e.g. recruit PCM, become competent for MT nucleation and duplication) during late M/early G1. This “centriole-to-centrosome conversion” can’t occur for our procentrioles at this stage since they are not even born during the mitosis that precedes MCC differentiation. Also, in our cells, such modification does not include the capacity to become competent for duplication since we know that procentrioles become basal bodies without making any round of duplication (Al Jord et al., 2014).

      Also, we have not done the experiments to tackle the question on when our centriole become “modified-like”. What we can say is that during A-stage, they become progressively positive for PCM (Fig. 5 Supplementary 2) and a weak signal shows that some MT are seen emerging from them (Fig. 5 and Fig. 5 Supplementary 2, and see point by point answer).

      What we do see is that, at the A-to-G transition, they increase their PCM recruitment, show clear and strong MTOC ability (sometimes as strong as the centrosomal centrioles), and that this is associated with migration and separation of centrosome/deuterosomes around the nuclear membrane (Fig. 5). We therefore connect this to what occurs at the G2/M transition which is an increased recruitment of PCM protein, an increased ability to nucleate MT, associated with centrosome migration and separation at the nuclear membrane. Since this process in the canonical cell cycle is called “centrosome maturation”, we therefore should refer to this term in our study. However, centrioles in the MCC variants are not organized in centrosomes, so we now compare what we see to the “centrosome maturation” of the canonical cell cycle with an associated reference (Joukov et al., 2018), but name it “centriole maturation”.

      We have modified the text (track changes visibles) and the schemes (Fig. 5, Fig. 5 Supplementary 1 and 2, Fig. 9, Fig. 9 Supplementary S1; new versions uploaded) accordingly.

      As regard to 1.5 or 2 cell cycles

      Except for the “MTOC conversion” that we have now changed, as explained above, we think our work does suggest (depicted on Fig. 9) what the reviewer states for centriole duplication: “In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs. Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium)”.

      We feel that going from early S to a G1 phase, after 2 mitosis, is what one can call “2 cell cycles”. One of the paper that inspired us a lot when studying how the cell cycle machinery can drive centriole amplification in MCC is a paper from Jadranka Loncarek team (Kong et al., 2014) where they also state that “nascent centrioles gradually mature through 2 cell cycles”. Very interestingly, in this study they show that when they enhance Plk1 activation, they could erase centriole age and new procentrioles are able to recruit PCM and appendages within only 1 cell cycle, without mitotic progression, like what we see in MCC. We have added the reference in our discussion.

      Point by point answer

      (1) Original work on canonical centriole disengagement and centriole-to-centrosome conversion should be cited (e.g. PMID: 16862117, PMID: 21576395)

      As explained earlier, we used the wrong term since the begining. We do not speak about the centriole-to-centrosome (nor MTOC) conversion since we do not test when centriole modification (Wang et al., 2011) occurs in the MCC cell cycle variant. We know that PCNT is present on the procentrioles during A-stage (as shown in Fig. 5 Supplementary 2B), but we do not know when it is recruited (UExM did not work properly with this antibody). We quantify a weak MT staining in regrowth experiment during A-stage and see that procentrioles can be connected to MT in both brain MCC and MEFs (as shown in Fig. 5D, E for brain MCC and Fig. 5 Supplementary 2F for MEFs) , but we do not know when during A-stage they become competent for nucleation. We therefore did not speak about this process that we do not document. What we clearly document/quantify is the enhanced MT nucleation capacities at the A-to-G transition, concomitent with the nuclear migration (easily defined with Cen2-GFP or GT335 stainings) and that we compare to centrosome maturation occuring at the canonical G2/M transition.

      (2) The authors state in several places that canonical centriole formation and maturation takes two iterations of the canonical cell cycle. This is imprecise. Based on the above work and work by others, the broadly accepted view is that it takes 1.5 cell cycles. This difference matters for the final proposed model (see below). Reviewed e.g. here: PMID: 20869612; PMID: 30601682

      Our answer is in the preamble.

      (3) "Centriole maturation cycle superposes with centriole elongation cycle in the MCC cell cycle variant": Your description of the canonical cycle differs from the current view in the field. In the current view, centriole biogenesis starts in early S, elongation proceeds through G2/M and by early G1 it is complete. During M/early G1 centrioles disengage and newly formed daughters recruit PCM (centrosome conversion). All this occurs in 0.5 cycles. Then these centrioles go through another complete cell cycle and when they reach early G1 again they have acquired DAs and SDAs (total of 1.5 cell cycles). Key here is that biogenesis and disengagement/centrosome conversion are separated by the first mitosis (ensuring duplication occurs only once), and acquisition of DAs and SDAs is separated by another mitosis (ensuring that cells only form a single cilium).

      (4) Fig 5A, B and Fig. 9

      (a) Are 2 separate figures needed for the model? They seem redundant.

      We find it easier not to wait Fig. 9 to have the first part depicted.

      (b) The model shows loss of SAS6 throughout G1, but this already occurs during M/early G1

      Thanks. It was already ok in Fig. 9, we have modified for Fig. 5.

      The model shows "MTOC capacity/conversion" during S phase, but this occurs during early G1

      Thanks a lot, as explained earlier, we used the term MTOC conversion occurring in G1 for what is normally called centrosome maturation occurring in G2/M, as explained earlier. We do not speak anymore of MTOC conversion since we have not tackled this question (explained above). We have therefore removed MTOC conversion in the texts and the schemes and replaced it by “centrosome maturation” for the duplication cycle, and by “enhanced MT nucleation capacity” for the MCC cycle. To be clearer and schematize that procentrioles are competent for MT nucleation before G2/M or A/G transitions, we have added some MT nucleated from G1 procentrioles during the canonical cycle, and from late A-stage procentrioles during the MCC cycle.

      The model shows disengagement only in the second M phase, but this occurs already at the first M phase, directly following centriole biogenesis, right before centosome conversion.

      This is a big edition error in both Fig. 5 and 9. Of course the daughter centriole disengage during the first M-phase. This has been changed. Thanks a lot for spotting it. This, however does not contradict the hypothesis of superposition.

      We also added the acquisition of distal appendage which was written in Fig. 5 but not in Fig.

      9 for duplication during the second M-phase.

      When the correct timings are incorporated in the figure, the proposed superposition of two cycles is not an accurate description of the events. Instead, your data seem consistent with a model where MCC incorporates all steps in one cell cycle variant that lacks mitoses, so that disengagement and MTOC conversion occur together with centriole elongation, followed immediately by acquisition of DAs and SDAs.

      We do agree with the acceleration of all steps into only one cycle, this is actually what we tried to propose. When putting the events in a scheme, this reveals that the events of the two canonical cycles nicely superpose, both in term of molecular composition and dynamics (Fig. 9). We therefore maintain that the null hypothesis is that the acceleration is done through a super opposition of events that; although driven by the same molecular machinery, are normally occurring in two consecutive cell cycle. This is notably consistent with the findings of Kong et al., 2014 cited previously.

      (5) While all reviewers felt that there was no need to introduce the new term "nest", they leave it to the authors to keep it. However, the authors may want to consider that the term is still not introduced and explained properly, which may confuse readers. For example, while this section reads like an introduction to the term: "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC", the term is already used two times before without explanation. The first mentioning is at the beginning of the results section and is followed by citations, which gives the impression that these studies describe the nest, which is not the case.

      The first mention of “nest” is in the end of introduction resuming the findings of the paper where the term is in the following context: “we found that centriole amplification emerges in a pericentrosomal “nest” concentrating core centriole/deuterosome elements”. We looked at nest definition in the Collins Dictionnary : “a structure or other place where creatures, esp. birds, give birth or leave their eggs to develop”, we felt this was clear. We added quotation marks around the term nest.

      Then, the result section opens with this sentence: “The origin of amplified centrioles in MCC remains controversial. Some live imaging experiments and electron microscopy suggest that the centrosome could constitute a nest for centriole and deuterosome biogenesis (Al Jord et al., 2014; Kalnins et al., 1972; Mori et al., 2017), but others have proposed that procentriole-loaded deuterosomes emerge independently from the centrosome location, all over the cytoplasm (Nanjundappa et al., 2019; Sorokin, 1968; Zhao et al., 2013, 2019).”. Here, the term nest is again used as a place of birth for centrioles and deuterosomes which is what is actually proposed in these papers. First, Kalnins el al., in 1969 (we made an error on the reference date, this has been changed), resume in their abstract “This observation suggests that all of the clusters may form initially in close association with the diplosomal centrioles”. Then, not to mention Al Jord 2014 which comes from our lab, the title of Mori et al. is “Cytoplasmic E2f4 forms organizing centres for initiation of centriole amplification during multiciliogenesis”, and in the paper, they show that E2F4 accumulates at the centrosome. This is now also proposed by collaborators for MCIDAS (Lu et al., 2025). We feel that these references, which are often omitted, are appropriated at this location.

      Then we continue with: “To test whether microtubules drive the organization of a centrosomal nest from which procentrioles emerge”, which keeps the notion of the place of birth.

      Then the title "Correlative DEUP1 live-imaging and EM highlights the existence of a pericentrosomal "nest" in brain MCC" arrives. In this section we first speak about a pericentriosomal cloud on which we zoom in using CLEM, to then conclude at the end of the section “Altogether live imaging mRuby-DEUP1/CEN2-GFP during early A-stage suggests that core deuterosome and centriole components are concentrated in a primordial cloud around the centrosome, which constitutes a nest where centrioles and deuterosomes concomitantly form before they move away from the centrosomal region (Fig. 2F)”.

      Finally, we begin the discussion section regarding the nest by: “We named this transitory compartment a “nest” since deuterosomes and procentrioles emerge specifically in this region and grow while moving away from it.”

      During the first revision, we tried to make it clearer. If this is still not the case after and the reviewer has another proposition of definitions/phrasing, we will be glad to consider it.

      As replied to the other reviewer, the term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the transitory region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      The following comments from Reviewer #3 may also provide further context regarding the editors' remaining concerns:

      The authors have done an excellent job addressing the points I raised overall, and the revision is substantially improved in focus and clarity. That said, some concerns raised by other reviewers, particularly regarding terminology and statistical analysis, could have been addressed more fully. One issue remains insufficiently resolved. Several quantitative analyses (for example Fig. 5C and 5E) still appear to rely on pooled single-event measurements collected across three independent experiments. This approach can overstate statistical significance. The authors indicate in their rebuttal that they use chi-square tests to compare proportions and to justify pooling across replicates. However, I am not convinced this addresses the issue for the intensity-based and single event distributions shown in the panels specified above. I recommend that these key analyses be represented with biological replicates shown explicitly (superplot-style, with replicates distinguished).

      Our reply was for the comparison of proportions and not the intensity-based and single event distributions shown in the panels Fig. 5C and Fig. 5E. We have now changed our plots to represent biological replicates explicitly (superplot-style, with replicates distinguished). As for the statistical analysis: we evaluated differences in marker intensity between A-stage and G-stage samples using a linear regression model, with stages as the main effect and replicate as a fixed covariate, to account for batch variation. Statistical significance was assessed using Type II ANOVA.

      Separately, I continue to feel that some newly introduced terminology (for example, the "nest") may not be necessary at this stage. It may be sufficient to describe these structures and focus on their spatiotemporal behavior, composition, and measurable features, rather than assigning new names. Having read the authors' response, I understand that they would like to retain this terminology, which is acceptable; however, it may not be readily adopted by the field.

      The term “nest” does not need to be retained as a new terminology. It is just a way for us to identify the region and to best define one of its function/characteristic which is to host the birth of new deuterosomes and centrioles.

      Minor correction (remove "in MCCs" part from the following sentence):

      In MCC, PCM1 depletion alters deuterosome formation and centriole production in brain and airway MCC (Hall et al., 2023; Zhao et al., 2021).

      Done

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      This work convincingly shows that, rather than gradually "evolving" throughout interphase, global chromatin architecture undergoes unexpectedly sharp remodeling at G1-S (and to a lesser extent, S-G2) transitions. By applying "standard" Hi-C analyses on carefully sorted cells, the authors provide an excellent temporal view of how global chromatin architecture is changed throughout the cell cycle. They show a surprisingly abrupt increase in compartmentation strength (particularly interactions between the "active" A compartments) at G1-S transition, which is slightly weakened at S-G2 transition. Follow-up experiments show convincingly that the compartment "maturation" does not require the DNA synthesis accompanying S phase per se, but the authors have not identified the responsible factors (work for future publications). The possible biological ramifications of these architectural changes (setting up potential replication "factories", and/or facilitating transcription-replication conflict resolution, both more pertinent for the active A compartments, which are most affected) have been well discussed in the article, but still remain speculative at this stage.

      We thank Reviewer #1 for their positive and constructive assessment of our work, and we agree that the questions of responsible factors and biological ramifications are important directions for future studies.

      My major criticism of this article is aimed more at the state of the field in general, rather than this specific article, but it should be discussed to give a more balanced view: what actually is a chromatin compartment? Chromosomal tracing and live tracking experiments have shown that the majority of "structures" identified from Hi-C experiments are statistical phenomena, with even "strong" interactions only being infrequent and transient. A-B compartments are "built up" from multiple very low-frequency "interactions", so ascribing causal effects for genome functions is even tougher. As a result, I have very little confidence in the results of the authors' polymer simulations and their inferred "peninsula" A compartment structures without any other supporting experimental data.

      We thank the reviewer for raising this important conceptual point. This issue extends beyond the scope of the present study but reflects an important ongoing discussion in the 3D genome field regarding the biological interpretation of chromatin compartments.

      We agree that Hi-C interactions should not be interpreted as stable pairwise contacts present in every cell. A growing body of evidence from chromatin tracing and live-cell imaging studies has demonstrated that many chromatin interactions identified by Hi-C are probabilistic and dynamic, with substantial cell-to-cell variability. Relatively speaking, however, A/B compartment organization represents a robust population-level property of genome organization that is highly reproducible across biological replicates and closely correlates with multiple independent genomic features. In particular, replication timing (RT) correlates very well with A/B compartment organization, with early and late RT domains corresponding to A and B compartment domains, respectively.

      Furthermore, single-cell DNA replication sequencing (scRepli-seq) analyses have revealed remarkably low cell-to-cell variability in RT, suggesting that RT profiles and A/B compartment organization reflect biologically meaningful and relatively stable features of nuclear architecture rather than purely statistical artifacts. Thus, while individual chromatin contacts may be transient and probabilistic, the megabase-scale compartment organization inferred from them appears sufficiently reproducible to support reproducible RT programs and other genome functions. Additional support comes from decades of work on DNA replication demonstrating that spatiotemporal replication patterns, visualized as replication foci following short EdU pulses, are remarkably reproducible between individual cells throughout S-phase progression. These patterns reveal clear spatial segregation between early-replicating A-compartment regions and late-replicating B-compartment regions even at the single-cell level.

      To directly address the reviewer’s concern that A/B compartment organization might represent only an ensemble-level statistical phenomenon without biological relevance at the single-cell level, we performed L1/B1-EdU DNA FISH on asynchronous mESCs and MC12 embryonic carcinoma cells. L1 elements are enriched in B compartment domains, while B1 elements are enriched in A compartment domains, allowing visualization of compartment segregation in individual nuclei across the cell cycle. This single-cell analysis confirmed our Hi-C findings: compartment segregation increased from G1 to early S, remained elevated throughout S phase with reduced cell-to-cell variability, and then weakened in G2. Thus, compartment segregation is detectable in single cells, and the temporal dynamics of compartment maturation identified by population Hi-C were independently recapitulated at single-cell resolution. We have added a new Results section describing these findings titled “Stepwise A/B compartment reorganization during interphase is conserved at single-cell resolution”, including new Figure panels 2D–H and Figure S5.

      Regarding the polymer simulations, we agree that these models should be interpreted with caution. We do not view them as direct representations of individual nuclei, but rather as heuristic models that help visualize structural trends present in the Hi-C data. To make this point explicit, we have added the following statement to the revised manuscript: “We note that these models are derived from population-averaged Hi-C data and should therefore be interpreted as a heuristic framework for understanding A/B compartment dynamics, rather than as definitive representations of individual nuclei.”

      That said, we did try to provide orthogonal experimental support for the "A peninsula" model by performing DNA FISH. In brief, we measured distances between probe pairs spanning two A domains on chromosomes 2 and 15 across different cell-cycle stages. We observed significant increases in inter-probe distances from G1 to early/mid S, with the most pronounced changes involving the central probes (i.e., probes located near the domain center), consistent with physical extension of the A domain during S phase. While these data do not prove the exact geometry depicted by the model, these findings provide independent experimental support for the peninsula model as a simplified but biologically grounded interpretation of the Hi-C data. These results are described in the Results section titled “A-compartment consolidation during S-phase involves enhanced long-range contacts and structural reorganization” and are presented in new Figure panels 5D–F and Figure S12.

      We thank the reviewer again for raising this important conceptual issue, which prompted us to better clarify both the biological interpretation and the limitations of our analyses.

      Specific minor points:

      (1) A better explanation for how Figure 1E was generated is required, because this figure could be very misleading. Figure 1F and all other cis-decay plots (and the Hi-C maps themselves) show that the strongest interactions are always at smaller genomic separations, so why should there be more "heat" at the megabase ranges in Figure 1E?

      We appreciate the reviewer's observation. The apparent discrepancy is simply due to the fact that the decay plot (Fig. 1E in the original submission, now Fig. S2C) does not include the shortest-range interactions. The lowest distance plotted is 25 kb, following the method originally described in Nagano et al. (Nature, 2017), which we used as a reference. The shortest-range interactions (below 25 kb) are indeed the most enriched, as seen on the diagonal of the Hi-C maps (Fig. 2A) and in the standard cis-decay plot (Fig. 1F in the original submission, now Fig. S2F). With the 25 kb cutoff in place, the "heat" observed at megabase distances (specifically 12–50 Mb) in early/mid G1 corresponds to the dark, non‑specific band around the diagonal visible in the Hi-C maps at the same time points. This is also reflected in the cis-decay plot (Fig. S2F), where distances in that range appear above the expected curve (a "bump" rather than a linear decay).

      To avoid confusion, we have updated the figure legend accordingly (Fig. S2C): “(C) Contact decay profiles for all cell cycle phases, plotted from 25 kb to 50 Mb, illustrating a continuum of cis-interactions and a progressive shift from long-range (> 12 Mb) to short-range (< 1 Mb) interactions during the G1-to-S phase transition.”

      We hope this explanation clarifies the figure.

      (2) An ultra-high-resolution Hi-C study (Harris et al., Nat Commun, 2023) identified very small A and B compartments, including distinctions between gene promoters and gene bodies, raising further questions as to what the nature of a compartment really is beyond a statistical phenomenon. It is unreasonable to expect the authors to generate maps as deep as this prior study, but how much do their conclusions change according to the resolution of their compartment calling? The authors should include a balanced discussion on the "meaning" of A/B compartments.

      We thank the reviewer for highlighting recent ultra-high-resolution work, such as Harris et al. (Nat Commun, 2023), which reveals compartment-like features at much finer genomic scales. We agree that these findings raise important questions regarding the scale-dependence and interpretation of A/B compartmentalization.

      In our study, we specifically focus on coarse-grained compartment organization, analyzed across multiple resolutions (from ~1 Mb to sub‑megabase scales). Importantly, the key conclusions, including the abrupt strengthening of compartmentalization at the G1/S transition, are robust across these resolutions.

      We also note that fine-scale compartment-like features likely operate under different rules than larger-scale compartments. Recent evidence suggests that these "micro‑compartments" are more dynamic and transient (Harris et al., Nat Commun, 2023; Goel et al., Nat Struct Mol Biol, 2025), whereas the large-scale compartments analyzed here capture more stable, global segregation patterns. Understanding how these two regimes relate to one another remains an important open question.

      We have added the following statement in the Discussion acknowledging the scale-dependent nature of compartmentalization: “At the same time, recent ultra-high-resolution Hi-C studies [36,37] have revealed compartment-like features at much finer genomic scales, emphasizing that A/B compartmentalization is, to some extent, inherently scale-dependent. Understanding how these fine-scale, often transient micro-compartments relate to the more stable, large-scale segregation patterns described here will be an important direction for future studies.”

      Reviewer #2 (Public review):

      Summary:

      This manuscript by Choubani et al presents a technically strong analysis of A/B compartment dynamics across interphase using cell-cycle-resolved Hi-C. By combining the elegant Fucci-based staging system with in situ Hi-C, the authors achieve unusually fine temporal resolution across G1, S, and G2, particularly within the short G1 phase of mESCs. The central finding that A/B compartment strength increases abruptly at the G1/S transition, stabilizes during S phase, and subsequently weakens toward G2 challenges the prevailing view that compartmentalization strengthens monotonically throughout interphase. The authors further propose that this "compartment maturation" is triggered by S-phase entry but occurs independently of active DNA synthesis, and that it involves a consolidation and large-scale reorganization of A-compartment domains.

      Strengths:

      Overall, this is a thoughtfully executed study that will be of broad interest to the 3D genome community. The data are of high quality, and the analyses are extensive, albeit not completely novel. In particular, previous work (Nagano et al 2017 and Zhang et al 2019) has shown that compartments are re-established after mitosis and strengthened during early interphase, and single-cell Hi-C studies have reported changes in compartment association across S phase. In particular, Nagano et al show that DNA replication correlates with a build-up of compartments, similar to what is presented here, with the authors' conclusion that compartment strength peaks in early S. The idea that it weakens toward G2, rather than continuing to strengthen, appears to be novel and differs from the prevailing framing in the literature.

      We thank Reviewer #2 for their thoughtful assessment and critique. We address their specific concerns below.

      Weaknesses:

      That said, several aspects of the conceptual framing and interpretation would also benefit from further clarification, and the mechanistic interpretation of the reported compartment dynamics requires more careful positioning relative to established models of genome organization. Specific concerns are outlined below:

      (1) One of the major conclusions of the study is that compartment maturation does not require ongoing DNA replication. However, the interpretation would benefit from more precise wording. Thymidine arrest still permits licensing, replisome assembly, and other S-phase-associated chromatin changes upstream of bulk DNA synthesis. Therefore, their data, as presented, demonstrate independence from DNA synthesis per se, but not necessarily from the broader replication program. Please clarify this distinction in the text and interpretations throughout the manuscript.

      We thank the reviewer for this important distinction. We agree with their point and have never claimed that compartment maturation is independent of the broader replication program. That is why we carefully used the term "active DNA synthesis" rather than "replication" throughout the manuscript.

      However, we acknowledge that one sentence in the text was ambiguous. The original sentence read: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a pre-replicative state where replication had not yet initiated, although cell-cycle markers indicated entry into S-phase.”

      We have now revised it to: “These results confirm that the cell population was successfully synchronized at the G1/S boundary, representing a state where the replication program (including origin licensing, replisome assembly, and helicase activation) has been initiated, as indicated by cell-cycle markers, but ongoing DNA synthesis (elongation) is blocked. ”

      This clarifies that compartment maturation is independent of active DNA synthesis (elongation) but not necessarily independent of upstream replication-associated processes. The change has been made in the manuscript.

      (2) A major conceptual issue that is not addressed at all is the well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartmentalization. Numerous studies have shown that loss of cohesin or reduced loop extrusion leads to stronger compartment signals, whereas increased cohesin residence or enhanced extrusion weakens compartmentalization. Given this framework, an obvious alternative explanation for the authors' observations is that the abrupt increase in compartment strength at G1/S, and its decline toward G2, could reflect cell-cycle-dependent modulation of cohesin activity rather than a compartment-intrinsic "maturation" program.

      The manuscript does not explicitly consider this possibility, nor does it examine loop extrusion-related features (such as loop strength, insulation, or stripe patterns) across the same cell-cycle stages. Without discussing or analyzing this widely accepted model, it is difficult to distinguish whether the reported compartment dynamics represent a novel architectural mechanism or an indirect consequence of known changes in extrusion behavior during the cell cycle. I strongly encourage the authors to analyze their data to determine if they observe anti-correlated loop changes at the same time they observe compartment changes. Ideally, the authors would remove loop extrusion during interphase using well-established cohesin degrons available in mESCs and determine if the relative differences in compartment dynamics persist.

      We thank the reviewer for raising this interesting point. We agree that there is a well-established anti-correlation between cohesin-mediated loop extrusion and A/B compartment strength in the literature.

      To test whether cell cycle compartment dynamics, particularly compartment maturation at the G1/S transition, could be explained by changes in loop extrusion, we analyzed insulation at RAD21/CTCF sites (mESC data from Hansen et al., eLife, 2017) across the cell cycle. During normal cycling, we indeed observed an anti-correlation: insulation dropped as compartment strength increased at the G1/S transition. However, in G1/S-arrested cells, insulation did not drop compared to late G1 (it even slightly increased) even though compartment maturation still occurred, indicating that the two processes can be uncoupled. This is consistent with other studies showing that loop extrusion and compartment dynamics are driven by independent mechanisms (Nora et al., Cell, 2017; Zhang et al., Nat Commun, 2021), although we cannot fully rule out some contribution from loop extrusion dynamics without direct cohesin degron experiments.

      We have added a new Results section describing these findings titled “Compartment maturation is independent of cohesin-mediated loop extrusion”, including new Figure panels 3H, I, and Figure S7.

      (3) The proposed "peninsula-like" A-domain structures are inferred from ensemble Hi-C data and polymer modeling, rather than directly observed physical conformations. That is, single-cell imaging data clearly have shown that Hi-C (especially ensemble Hi-C) cannot uniquely specify physical conformations and that different underlying structures can produce similar contact patterns. The "peninsula" language, as written, risks being interpreted as a literal structural model rather than a conceptual visualization. Instead of risking this as just another nuanced Hi-C feature in the field, the authors could strengthen the manuscript by either (i) explicitly framing the peninsula model as a heuristic description of contact redistribution rather than a definitive physical architecture, or (ii) discussing alternative structural scenarios that could give rise to similar Hi-C patterns. Clarifying this distinction would improve the rigor and help readers better understand what aspects of A-compartment consolidation are directly supported by the data versus model-based extrapolations. For example, it would be useful to clarify whether the observed increase in long-range A-A contacts reflects spatial extension of internal A regions, changes in loop extrusion dynamics, increased compartment mixing within the A state, or population-averaged heterogeneity across alleles.

      We thank the reviewer for this important clarification. We agree that the "peninsula" model should be framed as a heuristic description. As detailed in our response to Reviewer #1 (see above), we have added a disclaimer to the manuscript and provided orthogonal DNA FISH support for physical extension of A-domains during S phase. We have also ensured that the language emphasizes the conceptual nature of the model.

      (4) The extension of the analysis to additional cell types using HiRES single-cell data is a valuable addition and supports the idea that compartment maturation is not unique to mESCs. However, the limitations of these data, in particular, the limited phase resolution, in addition to the pseudo-bulk aggregation and variable coverage, should be emphasized more clearly in the main text. Framing these results as evidence for conservation in principle, rather than definitive proof of identical dynamics across tissues, would be a more appropriate framing.

      We agree with the reviewer. We have already explicitly acknowledged the limited temporal resolution and variable coverage of the HiRES dataset in the main text. To better reflect its supporting role, we have moved the HiRES figure (previously Fig. 4) to Fig. S10 and merged the corresponding results section with the previous one titled: “Formation of a consolidated A compartment in S-phase”.

      We have also revised the language to avoid overstatement. The original conclusion read: “Together, these findings strongly indicate that compartment maturation and the accompanying A compartment consolidation represent a robust and universally observed feature across different developmental contexts.”

      This has been changed to: “Together, these findings support the notion that compartment maturation and the accompanying A-compartment consolidation are not unique to mESCs and may represent a broadly conserved feature of mammalian chromatin organization.”

      Similarly, the abstract has been adjusted from: “Moreover, compartment maturation was not limited to mESCs but was also observed across different developmental contexts in mice.” to: “Moreover, compartment maturation was not limited to mESCs but was also evident across different developmental contexts in mice.”

      These changes frame the results as evidence for conservation in principle rather than definitive proof of identical dynamics across tissues.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Please address the minor points in the public review.

      In addition, on page 7, line 285: "In contrast, interactions showed minimal change across all distances though interphase". Do the authors mean "In contrast, B-B interactions..."?

      We thank the reviewer for catching this. The sentence has been corrected.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Lu and colleagues demonstrates convincingly that PRRT2 interacts with brain voltage-gated sodium channels to enhance slow inactivation in vitro and in vivo. The work is interesting and rigorously conducted. The relevance to normal physiology and disease pathophysiology (e.g., PRRT2-related genetic neurodevelopmental disorders) seems high. Some simple additional experiments could elevate the impact and make the study more complete.

      Strengths:

      Experiments are conducted rigorously, including experimenter blinding and appropriate controls. Data presentation is excellent and logical. The paper is well written for a general scientific audience.

      We thank the reviewer for these positive comments and for the thoughtful evaluation of our work.

      Weaknesses:

      There are a few missing experiments and one place where data are over-interpreted.

      (1) An in vitro study of Nav1.6 is conspicuously absent. In addition to being a major brain Na channel, Nav1.6 is predominant in cerebellar Purkinje neurons, which the authors note lack PRRT2 expression. They speculate that the absence of PRRT2 in these neurons facilitates the high firing rate. This hypothesis would be strengthened if PRRT2 also enhanced slow inactivation of Nav1.6. If a stable Nav1.6 cell were not available, then simple transient co-transfection experiments would suffice.

      We thank the reviewer for raising this point. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms.

      We have now performed new heterologous expression experiments to test whether PRRT2 modulates Nav1.6 slow inactivation. Consistent with our findings for other Nav isoforms, PRRT2 significantly enhances the slow inactivation of Nav1.6. We have incorporated these data into the revised Results and Figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) To further demonstrate the physiological impact of enhanced slow inactivation, the authors should consider a simple experiment in the stable cell line experiments (Figure 1) to test pulse frequency dependence of peak Na current. One would predict that PRRT2 expression will potentiate 'run down' of the channels, and this finding would be complementary to the biophysical data.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we performed a pulse-train protocol in the stable Nav1.2 cell line and quantified the use-dependent attenuation (“run-down”) of peak sodium current across successive depolarizations (Figure 1-figure supplement 1C). Compared with control cells, PRRT2-expressing cells exhibited a larger decline in peak current during trains, indicating greater reduction in channel availability during repetitive depolarizations (Figure 1-figure supplement 1C). This pattern is consistent with our observations above showing that PRRT2 enhances Nav channel slow inactivation. These new data have been incorporated into the revised manuscript. Please refer to Page 5, Lines 133-140; Figure 1-figure supplement 1C.

      (3) The study of one K channel is limited, and the conclusion from these experiments represents an over-interpretation. I suggest removing these data unless many more K channels (ideally with measurable proxies for slow inactivation) were tested. These data do not contribute much to the story.

      We agree with the reviewer’s assessment. To avoid over-interpretation and to maintain focus on PRRT2-dependent regulation of Nav channel slow inactivation, we have removed the potassium channel dataset and the associated conclusions from the revised manuscript.

      (4) In Figure 2, the authors should confirm that protein is indeed expressed in cells expressing each truncated PRRT2 construct. Absent expression should be ruled out as an explanation for the enhancement of slow inactivation.

      We thank the reviewer’s concern regarding expression of the truncated PRRT2 constructs in the Nav1.2 stable cell line, particularly PRRT2(1-266), which shows little effect on slow inactivation of Nav1.2 channels. In the revised manuscript, we conducted western blot to verify expression of the PRRT2(1-266)-HA construct in the Nav1.2 stable cell line. We have added these results to the revised manuscript, please refer to Page 6, Lines 171-173; Figure 2-figure supplement 1A and B.

      Reviewer #2 (Public review):

      Summary:

      As a member of DspB subfamily, PRRT2 is primarily expressed in the nervous system and has been associated with various paroxysmal neurological disorders. Previous studies have shown that PRRT2 directly interacts with Nav1.2 and Nav1.6, modulating channel properties and neuronal excitability.

      In this study, Lu et al. reported that PRRT2 is a physiological regulator of Nav channel slow inactivation, promoting the development of Nav slow inactivation and impeding the recovery from slow inactivation. This effect can be replicated by the C-terminal region (256-346) of PRRT2, and is highly conserved across species from zebrafish, mouse, to human PRRT2. TRARG1 and TMEM233, the other two DspB family members, showed similar effects on Nav1.2 slow inactivation. Co-IP data confirms the interaction between Nav channels and PRRT2. Prrt2-mutant mice, which lack PRRT2 expression, require lower stimulation thresholds for evoking after-discharges when compared to WT mice.

      Strengths:

      (1) This study is well designed, and data support the conclusion that PRRT2 is a potent regulator of slow inactivation of Nav channels.

      (2) This study reveals similar effects on Nav1.2 slow inactivation by PRRT2, TMEM233, and TRARG1, indicating a common regulation of Nav channels by DspB family members (Supplemental Figure 2). A recent study has shown that TMEM233 is essential for ExTxA (a plant toxin)-mediated inhibition on fast inactivation of Nav channels; and PRRT2 and TRARG1 could replicate this effect (Jami S, et al. Nat Commun 2023). It is possible that all three DspB members regulate Nav channel properties through the same mechanism, and exploring molecules that target PRRT2/TRARG1/TMEM233 might be a novel strategy for developing new treatments of DspB-related neurological diseases.

      We thank the reviewer for careful evaluation and insightful suggestions.

      Weaknesses:

      (1) Previously, the authors have reported that PRRT2 reduces Nav1.2 current density and alters biophysical properties of both Nav1.2 and Nav1.6 channels, including enhanced steady-state inactivation, slower recovery, and stronger use-dependent inhibition (Lu B, et al. Cell Rep 2021, Fig 3 & S5). All those changes are expected to alter neuronal excitability and should be discussed.

      We thank the reviewer for this suggestion. Although the present study focuses on PRRT2-dependent regulation of slow inactivation, we agree that PRRT2 may influence excitability through additional Nav-dependent mechanisms, including reduced current density and shifts in the voltage dependence of channel inactivation (Fruscione et al., 2018; Lu et al., 2021; Valente et al., 2023). Notably, because PRRT2 facilitates entry of Nav channels into slow-inactivated states both from closed states and from open states during prolonged depolarization, some of these previously reported effects may partly reflect enhanced slow inactivation and the resulting reduction in Nav channel availability. We have expanded the Discussion to integrate these prior findings and to clarify that these additional PRRT2-dependent effects may converge to shape neuronal excitability. Please refer to Page 16, Lines 445-452.

      (2) In this study, the fast inactivation kinetics was examined by a single stimulus at 0 mV, which may not be sufficient for the conclusion. Inactivation kinetics at more voltage potentials should be added.

      We thank the reviewer for this helpful suggestion. In the revised manuscript, we expanded our analysis of Nav1.2 fast-inactivation kinetics to include a range of test potentials (-20, -10, 0, +10, +20 and +30 mV) in the presence and absence of PRRT2. These experiments showed that PRRT2 expression did not significantly affect Nav1.2 fast-inactivation kinetics under these conditions. We have incorporated these new results into the revised manuscript. Please refer to Page 4, Lines 100-103; Figure 1C.

      (3) It is a little surprising that there is no difference in Nav1.2 current density in axon-blebs between WT and Prrt2-mutant mice (Figure 7B). PRRT2 significantly shifts steady-state slow inactivation curve to hyperpolarizing direction, at -70 mV, nearly 70% of Nav1.2 channels are inactivated by slow inactivation in cells expressing PRRT2 when compared to less than 10% in cells expressing GFP (Figure supplement 1B); with a holding potential of -70 mV, I would expect that most of Nav channels are inactivated in axon-blebs from WT mice but not in axon-blebs from Prrt2-mutant mice, and therefore sodium current density should be different in Figure 7B, which was not. Any explanation?

      We thank the reviewer for raising this point. In our axonal bleb recordings, although the holding potential was -70 mV, sodium current density was measured after a hyperpolarizing pre-pulse to -110 mV, which was applied before the test depolarization to relieve inactivation as much as possible (as described in the Methods). Therefore, the current density measurement in Figure 7B reflects the available current after this recovery step, rather than the steady-state availability at -70 mV. The lack of a difference in Figure 7B does not contradict the PRRT2-dependent shift in steady-state slow inactivation. In the revised manuscript, we have clarified this point explicitly in the Results and figure legend to avoid confusion. Please refer to Page 10, Lines 294-295.

      (4) Besides Nav channels, PRRT2 has been shown to act on Cav2.1 channels as well as molecules involved in neurotransmitter release, which may also contribute to abnormal neuronal activity in Prrt2-mutant mice. These should be mentioned when discussing PRRT2's role in neuronal resilience.

      We thank the reviewer for this suggestion. In addition to the Nav-dependent mechanisms, previous studies have shown that PRRT2 also regulates synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018) and presynaptic surface expression of Cav2.1 channels (Ferrante et al., 2021). These effects are also expected to influence neurotransmitter release and, consequently, neuronal and network excitability. In the revised manuscript, we have expanded the Discussion to acknowledge that these additional PRRT2-dependent mechanisms may also contribute to cortical resilience. Please refer to Page 16, Lines 452-457.

      Reviewer #3 (Public review):

      This paper reveals that the neuronal protein PRRT2, previously known for its association with paroxysmal dyskinesia and infantile seizures, modulates the slow inactivation of voltage-gated sodium ion (Nav) channels, a gating process that limits excitability during prolonged activity. Using electrophysiology, molecular biology, and mouse models, the authors show that PRRT2 accelerates entry of Nav channels into the slow-inactivated state and slows their recovery, effectively dampening excessive excitability. The effect seems evolutionarily conserved, requires the C-terminal region of PRRT2, and is recapitulated in cortical neurons, where PRRT2 deficiency leads to hyper-responsiveness and reduced cortical resilience in vivo. These findings extend the functional repertoire of PRRT2, identifying it as a physiological brake on neuronal excitability. The work provides a mechanistic link between PRRT2 mutations and episodic neurological phenotypes.

      We thank the reviewer for this positive evaluation of our work and for the constructive comments.

      Comments:

      (1) The precise structural interface and the molecular basis of gating modulation remain inferred rather than demonstrated.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav channel interface—through approaches such as targeted mutagenesis, crosslinking, and structural determination—will be required to define the binding interface and establish the molecular basis of gating modulation. Please refer to Page 16, Lines 465-468.

      (2) The in vivo phenotype reflects a complex circuit outcome and does not isolate slow-inactivation defects per se.

      We agree with the reviewer. Impaired slow inactivation in Prrt2-mutant mice is one plausible contributor to reduced cortical resilience. PRRT2 has also been reported to regulate surface exposure of Nav and Cav2.1 channels (Ferrante et al., 2021), as well as neuronal synaptic vesicle cycling (Valente et al., 2016; Coleman et al., 2018; Tan et al., 2018). Each of these PRRT2-associated processes could influence cortical excitability in vivo. We have therefore expanded the Discussion to clarify that the cortical phenotype likely reflects the combined contribution of multiple PRRT2-dependent mechanisms, rather than an isolated defect in slow inactivation alone. Please refer to Page 16, Lines 446-458.

      (3) Expression of PRRT2 in muscle or heart is low, so the cross-isoform claims are likely of limited physiological significance.

      We thank the review for this comment regarding physiological relevance. In the revised manuscript, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Please refer to Page 14 and 15, Lines 414-416; Lines 429-430.

      (4) The mechanistic separation between the trafficking effect of PRRT2 and its gating effects is not clearly resolved.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. Previous studies in heterologous overexpression systems have shown that PRRT2 can influence Nav channel trafficking and surface expression, raising the possibility that the observed effects on slow inactivation regulation might be secondary to altered channel abundance or localization. However, slow inactivation develops on a timescale of tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer intervals (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). These distinct temporal profiles argue against trafficking as the primary basis for the effects of PRRT2 on Nav channel slow inactivation described here, although direct quantification of dynamic changes in Nav channel surface expression will be required to fully exclude such a contribution (Liu et al., 2022; Tyagi et al., 2025). We have incorporated this point into the Discussion section. Please refer to Pages 13, Lines 378-388.

      (5) Additional studies with Nav1.6 should be carried out.

      We thank the reviewer for this suggestion. We have performed experiments to directly examine the effects of PRRT2 on Nav1.6 slow inactivation and incorporated these new data into the revised Results and figures, please refer to Page 8, Lines 211-215; Figures 4E and J.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Suggestions for future experiments (not for this paper)

      (1) Exploit the lower protein expression in V5-PRRT2 mice to examine the effects of a hypomorphic allele.

      We thank the reviewer for this insightful suggestion. We note that the V5 epitope knock-in reduced PRRT2 protein expression, which may functionally resemble a hypomorphic allele. Accordingly, in addition to its utility for biochemical experiments (e.g., co-immunoprecipitation), this line could serve as a genetic tool to interrogate PRRT2 dose-dependent effects in vivo. We have added this point to the revised manuscript, please refer to Page 9, Lines 265-267.

      (2) Examine disease-causing PRRT2 mutations.

      We thank the reviewer for this constructive suggestion. Testing disease-associated PRRT2 variants for their ability to regulate Nav channel slow inactivation would be an important next step to strengthen the disease relevance of the mechanism proposed here. Moreover, identifying missense variants that selectively disrupt slow-inactivation regulation could help pinpoint residues that are critical for PRRT2-Nav functional coupling and thereby inform future structure-function studies. We plan to pursue this direction in follow-up work.

      (3) Investigate spreading depolarization in PRRT2-deficient mice.

      We thank the reviewer for this suggestion. Although we have shown that PRRT2 deficiency facilitates spreading depolarization in the cerebellum, whether PRRT2 exerts similar control over spreading depolarization susceptibility in the cerebral cortex remains to be determined. We plan to address this in an independent study and to test how cortical spreading depolarization relates to other PRRT2-associated neurological disorders.

      Reviewer #2 (Recommendations for the authors):

      This study is, in general, well executed, and the manuscript is well written. However, I do have some questions.

      (1) The authors' previous works have shown that PRRT2 regulates both Nav1.2 and Nav1.6, considering the wide expression Nav1.6 in CNS and its role in neuronal activity, what makes the authors not include Nav1.6 in this study?

      We thank the reviewer for raising this question. In our previous work, PRRT2 produced broadly similar effects on Nav1.2 and Nav1.6. Therefore, in the initial version of this study, we focused primarily on Nav1.2 as a representative neuronal Nav channel isoform and placed greater emphasis on testing whether PRRT2-dependent regulation of slow inactivation extends across additional Nav isoforms. In response to reviewers’ concern, we have now performed new experiments to directly examine the effect of PRRT2 on Nav1.6 slow inactivation. These results have been incorporated into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      (2) Please explain why you chose 0 mV rather than -70 mV (closer to membrane potential) in the slow inactivation protocol.

      We thank the reviewer for raising this question. Nav channels can enter into slow inactivation from both resting/closed states and activated/open states. In our steady-state slow-inactivation assays, we found that PRRT2 enhances Nav1.2 slow inactivation under both conditions (Figure 1-figure supplement 1A and B). In whole-cell recordings, Nav1.2 channels typically begin to activate at command voltages more depolarized than approximately -60 mV. Accordingly, a conditioning voltage of -70 mV predominantly probes entry into slow inactivation from closed states, whereas 0 mV drives channel activation and more effectively induces slow inactivation. We therefore chose 0 mV as the primary conditioning potential because it is widely used in conventional slow inactivation protocols and induces slow inactivation more robustly than conditioning voltages at -70 mV. We have added this explanation in Methods section of revised manuscript, please refer to Page 20, Lines 569-571.

      (3) The authors mentioned that the insertion of V5 markedly reduced the PRRT2 protein level; thus, Prrt2-V5 knock-in mice could be considered as PRRT2 knock-down mice. Is there any noticeable difference in phenotype between Prrt2-V5 knock-in mouse and Prrt2-mutant mouse? In other words, is PRRT2 knockdown sufficient to affect neuronal excitability, or is a complete PRRT2 ablation required?

      We thank the reviewer for raising this concern regarding the functional consequences of reduced PRRT2 expression in the Prrt2-V5 knock-in mice. Given that PRRT2 protein levels are markedly reduced in this line, and that cerebellar stimulation-induced dystonia is a characteristic phenotype of PRRT2 deficiency, we tested whether Prrt2-V5 knock-in mice also exhibit this phenotype. We found that electrical stimulation of the cerebellar cortex induced dystonia-like attacks in a subset of Prrt2-V5 knock-in mice. These dystonic behaviors resembled those previously observed in Prrt2-mutant mice, whereas no such behaviors were induced in wild-type mice (Figure 6-figure supplement 1). These findings indicate that a substantial reduction of PRRT2 expression (approximately 80%) is sufficient to impair neuronal function and elicit a disease-relevant phenotype in a subset of animals, supporting the interpretation that the V5 knock-in allele is hypomorphic. We have incorporated these results into the revised manuscript, please refer to Page 9, Lines 265-267; Figure 6-figure supplement 1.

      (4) In Discussion (Page 13, lines 358-361), the authors mentioned a putative interaction between PRRT2 and the Nav channel by modeling, while there is no related data. Please either add modeling data or remove those sentences.

      We thank the reviewer for this suggestion. To avoid over-interpretation, we have removed the statements regarding the AlphaFold-based interaction model from the revised manuscript. We agree that the interaction interface remains to be demonstrated experimentally, and we now discuss this point in the Limitations section. Please refer to Page 16, Lines 465-468.

      (5) Typo: Page 14, line 399, "TMEM232" should be "TMEM233".

      We thank the reviewer for pointing out this typo. We have corrected it in the revised manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Mechanistic depth: While the functional data show altered slow-inactivation kinetics, the mechanistic explanation remains superficial. The AlphaFold-based prediction of PRRT2 interaction with DIV-S3 is speculative. The authors should clarify their illustrative rather than evidential intent and avoid over-interpretation.

      We thank the reviewer for this comment. To avoid over-interpretation, we have removed the AlphaFold-based interaction prediction from the revised manuscript. We have also expanded the Limitations section to emphasize that direct structural and biochemical mapping of the PRRT2-Nav interface, including targeted mutagenesis, crosslinking, and structural determination, will be necessary to elucidate the molecular basis of this interaction and its effect on channel gating. Please refer to Page 16, Lines 465-468.

      (2) Separation of trafficking vs. gating effects: Previous studies showed PRRT2 influences Nav trafficking and surface expression. Here, surface expression changes are not systematically quantified. Such an analysis would strengthen the argument that gating effects are not secondary to altered channel abundance or localization.

      We thank the reviewer’s concern regarding the possible contribution of trafficking effects to PRRT2-dependent regulation of Nav channel slow inactivation. We agree that direct analysis of Nav channel surface localization during prolonged depolarization and hyperpolarization would provide stronger evidence to distinguish gating effects from trafficking-dependent mechanisms. However, such experiments are technically challenging in this context: conventional surface biotinylation assays do not provide the temporal resolution required for these rapid protocols, and live-cell imaging approaches to monitor dynamic changes in Nav channel surface expression during slow-inactivation paradigms have not yet been established in our laboratory.

      Although PRRT2 has been reported to regulate Nav channel surface expression in heterologous systems, we consider it unlikely that trafficking is the major determinant of the slow-inactivation effects described here. Slow-inactivation develops on a timescale ranging from tens of milliseconds to seconds, whereas detectable changes in Nav channel trafficking and surface abundance generally occur over much longer timescales (minutes to hours) (Freal et al., 2023; Higerd-Rusli et al., 2023). We have expanded the Discussion in a revised manuscript. Please refer to Pages 13, Lines 378-388.

      (3) Isoform generalization: Data on other Nav channel subtypes are presented as evidence of a conserved mechanism. However, given tissue-specific expression of PRRT2, these findings may be of limited in vivo relevance. At the very least, additional studies with Nav1.6 should be carried out.

      We thank the review for this suggestion. In response, we conducted new experiments to examine the effect of PRRT2 on Nav1.6 slow inactivation. These results show that PRRT2 promotes entry of Nav1.6 channels into slow-inactivated states and delays their recovery, consistent with its effects on the other Nav isoforms examined in this study. We have incorporated these new data into the revised manuscript. Please refer to Page 8, Lines 211-215; Figures 4E and J.

      Furthermore, we clarify that the cross-isoform analysis was intended to assess mechanistic generality at the channel level, rather than to imply equivalent physiological relevance across tissues. The functional consequence of PRRT2 depend on the Nav isoform composition and cellular context of each tissue. We also note that the broad isoform activity of the PRRT2 should be considered in any future attempt to manipulate PRRT2 function therapeutically. Pages 14 and 15, Lines 414-416 and 429-430.

      (4) In vivo functional link: The EEG after-discharge threshold assay suggests decreased cortical resilience, but causality between slow-inactivation impairment and hyperexcitability remains indirect. Complementary in vivo recordings would strengthen the physiological link.

      We thank the reviewer for this helpful suggestion. To further link impaired slow-inactivation to the hyperexcitability, we applied a repetitive stimulation protocol in corpus callosum slices, a white-matter region of brain enriched in both PRRT2 and Nav channels. During high-frequency stimulation (e.g., 20 Hz), the amplitude of the compound action potential progressively decreased over the course of the stimulus train. This phenomenon, often referred to as adaptation, reflects activity-dependent reduction in Nav channel availability (Fleidervish et al., 1996; Mickus et al., 1999; Kim et al., 2012). Compared with wild-type mice, Prrt2-mutant mice exhibited less adaptation during high-frequency stimulation, consistent with impaired slow inactivation during repetitive activity, which may contribute to hyperexcitability (Figure 7-figure supplement 2). We have added these results to the revised manuscript. Please refer to Pages 11, Lines 311-322; Figure 7-figure supplement 2.

      (5) Structural interaction: It remains unclear whether PRRT2 binds the α-subunit directly or through accessory proteins. Crosslinking or detergent-solubilization controls of different stringencies could clarify this.

      We thank the reviewer for raising this important issue. We agree that our co-immunoprecipitation data do not distinguish whether PRRT2 associates with the Nav channel α-subunit directly or through other components of the protein complex. To avoid over-interpretation, we have revised the relevant text in the manuscript to remove any implication of direct binding and now describe the result as an association between PRRT2 and Nav channels.

      We have also expanded the Limitations section to note that additional experiments, such as crosslinking and structural studies, will be required to define the interaction interface between PRRT2 and Nav channels. Please refer to Page 16, Lines 465-468.

      (6) Comparisons to other regulators: The paper positions PRRT2 as distinct from FHFs and β-subunits. The data support this, but the discussion could more critically assess whether PRRT2 acts by stabilizing a pore-based inactivated conformation, as suggested for other slow-inactivation modulators.

      We thank the reviewer for this insightful suggestion. At present, relatively few modulators have been characterized in detail with respect to their effects on Nav channel slow-inactivation kinetics. Moreover, even for compounds such as lacosamide, which has been proposed to act as a slow-inactivation modulator, the underlying mechanism remains under debate (Errington et al., 2008; Jo and Bean, 2017). Therefore, in the revised manuscript, we discussed the possible mechanism of PRRT2 in the context of current models of Nav channel slow inactivation.

      Previous studies suggest that entry into the slow-inactivated state involves at least two coupled processes: conformational changes in the voltage-sensing domains and structural rearrangements in the pore region, including the selectivity filter and intracellular activation gate (Catterall et al., 2024; Silva, 2014). During prolonged depolarization, voltage sensors become stabilized in the up-state, while the pore undergoes progressive rearrangements associated with slow inactivation (Balser et al., 1996; Vilin et al., 1999). Thus, mechanisms that further stabilize voltage sensors in the up-state and/or facilitate pore-based inactivated conformations could enhance slow inactivation.

      Within this framework, PRRT2 may enhance slow inactivation by facilitating one or both of these processes, although direct evidence is still lacking. We have incorporated this discussion in relative section of revised manuscript. Please refer to Page 14, Lines 389-404.

      Response references:

      Jo S, Bean BP. Lacosamide Inhibition of Nav1.7 Voltage-Gated Sodium Channels: Slow Binding to Fast-Inactivated States. Mol Pharmacol. 2017 Apr;91(4):277-286.

      Errington AC, Stöhr T, Heers C, Lees G. The investigational anticonvulsant lacosamide selectively enhances slow inactivation of voltage-gated sodium channels. Mol Pharmacol. 2008 Jan;73(1):157-69.

      (7) Behavioral/clinical link: Given the strong human genetics background of PRRT2 disorders, a brief analysis or reference to electrophysiological phenotypes in patient neurons would contextualize the cortical findings.

      We thank the reviewer for this suggestion. Previous studies showed that iPSC-derived excitatory neurons from a patient carrying a homozygous PRRT2 mutation exhibited increased sodium currents and neuronal hyperexcitability (Fruscione et al., 2018). Given that slow inactivation regulates Nav channel availability and thereby influences neuronal excitability, these electrophysiological abnormalities in patient-derived neurons may, at least in part, reflect impaired PRRT2-dependent regulation of Nav channel slow inactivation. We have added this point to the relative section of the revised manuscript. Please refer to Pages 15, Lines 432-437.

      Minor comments

      (1) Figures should include statistical sample sizes (n) and ideally overlay data points rather than only means {plus minus} SEM.

      We thank the reviewer for this suggestion. In the revised manuscript, we present both individual data points and mean ± SEM in the column graphs. For the line graphs, individual data points were not overlaid because of space and readability constraints, and these panels therefore display mean ± SEM only. Sample sizes for each group are provided in the corresponding figure legends.

      (2) The AlphaFold model should be provided as a supplementary figure with confidence scores indicated.

      We thank the reviewer for this suggestion. However, because the predicted Nav1.2-PRRT2 interaction interface has not yet been experimentally validated in our study, we chose to remove the AlphaFold-based model from the revised manuscript to avoid over-interpretation.

      (3) Clarify whether TTX sensitivity was verified in the axonal bleb preparation.

      We thank the reviewer for raising this point. We verified the identity of the sodium currents in the axonal bleb preparation by their sensitivity to TTX, and this information has now been added to Figure 7A in the revised manuscript. Please refer to Page 10, Line 290; Figure 7A.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      In this valuable study, the authors developed long-term imaging tools to simultaneously monitor the temporal and spatial dynamics of excitatory and inhibitory synapses and reported that excitatory and inhibitory synapses need to develop synergistically during synaptogenesis to maintain balance. While the analysis and quantification of the imaging data are incomplete, there is convincing evidence that the developed tools are feasible. If these tools can function stably in vivo, their applications will be much broader.

      We have completely overhauled our analysis and quantification methods and generated custom-made drift correction and tracking pipelines. Also, we have tested these tools ex vivo.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      By imaging the dynamics of synaptic proteins in cultured neurons, this study presents significant findings regarding the dynamics of excitatory and inhibitory synaptic proteins during development. The evidence shows that the ratios of excitatory and inhibitory synaptic proteins are stable during synapse development. This discovery advances our understanding of the complex mechanisms governing synapse formation. The strength of the evidence is robust, as it is supported by a combination of biological assays and endogenous labeling.

      Strengths:

      This research sheds light on the dynamics of the excitatory and inhibitory synapses during development. It is crucial to understand that while excitatory synapses and inhibitory synapses are developed independently, the ratio of their number is relatively stable during development, maintaining a stable excitatory/inhibitory ratio.

      Important findings and implications in the research include:

      (1) Persistent Synapse Dynamics: Excitatory and inhibitory synapses remain highly dynamic even in mature neurons (DIV12-14), challenging the dogma that synaptic structures are stable after the synaptogenesis stage.

      (2) Maintained E/I Balance: Despite ongoing synapse turnover (formation/elimination) and presynaptic terminal reduction, the overall density and ratio of excitatory-to-inhibitory synapses remain relatively stable during circuit maturation (Figure 7).

      (3) Developmental Shifts: While presynaptic compartments decrease over time, postsynaptic sites increase, suggesting independent regulation of pre- and postsynaptic elements within a stable E/I framework.

      We thank the Reviewer for their positive feedback and careful review of our study.

      Weaknesses:

      This study focuses on specific synaptic proteins within synapses, which may not fully represent the dynamics of other synaptic machinery; also, whether similar observations exist in vivo is still unknown. Further research is needed to explore the implications of these findings in more complex neuronal environments.

      We also thank the Reviewer for their insights and suggestions. We have added discussion of this important point to the Discussion section. Furthermore, we have tested the applicability of our tools ex vivo (new Figures 1, 4, and 6). While using these tools in vivo for live imaging is the eventual goal, we started in a reduced culture system given the relative simplicity. Our current study now provides a framework for future experiments applying these approaches in more complex in vivo systems.

      Reviewer #2 (Public review):

      Summary:

      The Garbett et al. identified a critical need to begin to understand the interplay between the assembly, maturation, and elimination of excitatory and inhibitory synapses. They also detail the lack of reliable tools to address this gap in knowledge. Here, the authors developed synaptic reporters expressed by lentiviruses (mClover3-Homer1c, HaloTag-Syb2, and tdTomatoGephyrin). They combined these reporters with resonance scanning confocal imaging to measure synapses over a 15-hour period during neuron development and in mature neurons in primary hippocampal cultures. Using these reporters in the same neuron, the authors compared the ratios of postsynaptic excitatory and inhibitory specializations that co-localize with presynaptic terminals during development and in mature neurons and found that they are stable across time points. Finally, the authors developed CRISPR/Cas9 tools (TKIT) to knock-in endogenous fluorescent tags (GFP/tdTomato-Gephyrin) or epitope tags (HA-Bassoon and HAHomer1) to begin to study synapse dynamics using endogenous proteins. I believe this paper highlights an important gap in knowledge and begins to offer methodologies to determine the dynamic coordination between excitatory and inhibitory synapses.

      Strengths:

      (1) The experiments are well-designed and carefully controlled.

      (2) The authors carefully validated the reporter and TKIT constructs.

      (3) The authors provide strong proof-of-principle for the use of the reporter constructs to track synapse formation, maintenance, and elimination over a 15-hour period.

      (4) Ingenious use of technologies (reporters, TKIT, and resonance scanning confocal microscopy) to develop a platform for future studies of synapse dynamics.

      (5) Strong evidence supporting that the ratio of excitatory and inhibitory synapses (those that oppose syb2) stays constant through development.

      We thank the Reviewer for their positive assessment of our study.

      Weaknesses:

      Overall, this is a well-executed study that develops tools to simultaneously image excitatory and inhibitory synapse dynamics and represents an important first step to address the fundamental question regarding the coordination between these two types of synapses.

      Minor weaknesses of the manuscript include:

      (1) The lack of a characterization of endogenous Homer1-positive excitatory synapses using TKIT.

      We attempted to perform live imaging of endogenous Homer1-positive synapses using the TKIT approach by tagging endogenous Homer1 with mClover3 but encountered low signal/noise while live imaging. This prompted us to focus our current study on live imaging endogenous Gephyrin. Future studies using more robust tags (e.g. StayGold, HaloTag) for TKIT tagging of endogenous Homer1 will likely help circumvent this issue.

      (2) Discussion about other approaches to study excitatory and inhibitory synapses using endogenous proteins (e.g., intrabodies - FingR or nanobodies) should be included.

      This important point was also raised by other Reviewers. We have now significantly expanded the Discussion section, including discussion of this point.

      (3) The activity state of a neuron and/or a synapse might alter the dynamic properties (formation, maintenance, and/or elimination). A discussion on whether the overexpression of Homer1 and/or gephyrin might alter synapse/neuron activity would provide greater interpretability of the results. A discussion of the potential limitations and benefits of the reporter and TKIT approaches would be beneficial.

      We agree and have added discussion of these points to the Discussion section.

      (4) A description and interpretation of the computational approach to calculate particle tracking would be helpful. I found that particle tracking figures, while elegant, are difficult to interpret.

      As discussed in more detail below, we have generated drift correction and particle tracking approaches for the revised manuscript. We now elaborate on these new approaches in the paper.

      We thank the Reviewer again for their very helpful input and suggestions.

      Reviewer #3 (Public review):

      In the present study, the authors describe the development of new tools and imaging strategies to assess the concomitant development of excitatory and inhibitory synapses in dissociated neuron cultures. To this end, they generate fluorescently tagged constructs of excitatory and inhibitory synapse marker proteins using either conventional overexpression or CRISPR-based strategies. They then image these marker proteins over a timespan of 15 hours to assess synaptic dynamics at different developmental timepoints. Based on their data, they conclude that excitatory and inhibitory synapse development occur in concert to maintain a functional balance despite individual synapse turnover.

      Overall, this study addresses an interesting question, i.e., the interplay between the development of excitatory and inhibitory synapses, which has important implications, particularly for neurodevelopmental disorders in which the balance of excitation and inhibition is disrupted. The experiments are technically solid and well-executed, and the individual images are highly compelling.

      We thank the Reviewer for their positive assessment of our study.

      However, a number of aspects remain to be addressed in order for the study to support the claims made by the authors. First, the novelty aspect of the development of the fluorescently tagged synaptic proteins is unclear, since reporters of this nature are in routine use in many labs. Second, the analysis of the acquired images often seems incomplete, with only example images but no quantification shown, or the distinction between spatial and temporal dynamics appearing unclear. Third, given this incomplete analysis, the interpretations of the authors are not always convincingly supported by the data presented. In conclusion, substantial improvements are required to render the main messages of the study clear and compelling.

      We agree and have incorporated all of the Reviewer’s suggestions in the revised manuscript (please see below).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This is an interesting study. This reviewer has the following questions/comments for the authors:

      (1) Please provide evidence that the gRNAs targeting each gene of synaptic protein have no offtarget effects.

      We now include analysis of off-target effects for the TKIT tools (new Figure S6).

      (2) While structural E/I balance is shown, functional electrophysiological validation (e.g., mEPSC/mIPSC ratios) is absent. It is interesting to know whether the balanced functional structural changes translate to functional?

      We thank the Reviewer for this insightful suggestion and now include these recordings in the revised paper (new Figure 8).

      (3) In lines 217-218, please define thresholds for "stable" vs. "dynamic" puncta (e.g., temporal and spatial criteria).

      We more clearly define our categorization parameters (e.g. new Figure 2).

      (4) In Figure 5B: The low co-localization between endogenously tagged Bassoon and antibodystained Bassoon is likely due to the low TKIT efficiency. Quite a few HA-tagged Basson signals are insensitive to Basson-antibody. The authors are suggested to explain those.

      We thank the Reviewer for identifying this and add discussion to the Results section.

      (5) For the data analysis. If each n represents an independent neuronal culture, should the authors are suggested to provide the number of neurons/dendrites analyzed for each independent culture?

      We have added these important details to the manuscript.

      (6) Regarding the title, the author used the term "coordinated dynamics". This reviewer finds it is a bit over-claim because the stable ratios of the number of excitatory synapses and inhibitory synapses are likely an association, not actively "coordinated". I suggest that the authors rephrase this.

      We agree that we cannot argue that excitatory and inhibitory synapses are causally coordinated in our current study. Their levels are likely associated by either association or direct coupling, which we now discuss further in the first paragraph of the Discussion. We have rephrased the title accordingly.

      Reviewer #2 (Recommendations for the authors):

      I have only minor suggestions that I think will improve the manuscript:

      (1) Please define Syn1/2 on line 129.

      We have defined this in the revised paper.

      (2) For Figures 2B, C, and 4B, C: are the puncta in panel C from the dendrites in panels B? If so, it would be helpful to identify the ROIs selected in panels C.

      We now include this in new Figure 2.

      (3) For the particle tracking figures, while the ability to track all synaptic puncta is very impressive, it is sometimes difficult to clearly track the lifespan of a synaptic puncta from the current figures. I believe that it would be helpful if the authors selected specific examples of synapses formed, maintained, and eliminated.

      We agree and now include more examples.

      (4) I believe that more detail about the computational approach and analysis for the particle tracking (Figs 2E and 4E) would help the interpretability of the figure.

      This important point was also raised by the other Reviewers. We generated custom tools during the revision that significantly expand the capabilities of our tracking approaches and more clearly describe them in the revised manuscript.

      (5) Similar to the rigorous gephyrin TKIT analysis (Fig. 6), did the authors perform a similar analysis for Homer1c TKIT? This might be valuable to confirm that overexpression of the Homer1 reporter does not indirectly alter synapse dynamics.

      We attempted to perform live imaging of mClover3 TKIT-tagged endogenous Homer1 but encountered low signal/noise with live imaging. We now add discussion that optimization of more robust tags (e.g. StayGold, HaloTag) will likely be necessary for live imaging of different target proteins.

      (6) The tools developed by Garbett et al. have the potential to be broadly utilized in the field to provide new insight into the coordination of excitatory and inhibitory synapses. It would thus be helpful for the authors to include a discussion about the strengths and limitations of the reporter and TKIT methods relative to other approaches used to live image synapses (e.g., intrabodies (FingR and nanobodies)).

      We have now significantly expanded the Discussion to include these important points.

      (7) In the discussion, can the authors elaborate on whether it is experimentally feasible to apply their TKIT labeling of gephyrin and Homer1c in the same neuron to assess the endogenous excitatory and inhibitory synapse dynamics from the same neuron?

      We have added discussion of this point and also proof-of-concept data supporting tagging of two postsynaptic targets within the same neuron (new Figure S5D).

      Reviewer #3 (Recommendations for the authors):

      (1) While the new tools described in the current manuscript can undoubtedly be used for the described purposes, the novelty of these tools is unclear to me. Viral vectors expressing fluorescently tagged versions of Homer1, synaptobrevin, and gephyrin are commercially available, e.g., via Addgene, and they are in routine use in many labs. CRISPR-mediated strategies for this purpose have also been previously reported (e.g., Willems et al. 2020, PLOS Biology; Fang et al. 2021, eLife). It is not clear to me how the tools reported here present a significant improvement over existing resources, other than that they use different fluorescent tags. If this aspect is a central part of the current manuscript, it should be expanded on in the discussion, including a direct comparison with available tools to highlight the novel aspects.

      We agree and have significantly expanded the Discussion to include these important points. Also, rather than argue that our tools are superior to pre-existing approaches, we adjust the text to argue that our tools and analytical approaches have been designed and optimized for the purposes we apply them to.

      (2) In addition to generating new tagged constructs, the authors also state that they have developed new imaging and analysis strategies to facilitate long-term assessment of synaptic dynamics. However, in many figures, they present only sample images, with little quantification to allow assessment of the wider relevance of the imaged synapses. For example, in Figures 2C and 4C, they present one example each of, e.g., a stable, nascent, transient, or eliminated synapse. However, they do not provide any quantification on how frequently any of these events occur, or whether they can be reliably quantified at all. These quantifications (i.e., percentage of each event type across a large population of synapses) would be necessary and should be added to demonstrate that this tool can be used for more than single example images.

      We have generated custom-made drift correction and particle tracking approaches for the revised manuscript. Based on the reviewer’s suggestion, we have quantified the relative frequencies of stable, nascent, transient, and eliminated synapses (Fig 2B-G, Fig3A-F, Fig 5A-F, Fig 7B-C). These metrics greatly enhance the biological interpretation of our results. We have also added a supplemental movie with an example image with corresponding categorized tracks for each puncta type (Movie S3)

      (3) The authors do present an automated visual representation of spatial track length across the neuron, e.g., in Figure 2E and 4E, although this is also not quantified. Moreover, the track lengths appear surprisingly short, despite the authors' claims that their analyses 'highlight the dynamic nature of excitatory synapses over these timescales'. It is not clear to me whether these short tracks are more than just jitter, either in the synapses themselves or in the images due to technical limitations. E.g., in panel 2E, I see very few examples in which the track is not simply centered around one point, but actually expands over a distance. Quantification of the distance between start and end points of the tracks would be important to support the claim that these synapses are dynamic in terms of spatial translocation (if that is what the authors meant). Or if the 'dynamic nature' of the synapses referred to temporal dynamics, it is unclear to me how this information can be gained from the represented tracks.

      We thank the reviewer for these excellent points. To accurately access spatial motion, we drift-corrected our images with a custom correction algorithm to eliminate stage or microscope drift as a source of contaminating motion (See Methods, Movie S2), in addition to collecting time-lapse imaging with Nikon perfect focus. We noticed heterogeneity in our cultures such that some areas contained very mobile neurites, while other remained stationary (Fig. S1). We binned movies into either moving or still neurites and assessed spatial metrics as suggested (Fig. S1A). Consistent with our binning, puncta on moving neurites showed larger net displacement (distance between start and end points), but puncta on still neurites also showed ~1 µm net displacement (Fig. S1D). We also quantified puncta speed and found that puncta on moving neurites generally moved faster (Fig. S1C). We appreciate the reviewer’s insight that track length were surprisingly short, and after employing our drift correction and revised tracking methods, we now see substantially longer track lengths (Fig 2E, Fig 3C & F, Fig S2B & C). We additionally see a large fraction of tracks that persist throughout the imaging session (Fig 2E, Fig S2B & C).

      (4) In Figure 3, the authors now quantify track length, but in this case in the unit 'minutes', from which I would interpret that this is now meant to assess the temporal dynamics rather than the spatial dynamics. The lack of a clear distinction between spatial dynamics and temporal dynamics is very confusing to me, since these are entirely independent measures. 'Track length' to me indicates spatial dynamics, and I would expect the units to be a measure of distance. 'Track duration', which the authors also use in some places, but inconsistently as far as I can tell, makes sense to me for the assessment of temporal dynamics, with the units being a measure of time. I would strongly recommend being very clear about this distinction, since the current representation of the data is very difficult to follow and interpret.

      In addition to new spatial metrics, we have clarified in the text when we are referring to spatial dynamics (distance) versus temporal dynamics (time). As suggested, we use duration when referring to time, and speed or distance when referring to spatial metrics.

      (5) The images from the newly generated CRISPR-based tags in Figures 5-7 are striking and very compelling - these will be very useful tools. However, here too, it seems that the interpretation of the data does not really match the results. All quantification indicates that there is very little change in synapse density or other assessed parameters over the time course of the imaging, and yet the authors emphasize the dynamic nature of visualized synapses. More compelling quantification would be needed to support this claim.

      We have quantified spatial and temporal metrics for live neuron culture imaging for all tools developed including CRISPR-based tags (Figure 7).

      (6) The discussion is extremely short and provides almost no integration of the results of the study into the framework of existing knowledge. Instead, it focuses almost exclusively on unanswered questions and future perspectives, which are also important, but not helpful in interpreting the findings from the current study. The latter aspects should be added to provide essential context for the current findings.

      We agree and have added additional discussion of our current findings to help contextualize their significance.

      We thank the Reviewers again for their positive feedback and insightful input, which has undoubtedly strengthened our study.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates how human temporal voice areas (TVA) respond to vocalizations from nonhuman primates. Using functional MRI during a species-categorization task, the authors compare neural responses to calls from humans, chimpanzees, bonobos, and macaques while modeling both acoustic and phylogenetic factors. They find that bilateral anterior TVA regions respond more strongly to chimpanzee than to other nonhuman primate vocalizations, suggesting that these regions are sensitive not only to human voices but also to acoustically and evolutionarily related sounds.

      The work provides important comparative evidence for continuity in primate vocal communication and offers a strong empirical foundation for modeling how specific acoustic features drive TVA activity.

      Strengths:

      (1) Comparative scope: The inclusion of four primate species, including both great apes and monkeys, provides a rare and valuable cross-species perspective on voice processing.

      (2) Methodological rigor: Acoustic and phylogenetic distances are carefully quantified and incorporated into the analyses.

      (4) Neuroscientific significance: The finding of TVA sensitivity to chimpanzee calls supports the view that human voice-selective regions are evolutionarily tuned to certain acoustic features shared across primates.

      (4) Clear presentation: The study is well organized, the stimuli well controlled, and the imaging analyses transparent and replicable.

      (5) Theoretical contribution: The results advance understanding of the neural bases of voice perception and the evolutionary roots of voice sensitivity in the human brain.

      Weaknesses:

      (1) Acoustic-phylogenetic confound: The design does not fully disentangle acoustic similarity from phylogenetic proximity, as species co-vary along both dimensions. A promising way to address this would be to include an additional model focusing on the acoustic features that specifically differentiate bonobo from chimpanzee calls, which share equal phylogenetic distance to humans.

      (2) Selectivity vs. sensitivity: Without non-vocal control sounds, the study cannot determine whether TVA responses reflect true selectivity for primate vocalizations or general auditory sensitivity.

      (3) Task demands: The use of an active categorization task may engage additional cognitive processes beyond auditory perception; a passive listening condition would help clarify the contribution of attention and task performance.

      (4) Figures and presentation: Some results are partially redundant; keeping only the most representative model figure in the main text and moving others to the Supplementary Material would improve clarity.

      We thank the reviewer for contributing to the improvement of the present study and for the extremely constructive criticism. Concerning the identified weaknesses of our work, we provide here some general answers while the detailed review (below) addresses point-by-point the reviews in high detail.

      (1) We totally agree that acoustics and phylogeny cannot be disentangled in our study, which is a limitation. We now provide the suggested analysis on the acoustic specificities of chimpanzee and bonobo calls.

      (2) This point on selectivity vs. specificity is indeed crucial, and we now provide a more careful viewpoint and phrasing on this aspect, since our study can only provide partial arguments for this important distinction.

      (3) Task demand following species categorization might rightfully yield to the engagement of distinct brain network compared to merely listening to the stimuli. We discuss this aspect and put forward the argument that, while we cannot control for this aspect, our attentional control study performed by an independent sample, N=28 provides clear evidence that no species triggered an attention bias. In other words, task demand might play a role, but at least in the study we know that attentional resources were not biased towards one species in particular since no effects were observed.

      (4) We agree that results were not articulated in a clear fashion and that figures were redundant. We addressed this aspect and regrouped the figures where appropriate while we include the rest in the supplementary material now.

      Reviewer #2 (Public review):

      Summary:

      This study investigated how the human brain responds to vocalizations from multiple primate species, including humans, chimpanzees, bonobos, and rhesus macaques. The central finding - that subregions of the temporal voice areas (TVA), particularly in the bilateral anterior superior temporal gyrus, show enhanced responses to chimpanzee vocalizations - suggests a potential neural sensitivity to calls from phylogenetically close nonhuman primates.

      Strengths:

      The authors employed three analytical models to consistently demonstrate activation in the anterior superior temporal gyrus that is specific to chimpanzee calls. The methodology was logical and robust, and the results supporting these findings appear solid.

      Weaknesses:

      The interpretation of the findings in this paper regarding the evolutionary continuity of voice processing lacks sufficient evidence. A simple explanation is that the observed effects can be attributed to the similarity in low-level acoustic features, rather than effects specific to phylogenetically close species. The authors only tested vocalizations from three non-human primate species, other than humans. In this case, the species specificity of the effect does not fully represent the specificity of evolutionary relatedness.

      We want to thank the reviewer for the constructive criticism and for evaluating the manuscript.

      Concerning the principal weakness highlighted, we provide new analyses behavioral, acoustics, model-based fMRI that improve our understanding of the influence of both phylogeny and bioacoustics in our data. We argue that the explanation proposed by the reviewer cannot explain our results, as also observed in several other research from us and others. We discuss this aspect and emphasize that including stimuli from more species would greatly improve the understanding of phylogeny and bioacoustics in this context.

      Reviewer #3 (Public review):

      Summary:

      Ceravolo et al. employed functional magnetic resonance imaging (fMRI) to examine how the temporal voice areas (TVA) in the human brain respond to vocalizations from different nonhuman primate species. Their findings reveal that the human TVA is not only responsible for human vocalizations but also exhibits sensitivity to the vocalizations of other primates, particularly chimpanzee vocalizations sharing acoustic similarities with human voices, which offers compelling evidence for cross-species vocal processing in the human auditory system. Overall, the study presents intellectually stimulating hypotheses and demonstrates methodological originality. However, the current findings are not yet solid enough to fully support the proposed claims, and the presentation could be enhanced for clarity and impact.

      Strengths:

      The study presents intellectually stimulating hypotheses and demonstrates methodological originality.

      Weaknesses:

      (1) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy.

      We thank the reviewer for evaluating our manuscript and for the constructive criticism as well as the many suggestions. Concerning the weaknesses of the study, we provide here some quick answers while more detailed responses can be found below.

      (1) We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator.

      (2) We totally agree that figure redundancy was a problem and we now reduced confusion by combining congruent aspects while pushing other results to the supplementary material.

      Recommendations for the authors:

      Reviewing Editor Comments:

      With additional analyses and discussions, the work has the potential to offer important insight into the evolutionary continuity of voice processing.

      We thank the Reviewing Editor for this additional motivation and for offering us the possibility to revise our manuscript. We will now provide our point-by-point reviewing, referring to manuscript modifications by section and/or line number(s). All modifications are also highlighted in light grey in the text.

      Reviewer #1 (Recommendations for the authors):

      The manuscript is clearly written and addresses an important comparative question about the specificity of human TVA responses. The acoustic analyses are well designed, and the imaging work is careful and thorough. However, several conceptual and methodological issues need clarification or tempering of claims, particularly regarding (i) the distinction between sensitivity and selectivity, (ii) the confounding of acoustic and phylogenetic factors, and (iii) the interpretation of "chimpanzee-specific" TVA activity.

      (1) Introduction

      Line 48: cite more recent infant EEG evidence for early voice sensitivity (Calce, Curr Biol).

      The reference and explanation were added, lines 46-48.

      Line 53: mention recent data on voice processing in marmosets (Jafari, Cell Rep; Dureux, Curr Biol).

      We added the references and the mention of these interesting studies on common marmosets, lines 53-54.

      Line 59: Fecteau et al. (2004) already explored cross-species selectivity; please integrate and discuss.

      We now mention here the work from Fecteau and colleagues and its relevance, see lines 57-59.

      Line 70: clarify that in [27] (Bodin et al., 2021) human TVA responded similarly to human nonverbal vocalizations and macaque coos, likely due to acoustic similarity.

      We added this important aspect, thank you for this precision. See lines 71-72.

      Clarify why an active species-categorization task was chosen instead of passive listening, which is standard in TVA research. Were participants familiarized with stimuli beforehand?

      We added a sentence on this aspect, but basically to summarize it here: we wanted to be able to test human recognition of nonhuman primate species’ calls. From the start, we wanted to test the frontal mechanisms related to decision-based processes of humans when categorizing non-human primate calls hence the 2023 article we published. See lines 75-77 and we also added information on familiarization to the stimuli in the Methods, lines 679-682.

      The 16 acoustic features mentioned should be briefly defined earlier, as they are central.

      We feel like describing 16 acoustic parameters in the introduction would be heavy on the reader, so we instead added a reference to the supplementary table (Table S1) in which these are named and described. See line 80.

      Explain why only chimpanzees and bonobos were selected among the great apes, and discuss the value of including both, given their equal phylogenetic proximity but largely dissimilar acoustics.

      The stimuli were obtained by Thibaud Gruber and his team and through collaborations with Katie Slocombe and Zanna Clay. Unfortunately, at the time we could only use chimpanzee and bonobo calls for the great apes. Therefore, it was mainly a material constraint rather than a deliberate choice to exclude other great apes. We now discuss this aspect and present the absence of other great apes as a limitation (lines 587-591).

      Rephrase references to "recruitment" of TVA - this term implies general activation, while the key question concerns selectivity (stronger responses to voices vs. non-vocal controls).

      We rephrased throughout the manuscript, thank you for this suggestion.

      The hypothesis section should more clearly separate the acoustic and phylogenetic predictions, and clarify which earlier data motivate each.

      We now explicitly categorize the hypotheses according to either Bioacoustics or Phylogeny to clarify. We also added references motivating each hypothesis. See lines 114-120.

      (2) Methods

      Clarify whether stimuli were RMS-normalized or otherwise balanced for energy (line 128).

      Sound pressure level was kept constant but the stimuli were not normalized, specifically to avoid a negative impact on their naturality. We added a sentence (lines 131-132) including a reference on this aspect.

      The task design could benefit from reporting accuracy in addition to reaction times for the 4AFC species classification task.

      We agree this aspect was missing. We now report accuracy data (controlled for reaction times and acoustics of Model 3) for the species categorization task (lines 147-165; Fig.1B), and in the Methods (lines 769-786). The fitted regression values of this analysis are also used for a new fMRI model (Model 4), to uncover within-TVA correlates of the probability of correct species categorization (lines 309-325; Fig.4).

      Please note that previously, the behavioral data of the species categorization task were completely absent (N=23), and the reaction times data previously part of Fig.1 were for the species attentional bias task (independent sample of N=28). Since this aspect was not clear at all (same remark by all reviewers—apologies for that), we now include a clear separation in Fig.1, with newly added panels D & E part of a distinct figure area named: “Control task: Testing for Species attentional bias (N=28)”. Panel D illustrates the control task paradigm (each species as exogenous cue; “dot-probe” paradigm) while panel E shows the results (target sine wave tone or “bip” detection reaction times), showing that no species triggered more attentional capture than the others (Species effect non-significant).

      The acoustic parameters used in Models 2 and 3 should be explicitly listed in the Methods (even if already published elsewhere).

      In addition to their description in Table S1, we now include the 16 acoustic parameters used to calculate acoustic distance between the species in the Methods, see lines 828-844.

      Consider simplifying the presentation of the three models: a figure summarizing their relationships would help.

      We now include only one figure (Fig.2) for Model 3, and we pushed model 1&2 to the supplementary material. We also simplified Fig.3 for a clearer view of the overlaps between the 3 models within the TVA.

      The description of “systematic and thorough control of phylogeny” (line 119) is overstated, given that only three nonhuman species were included.

      We agree with the reviewer and we suppressed both “systematic” and “thorough” from the sentence.

      Provide rationale for not including a nonvocal control category (e.g., scrambled vocalizations or environmental sounds) to assess TVA selectivity.

      The main objective of the study was to uncover whether human participants could recognize the vocalizations from nonhuman primates—from both great apes and monkeys—as compared to the human voice. We therefore did not include nonvocal or noise stimuli. We added this point as a limitation in the Discussion (lines 593-596 and 609-611).

      Even though we did not include such stimuli for the reason mentioned above, the delineation of subtypes of nonvocal material within the TVA of our participants (Fig.2) are, in our opinion, clarifying the message: chimpanzee-selective activations are fully within ‘voice vs. animal’ and ‘voice vs. nature’ TVA subareas, while it is not the case in ‘voice vs. music’ and ‘voice vs. noise’ TVA subareas.

      Clarify if participants were trained or had a practice session to recognize the four species before scanning.

      The participants were indeed trained on 3 stimuli per species before entering the MRI scanner. These stimuli were discarded from the species categorization task. We added a sentence about this aspect, see lines 131-132.

      Specify what is meant by "no good or bad response" in the attentional control task (line 724).

      We suppressed this wording as it was highly confusing.

      (3) Results

      Behavioral accuracy should be reported to complement reaction times.

      We now added behavioral data for the species categorization task as well as the neural correlates of accurate species categorization. See our previous response above (‘‘‘).

      Figures 2-4 largely overlap; consider merging or simplifying to reduce redundancy.

      We agree and this point was raised by the other reviewers as well. Task-based results are now presented only for Model 3 as Fig.2, while Fig.3 (previously Fig.5) summarizes the overlap between the three models. Figures for Models 1 & 2, previously labelled Fig.3 and Fig.4, were moved to the supplementary material.

      Figure 2: Please indicate more clearly where "chimp-selective" areas are located (perhaps with zooms).

      We agree, we now modified Fig.2 with zoomed-in panels and a clearer outline of chimp-selective areas (solid blue outline). This outline is also referenced in the text (lines 236-237).

      Correction for multiple contrasts: With many pairwise tests, adjustments (Bonferroni or FDR) should be mentioned explicitly.

      We now specify ‘FDR correction at the voxel level’ at the beginning of the Results section (lines 195-198) as well as in each figure.

      Replace "specific to chimpanzee" with "selective for chimpanzee" to avoid implying exclusivity.

      We made the suggested replacement throughout the manuscript.

      Discuss whether the small macaque-related clusters might simply reflect acoustic overlap rather than true category selectivity.

      We added a section on this important aspect, including results that support the role of mid-STG/STS regions for more noise-like stimuli, including the use of macaque coos. See lines 450-461.

      (4) Discussion

      The discussion overstates claims of "chimpanzee-selectivity" in TVA. The evidence shows relative preference, not absolute selectivity.

      We now specify from the start of the Discussion that we are not interpreting the results as absolute selectivity but rather as more relative preference, see lines 371-373.

      The authors repeatedly conflate acoustic and phylogenetic factors; this should be explicitly acknowledged as a limitation.

      We agree, and we completed the limitations section already dedicated to this aspect by a more explicit account of the confound, see lines 609-611.

      Clarify what is meant by "recruitment" and "selectivity" (lines 411-419, 577). TVA activity often reflects enhanced responses to voices compared to non-vocal sounds, not exclusive activation.

      We clarified this wording in the Discussion (lines 377-378) and replaced another instance by “activated the […]” to make it clearer what we imply, namely enhanced activity triggered by chimpanzee calls within human TVA.

      The lack of non-vocal control conditions should be discussed as a major interpretive limitation.

      We added this point as a limitation in the Discussion (lines 593-596).

      The statement that "chimpanzee-selective activity" arose in humans who have never been exposed to chimp calls (line 450) invites evolutionary speculation but should be more cautiously phrased.

      We agree, and we rephrased by: “[…] with chimpanzee calls triggering responses in the anterior STG/TVA of our human participants […]”. See lines 432-433.

      The comparison to recent macaque data (Giamundo et al., 2024 PNAS) is crucial: these findings of human-voice-selective neurons in macaques directly parallel the present human-chimp result.

      We agree with the reviewer, and we are hopeful to read similar results for other apes/great apes in the future.

      Reviewer #2 (Recommendations for the authors):

      (1) The primate vocalizations used in this study were recorded in diverse social and emotional contexts, which may have contributed to the observed differences in TVA activation. Since the temporal voice areas are known to be sensitive to affective and socially relevant cues, these contextual differences could confound the interpretation of species-specific neural responses. Therefore, I suggest that the authors conduct a post-hoc analysis to quantify and compare the affective valence, arousal levels, and social contexts associated with each stimulus set.

      We agree that the TVA are sensitive to social—or socially relevant—cues, motivating the very thorough work of the expert reserve personnel on-site to accurately categorize the calls according to the very specific context they were produced in. If the reviewer meant presenting these stimuli to non-expert participants and asking them to categorize the context or valence, we think it would make no sense since the ratings would be completely below chance level and therefore uninformative. The newly added behavior—and model-based fmri—data include this crucial point, a factor that we named ‘Context’ in our analyses. In fact, for each species’ 18 stimuli, we control for agonistic and affiliative production context—split evenly, per species. Also, computing an additional posthoc analysis by splitting the stimuli according to Context would result in too few trials to get sensible and reliable fMRI results.

      That being said, our study targets this specific aspect by extracting the acoustic features that characterize our stimulus set the best, across context-species-valence-arousal, which is exactly what we want. Through the three types of modeling we used—from more simplistic to more elaborate the results converge only for one species: chimpanzee calls.

      We think the addition of behavioral data, model-based fMRI data, and the specific analysis on acoustic differences between chimpanzee and bonobo calls strengthens the message and the validity of our findings.

      (2) Although the author mentioned that the behavioral effects triggered by these vocalizations have been reported previously, the behavioral responses of the participants in the current study are also crucial for our understanding of the results. If the MRI data can be combined with the participants' behavioral responses for comprehensive analysis, the conclusions of this study will be more compelling.

      We agree with the reviewer, and we added the behavioral data—controlling for reaction times, production context and acoustics of interest—and we also included a model-based fMRI modeling of the probability of correct species categorization as Model 4, Fig.4. See, respectively: lines 147-165, Fig.1B; Methods, lines 769-786; Neuroimaging results, lines 309-325.

      (3) I am still not convinced that phylogenetic proximity drives the observed neural selectivity. While chimpanzee vocalizations do elicit stronger responses in anterior STG, the claim that this reflects evolutionary relatedness lacks evidence. If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree with the reviewer that generalizing our results in terms of phylogenetic proximity alone is not a viable option. Including many more primate species including other great apes would be necessary, and we mention this crucial aspect in the limitations section. We also insist in the Discussion on the interdependence between phylogeny and acoustics in our data, since: 1) we cannot fully disentangle these factors here, 2) we cannot attribute our results to either one or the other. See lines 387-390, 410-411, 473-477, 587-591.

      If the acoustic features of a certain call from a particular species are similar to those of human voices, it may also lead to similar effects.

      We agree, and nobody could disagree: if an auditory object is extremely similar to the human voice in terms of acoustics, it would therefore potentially activate the TVA. This is exactly our message: in the natural ‘auditory world’, the calls from chimpanzees seem to be among the very few animal auditory signals that are sufficiently close, acoustically, to the human voice and therefore trigger TVA activity. They also happen to be the calls from a species which is phylogenetically the closest to humans with minimal differences with other great apes. Our results are in that sense very aligned with work from the laboratory of Pascal Belin, namely on ‘voice patches’ in the primate brain located in the (anterior) TVA, cited in our manuscript.

      We therefore think our interpretation does not exclude that in the near future, similar results within the TVA could be observed for other auditory objects, and if animal, from a species potentially much more distant phylogenetically or from vocal signals of other great apes.

      We added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as scrambled or spectrum shifted per-species stimuli, which would have made the interpretation clearer identical acoustics but alteration/destruction of the species auditory object. See lines 593-596 and 609-611.

      Reviewer #3 (Recommendations for the authors):

      While the manuscript presents intriguing results, several concerns are raised for further consideration, detailed below.

      We thank the reviewer for evaluating the manuscript and for the constructive criticism and suggestions.

      Major concerns:

      (1) This study claims that bilateral anterior superior temporal gyrus (aSTG) in humans can be specifically activated by chimpanzee vocalizations rather than all other primate species after regressing out relevant acoustic parameters using three distinct analyses. I am wondering if a control stimulus (e.g., scrambled chimpanzee vocalizations) were presented, would the activation patterns in these same temporal voice areas (TVA) exhibit significant differences compared to the natural chimpanzee vocalizations?

      We completely agree with the reviewer, and this point was also raised by the other reviewers. We therefore added a key limitation point in the Discussion on the absence of auditory control stimuli in our design, such as per-species scrambled or spectrum shifted stimuli, which would have made the interpretation clearer—identical acoustics but alteration/destruction of the species auditory object. See lines 609-611.

      (2) The figure organization/presentation requires significant revision to avoid confusion and redundancy. E.g:

      Figure 1C is the same as Figure S1. In addition, Figure 1C lacks a figure legend and descriptive label.

      The scatter plots in Figures 2D, 2H, 3D, 3H, and 4D, 4H are same as those in Figures S2, S3, and S4. However, some of these duplicate plots even have inconsistent axis labels.

      In several panels, the main figures appear to be summaries derived from the supplementary figures. The authors should organize these figures well to eliminate redundancy.

      Please double-check all the figures to make sure of accuracy.

      We agree that the figures were badly organized and were too crowded and redundant. We now suppressed the redundancy between Fig.1 and Fig.S1, and we reduced fMRI results to one figure for statistical Model 3 while the other models are in the supplementary data—we also justify this decision in the text by highlighting that model 3 is the most elaborate and sensitive one. Fig.3 (previously ‘Fig.5’) shows the overlaps between models and was simplified and clarified as well.

      (3) The analysis of the fMRI data does not account for the participants' behavioral performance, specifically their reaction times (RTs) during the species categorization task. It is possible that processing vocalizations from certain species requires more cognitive effort or induces higher decision uncertainty. Could the observed neural effects be confounded by the decision-making process itself?

      We now include behavioral data analysis (accuracy data controlled for reaction times and acoustics of existing Model 3, using mixed-effects logistic regression) in addition to a new, 4th model for fMRI data. This 4th model was computed in a model-based fashion by modeling the probability of correct categorization within the TVA (fitted regression coefficients, per Participant, Species, Trial) and revealing the neural correlates of this modulator. We now display these results in Fig.4 and we introduce the motivation factor for including a categorization task rather than more traditional passive listening (lines 75-77), as well as limitations, lines 595-596.

      (4) One interesting attempt of this study is to dissociate biologically salient information in animal vocalizations from their low-level acoustic properties. This presents a fundamental conceptual challenge: how to rigorously disentangle a vocalization's species-specific attributes from its inherent acoustic correlates. More precisely, what essential biological information persists in a species' vocal signal after statistically accounting for all quantifiable acoustic features? I recommend that the authors address it in the discussion.

      We thank the reviewer for this very important comment, and for suggesting we discuss it in the manuscript. We completely agree: we cannot fully orthogonalize species and acoustics, and this aspect relates also more broadly to cognitive and affective neuroscience studies involving vocal material. Namely: “What is an auditory object without acoustics?”

      We included a full paragraph on this aspect, see Discussion, lines 570-584.

      (5) If a brain region, such as TVA, is responsive to both acoustic parameters and biological meanings of animal vocalizations, the method used in this study might be inadequate by setting covariates to zero. It is possible that species information is embedded within a specific acoustic pattern. The current modeling approach may not capture such complex information and could potentially introduce bias when estimating the species effect. I recommend that the authors address this issue in the discussion.

      We thank the reviewer for this point once again, we addressed it in the Discussion, lines 581-584, and also in the section dedicated to study limitations, lines 609-613.

      (6) In the discussion, non-human primate vocalizations are "unreadable" to humans. If this is the case, what is the fundamental perceptual difference between these vocalizations and those from the other animal species? An alternative and highly plausible explanation for the findings is the differential familiarity of the participants with the various species, driven by media exposure (e.g., documentaries) or zoo visits and interactions. The authors need to provide a stronger justification for their control stimuli and directly address, either through discussion or additional analysis, how the factor of familiarity might explain their results better than the proposed "evolutionary distance" hypothesis.

      We now discuss this important aspect, see lines 560-569.

      We thought about doing additional analyses on this aspect but we concluded that we did not have any reliable indicators of familiarity for our participants, and additionally they were all recruited for being ‘unfamiliar’ with great apes or old-world monkeys’ vocalized communication.

      Also, frequent mismatches in the media between images of apes and the associated vocal signals (for instance, the depiction of a chimpanzee but with background audio of macaque coos) are not helping this cause.

      Minor:

      (1) No figure legend and result description for Figure 1.

      Figure 1 has a legend, maybe it was cut out during the uploading process, but it is present and verified now.

      (2) In the main text, three statistical models were referenced. Was the data used in each subsequent statistical model derived from the processed data of the preceding model? Please clearly explain this in the main text.

      We now specify this aspect in the Methods and the Results section to clarify that each model is independent from the others (lines 964-966 and 189-191, respectively).

      (3) In Figure 5, the two dashed lines representing Model 1 and Model 2 are confusing for readers.

      We modified the figure (now Fig.3) and simplified it by removing some outlines and clarifying the colors, therefore improving readability.

      (4) Lack of reaction times in the species categorization task.

      We clarified behavioral data, including the results for the species categorization task and for the control, exogenous cueing task, see modified Fig.1 and behavioral results section of the Results.

      (5) Figures 2, 3, 4, 5, Please keep the font size of the figure title consistent.

      Figure title font size were uniformized.

      (6) Line 201, Line 224, and so on, (EFG) → (E, F, G).

      We modified this aspect in every figure legend, including the supplementary material.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Argunşah et al. describe and investigate the mechanisms underlying the differential response dynamics of barrel vs septa domains in the whisker-related primary somatosensory cortex (S1). Upon repeated stimulation, the authors report that the response ratio between multi- and single-whisker stimulation increases in layer (L) 4 neurons of the septal domain, while remaining constant in barrel L4 neurons. The authors attribute this divergence to differences in short-term synaptic plasticity, particularly within somatostatin-expressing (SST<sup>+</sup>) interneurons. This interpretation is supported by

      (1) The increased density of SST+ neurons in L4 of the septa compared to barrel domain,

      (2) The stronger response of (L2/3) SST+ neurons to repeated multi- vs single-whisker stimulation and

      (3) the reduced functional difference in single- versus multi-whisker response ratios across barrel and septal domains in Elfn1 KO mice, which lack a synaptic protein that confers characteristic short-term plasticity, notably in SST+ neurons.

      Consistently, a decoder trained on WT data fails to generalize to Elfn1 KO responses. Finally, the authors report a relative enrichment of S2- and M1-projecting cell densities in L4 of the septal domain compared to the barrel domain, suggesting that septal and barrel circuits may differentially route information about single vs multi-whisker stimulation downstream of S1.

      Strengths:

      This paper describes and aims to study a circuit underlying differential response between barrel columns and septal domains of the primary somatosensory cortex. This work supports the view these two domains contribute distinctly to the processing single versus multi-whisker inputs and highlight the role of SST+ neuron and their short-term plasticity. Together, this study suggests that the barrel cortex multiplexes whisker-derived sensory information across its domains, enabling parallel processing within S1.

      Weaknesses:

      Although the divergence in responses to repeated single- versus multi-whisker stimulation between barrel and septal domains is consistent with a role for SST<sup>+</sup> neuron short-term plasticity, the evidence presented does not conclusively demonstrate that this mechanism is the critical driver of the difference. The lack of targeted recordings and manipulations limits the strength of this conclusion: SST<sup>+</sup> neuron activity is not measured in L4, nor is it assessed in a domain-specific manner. The Elfn1 knockout manipulation does not appear to selectively affect either stimulus condition, domain or interneuron subtype. Finally, all experiments were performed under anesthesia, which raises concerns about how well the reported dynamics generalize to awake cortical processing.

      We thank the reviewer for their careful reading of the manuscript and their balanced assessment of both its strengths and limitations. We acknowledge the reviewer’s concerns regarding the lack of direct, layer- and cell-type–specific recordings and manipulations of SST<sup>+</sup> interneurons, as well as the use of anesthesia. As noted in the Discussion, these factors limit the extent to which causal mechanisms can be established and the degree to which the reported dynamics can be generalized to awake cortical processing. For this reason, we intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, developmental, physiological, and genetic evidence, rather than as definitive proof. We believe this conceptual framing appropriately reflects the scope of the current data while highlighting clear directions for future work.

      Reviewer #2 (Public review):

      Summary:

      Argunsah and colleagues demonstrate that SST expressing interneurons are concentrated in the mouse septa and differentially respond to repetitive multi-whisker inputs. Identifying how a specific neuronal phenotype impacts responses is an advance.

      Strengths:

      (1) Careful physiological and imaging studies.

      (2) Novel result showing the role of SST+ neurons in shaping responses.

      (3) Good use of a knockout animal to further the main hypothesis.

      (4) Clear analytical techniques.

      Comments on revisions:

      The authors have effectively responded to my initial critiques - I have no further concerns.

      We thank the reviewer for their positive evaluation of our work and for recognizing the novelty of the findings, the careful physiological and imaging approaches, the use of the Elfn1 knockout model, and the clarity of the analytical framework. We are pleased that the reviewer has no further concerns and appreciates the contribution of this study to understanding the role of SST<sup>+</sup> interneurons in shaping sensory processing in the barrel cortex.

      Reviewer #3 (Public review):

      Summary:

      This study investigates the functional differences between barrel and septal columns in the mouse somatosensory cortex, focusing on how local inhibitory dynamics (particularly involving SST<sup>+</sup> interneurons) may mediate temporal integration of multi- whisker (MW) stimuli in septa. Using a combination of in vivo multi-unit recordings, calcium imaging, and anatomical tracing, the authors propose a model in which Elfn1-dependent synaptic facilitation onto SST<sup>+</sup> interneurons contributes to the distinct sensory responses to MW input in barrels and septa, enabling functional segregation between these domains.

      Strengths:

      The study presents a thought-provoking and useful conceptual model for understanding sensory processing in the somatosensory cortex. While barrel columns have been widely studied, septal regions remain relatively understudied in mice. If septa indeed act as selective integrators of distributed sensory input, this would suggest a novel computational role for cortical microcircuits beyond the classical view focused on barrels. Although still hypothetical, the proposed model in which SST<sup>+</sup> interneurons contribute to domain-specific sensory responses between barrel and septal domains is intriguing and opens new avenues for investigating inhibitory circuit mechanisms.

      Weaknesses:

      The primary limitation of this study lies in the spatial and cellular specificity of the recording techniques. The physiological data rely predominantly on unsorted multi-unit activity (MUA) recorded with lowchannel-count silicon probes. Because MUA aggregates signals from multiple neurons over a radius of approximately 50-100 µm (often wider than the typical septal width in mice), this approach makes it difficult to confidently isolate activity originating strictly from within septal domains. The manuscript would benefit from additional analyses to validate the spatial specificity of these recordings, such as systematically varying spike detection thresholds to test the robustness of domain attribution, as suggested by the reviewer. Furthermore, although the authors now appropriately frame their findings in the Elfn1 knockout mice as indirect evidence, it is worth emphasizing that the study lacks direct in vivo, cell-type-specific recordings and manipulations to more definitively test the proposed mechanism.

      We thank the reviewer for their thorough and constructive evaluation of the manuscript and for highlighting both the conceptual strengths of the study and its technical limitations. We agree that the spatial and cellular specificity of unsorted multi-unit recordings imposes inherent constraints on the interpretation of domain-specific activity, particularly given the narrow width of septal compartments in mice. As now clarified in the manuscript, we do not claim absolute cellular specificity of “septal” recordings but rather interpret them as septal-enriched populations. To directly address this concern, we performed additional threshold-based analysis demonstrating that the key domain-specific effects persist selectively in Layer 4 under stricter spike-detection criteria, supporting a local circuit origin of the critical findings. Further, the more stringent detection criteria (Suppl Fig 3A) collapse the divergence seen in Layer2/3 (Suppl Fig 4C), suggesting that this divergence arises in Layer 4, where SST+ interneuron distributions diverge between barrel and septa.

      We further agree that the Elfn1 knockout results provide indirect, rather than definitive, evidence for causal involvement of SST<sup>+</sup> interneurons and therefore intentionally frame the Elfn1–SST mechanism as a working model supported by converging anatomical, physiological, developmental, and genetic observations. We believe this explicitly moderated interpretation appropriately reflects the scope of the current data while establishing a clear conceptual framework and motivation for future studies employing cell-type-specific recordings and manipulations to directly test the proposed mechanism.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Major comments

      (1) Interpretation of "septal" recordings: The authors claim that the activity recorded from electrodes placed in the septa can be confidently attributed to septal neurons. In my previous review, I raised a major concern that such "septal" recordings likely include spikes from adjacent barrels, given the broad spatial resolution of MUA and the narrowness of the septa in the mouse S1. In fact, the intermediate properties observed in septal recordings from wild-type mice could be explained by a mixture of activity from principal and neighboring barrels-an interpretation that contrasts with the authors' conclusion. Upon reviewing the probe model used (A8x8-Edge-5mm-100-200-177), I noticed a discrepancy between the manufacturer's design and the schematic provided in the manuscript. The electrodes are located near the right edge of the probe rather than the center, suggesting that neurons in adjacent barrels could easily be sampled. In my previous review, I therefore suggested alternative approaches, such as calcium imaging, to more convincingly support the authors' claims. However, the revised manuscript does not include new experiments or additional analyses addressing this issue. Instead, the authors argue that using a high spike detection threshold (SD > 7.5) ensures that recorded activity originates from septal neurons, even though this value does not appear particularly conservative, as it was merely adopted from a previous study without justification in the present context. While I agree that a higher threshold may reduce contamination from distant sources, it does not guarantee that only septal neurons contribute to the signal. By nature, MUA reflects activity from multiple neurons within a radius of at least 50-100 µm. To more rigorously support the claim of spatial specificity, I strongly encourage the authors to reanalyze their existing dataset by systematically varying the spike detection threshold and quantifying how the properties and selectivity of detected units change. If neurons closer to the electrode indeed exhibit distinct domain-specific properties, they should become more prominent as the threshold increases. Such an analysis would strengthen the authors' interpretation and improve the manuscript's impact, even in the absence of new experimental data. Alternatively, the authors could revise their claims to acknowledge that the "septal" electrodes likely record from a population that includes septal neurons as well as neurons located at the periphery of principal and adjacent barrels.

      We agree with the reviewer that, by nature, MUA reflects the activity of multiple neurons within a spatial radius and that recordings obtained from electrodes positioned in the septa may include contributions from neurons located at the periphery of adjacent barrels. This concern is further compounded in superficial layers by probe geometry and orientation: given the narrow width of septa and the lateral spread of processes in upper cortical layers, recordings in L2/3 are inherently more susceptible to spatial mixing than those in layer 4, where columns are more compact and cytoarchitecturally distinct. To directly address these issues, we reanalyzed the same dataset using a more stringent spike detection threshold (SD > 9.5), compared to the originally reported SD > 7.5. Importantly, increasing the threshold selectively reduced or eliminated effects in L2/3, while the key domain-specific differences in L4 responses both the differential MW/SW dynamics in wild-type animals and their attenuation in Elfn1 knockout mice remained robust (the new Supp. Fig. 3. In the manuscript). This threshold-dependent dissociation is consistent with the interpretation that the critical effects reported in L4 arise from neurons spatially closer to the electrode and are less influenced by probe orientation or distant sources, rather than reflecting simple mixing of barrel signals. While this analysis does not claim absolute cellular exclusivity of septal neurons, it provides empirical support that the principal conclusions of the study are robust to stricter spatial sampling criteria and are particularly anchored in L4 circuitry. Accordingly, we now explicitly acknowledge in the manuscript that “septal” recordings likely represent septal-enriched populations rather than purely septal neurons, while emphasizing that the persistence of L4 effects under higher spike-detection thresholds strengthens the conclusion that local L4 inhibitory dynamics underlie the reported functional differences between barrel and septal domains.

      The greater sensitivity of L2/3 results to spike-detection threshold is also expected based on both anatomical considerations and probe geometry. Neurons in L2/3 possess broader horizontal dendritic and axonal arbors and participate in more laterally distributed integration across columns, making population signals in these layers intrinsically less spatially focal. As a result, conservative spike-detection criteria preferentially suppress L2/3 effects, particularly when recordings are obtained with probes optimized for deeper layers. Importantly, our two-photon calcium imaging data while similarly limited to L2/3 demonstrate that SST<sup>+</sup> interneurons show locally measurable and stimulus-specific responses at the single-cell level, providing independent support that L2/3 SST<sup>+</sup> activity is stimulus-modulated rather than artifactual. Taken together, these observations suggest that L2/3 results reflect more distributed and integrative network activity, whereas the L4 effects that persist across thresholds are more directly attributable to local circuit mechanisms. This layer-specific dissociation further supports our interpretation that the central findings of the study are driven by local inhibitory dynamics in L4, with L2/3 activity reflecting downstream integration rather than primary domain-specific computation.

      (2) Interpretation of the Elfn1 KO data: The authors' interpretation that Elfn1-dependent facilitation of SST<sup>+</sup> interneurons underlies the differential sensory responses between barrel and septal domains is conceptually appealing and supported by several converging, albeit indirect, lines of evidence. Specifically, the consistent correspondence among the differential activation of SST<sup>+</sup> neurons upon SWS and MWS, the late development of the barrel-septa differences in the responses to SWS and MWS, and the attenuation of this difference in Elfn1 knockout mice lends plausibility to the proposed model. However, it should be emphasized that the data remain indirect: the study does not include direct recordings of SST<sup>+</sup> neuronal activity from the knockout mice, nor cell-type- specific manipulations to demonstrate causal involvement. The mechanistic explanation therefore represents a hypothesis rather than definitive proof. That said, the authors clearly acknowledge these limitations in the Discussion and appropriately moderate their claims by presenting the SST-Elfn1 mechanism as a working model. Given this careful framing, the current manuscript can be regarded as a valuable conceptual contribution that advances our understanding of how inhibitory dynamics may shape temporal processing in the barrel cortex. Further experiments, as mentioned above, will be essential to test the causal role of this mechanism directly.

      We thank the reviewer for this thoughtful and balanced assessment. We fully agree that the Elfn1 knockout experiments provide indirect rather than definitive evidence for a causal role of SST<sup>+</sup> interneurons in mediating the domain-specific MW/SW response dynamics between barrels and septa For this reason, throughout the revised manuscript we explicitly frame the Elfn1–SST mechanism as a working model rather than a proven mechanism.

      Minor comments:

      The authors have adequately addressed my previous minor comments. In this round, I carefully reviewed the revised manuscript and identified several issues related to references. I would also like to add a brief comment regarding the Discussion section:

      (1) Stachniak et al., 2021 is included in the reference list but is not cited anywhere in the main text. Please either remove this entry or cite it appropriately in the manuscript.

      Removed.

      (2) Yamashita et al., 2018 is cited in the main text (Line 767), but it is not included in the reference list.

      Fixed.

      (3) Sylwestrak and Ghosh, 2012 is cited at Line 261 and Line 270, but likewise absent from the reference list.

      Fixed.

      (4) At Line 497, Chen et al., 2015 is cited, but, the appropriate and original reference would be Chen et al., 2013 (PMID: 23792559), which should either replace or precede the 2015 citation.

      Added.

      (5) At Line 221, El-Boustani et al., 2018 is cited. However, this study is based on the visual cortex, whereas the manuscript concerns the barrel cortex. A more relevant citation (e.g., Lefort et al., 2009 [PMID: 19186171]) would better support the discussion of cellular organization in the barrel cortex. Please consider updating the citation.

      Thank you for this suggestion. We agree with the reviewer and now we have changed El-Boustani with Lefort et al. 2009 as suggested by the reviewer.

      (6) Furthermore, Chakrabarti & Alloway (2006) performed tracer-based mapping of projections from barrel and septal columns in rat S1 and similarly suggested differential organization of M1- and S2projection neurons in the barrel and septal regions.

      Although the current study thoroughly analyzed the layer-specificity of the location of these projection neurons, the lack of explicit discussion of this relevant prior work is a notable omission.

      The authors should incorporate a comparison with these results to better contextualize their findings.

      The following text is added to the discussion: “Our retrograde labeling data supports and expands on previous work proposing similar models (Alloway, 2008; Chakrabarti and Alloway, 2006).”

    1. Author response:

      (1) Introduction & Roadmap

      We are grateful to the Reviewers for engaging with outstanding questions relating to our findings’ connections to multiple subdisciplines of cognitive neuroscience. Noting that Reviewers 1 and 2 interpreted our findings differently, we welcome the opportunity to engage in what Reviewer 1 characterised as “an interesting debate”. To promote a shared understanding and discussion of our findings, we have organised our response to address more technical comments first.

      Our provisional response is organised as follows: Section 2 addresses selected technical comments relating to our Results. Section 3 addresses comments related to the design of our behavioural paradigm. Section 4 focuses on the broader interpretation of our findings. Section 5 concludes our provisional response with potential future directions and a summary of the significance of our findings.

      (2) Selected technical comments related to our Results

      We apologise to Reviewer 2 for the confusion in relation to the meaning of “attended” and “unattended” trials. What we said was “Positive Pref values indicate a higher response rate to the contralateral side than the ipsilateral side (relative to the electrode)” (Figure 3c caption), “we indexed all contralateral whisker vibrations according to their associated Perf and Pref” (Results text), and “we divided trials into (contralaterally) attended (Pref<sub>C/L</sub>: Pref>0) and unattended (Pref<sub>I/L</sub>: Pref<0) groups” (Results text). We can confirm that we defined an “unattended trial” (Pref<0) as a contralateral stimulus trial in the centre of an epoch (10-15 trials) within which the mouse responded (licked) more frequently to ipsilateral stimuli. Critically, we did not define an unattended trial as an ipsilateral stimulus trial. Furthermore, attention thus defined (i.e. Pref>0) can vary independently of the whisker stimulus associated with rewards. Indeed, while we initially did not include this result in our paper for the sake of brevity, even unrewarded “attended” trials (Pref>0) evoked significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials. We note that this is an analysis suggested by Reviewer 1, and we will include and discuss this result in our revised manuscript (e.g. in relation to literature suggested by Reviewer 2). For additional clarity, we use “Performance” (Perf) in relation to overall stimulus detection, consistent with the analysis of Lee et al. (2020), which found this measure was correlated with pupil diameter in a vibrissal target detection task.

      We thank Reviewer 1 for noticing that the axes on Figure 3e should be labelled “Pref>0” (Y axis) and “Pref<0” (X axis), as suggested by the figure caption. We will correct this in our revised submission. The yellow point on Fig 3e shows the unit from Fig 3d, while the yellow line in Fig 3e shows the magnitude of that unit’s (non-normalised) gain modulation. While this is alluded to in the Results text (“The example unit in Figure 3d is in the 93rd percentile of units for raw modulation depth (ΔHits(attended – unattended) = 3.3 spikes/second; yellow line in Fig.3e)”, this should be explained in the Figure caption, and it will be in our revised manuscript. We would also like to clarify that Figures 3g–3h display results for all units, not just the top 25%. We agree this is not sufficiently clear and we will rectify this in our revised manuscript. Addressing Reviewer 2, while we acknowledge that mice responded less to both stimuli in the second block, they also meaningfully adjusted their behaviour to the reversal in reward contingencies: their responses to the previously rewarded stimulus reduced significantly more than those to the previously unrewarded stimulus.

      (3) Design of the behavioural paradigm

      We made a deliberate design choice to maximise the ecological validity of our behavioural paradigm, and note that there are advantages to doing so. For example, our paradigm can be used to show that even unrewarded “attended” trials (Pref>0) evoke significantly greater neuronal responses than unrewarded “unattended” (Pref<0) trials (see Section 2, above). Indeed, it is precisely this finding that makes our paradigm uniquely suited to the investigation of value-driven attentional capture (Anderson et al., 2011): in this instance attention directed to stimuli that are no longer rewarded despite equal availability of rewarded stimuli. This finding also demonstrates that our paradigm dissociates attention from stimulus-reward contingency at least as well as other paradigms which have been successfully used to study spatial attention in mice. As noted in Section 2, we will discuss this result in relation to other relevant research (e.g. Ramamurthy et al., 2025) in our revised manuscript.

      Briefly, the direct manipulation of reward contingencies is one of two noteworthy methodological distinctions between our own paradigm and that of Ramamurthy and colleagues (2025). The task of Ramamurthy et al. (2025) associated all whisker stimuli with rewards and delivered stimuli to different whiskers on a single whisker pad. These methodological distinctions may have reduced the relevance of the spatial differences between stimuli to the mice undertaking the task. Indeed, it is not certain that a mouse would treat the unilateral variation in whisker stimulation Ramamurthy and colleagues delivered as primarily spatial or featural. The psychophysical and neural differences between spatial and featural attention in humans suggest dissociable underlying mechanisms, and the same may be true in mice. Thus, our own paradigm may more effectively isolate spatial attention from featural attention. Conversely, to the extent that the findings of Ramamurthy and colleagues do reflect spatial attention, our combined findings and paradigms help elucidate the associated mechanisms across spatial scales in mice.

      We acknowledge that spatial cueing is well-suited to isolating the effects of covert attention from other forms of attention. However, it should be noted that spatial cueing in rodents is subject to its own challenges, including limitations in trial numbers due to the required manipulation of stimulus intensity (Reynolds et al., 2000; Herrmann et al., 2010), cue validity and associated trial probabilities (Peterson & Gibson, 2011; Girardi et al., 2013). Such experiments are further complicated by the duration and efficacy of training (i.e. the number of mice that learn the task; Wang & Krauzlis, 2018; Hu & Dan, 2022). It is also worth noting that trial probability manipulations introduce the same limitation in trial numbers with block-type attention tasks (You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022).

      While there are clear differences between our own paradigm and those mentioned above, there are also important similarities. First, these tasks are all goal-directed, stimulus-driven, and reliant on learned task contingencies (e.g. Peterson & Gibson, 2011; Girardi et al., 2013). Furthermore, these paradigms are all operant conditioning protocols which leverage learned stimulus-reward contingencies to train attention-related behaviours in mice. A noteworthy similarity between our findings and those of authors using block-type attention tasks in particular (e.g. You & Mysore, 2020; Kanamori & Mrsic-Flogel, 2022) is the observation of apparent attentional biases in behavioural responses independent of the experimental manipulations (i.e. stimulus probability / reward contingency).

      (4) Comments relating to the broader interpretation and discussion of our findings

      Fundamentally, attention involves dedicating limited processing resources to some stimulus events at the expense of others. The design of our behavioural paradigm was informed by existing literature on spatial attention in humans, non-human primates, and mice. Our choice of behavioural and neuronal measures as proxies for attention in mice is consistent with this literature. It is technically possible “an animal could pay ‘more’ (rather than less) attention to the stimulus delivered on the unrewarded side, to make sure it suppresses the incorrect response”, but this seems unlikely given what is known about how attention is typically allocated in such tasks, based on the previously mentioned literature.

      With respect to the interpretation and discussion of our findings, Reviewer 1 describes them as “a behavioral phenomenon that can reasonably be interpreted as spatial attentional capture” but suggests they do not clearly distinguish whether this attentional capture is covert or overt. We respectfully disagree for three reasons. First, as discussed in our paper, whisker motion during detection tasks has consistently been associated with reduced detection performance (Ollerenshaw et al., 2012; Kyriakatos et al., 2017; Vandevelde et al., 2023), suggesting that a “receptive” strategy (Diamond & Arabzadeh, 2013) of whisker immobilisation is more applicable to the current data than a “generative” strategy of asymmetric whisker movement (O'Connor et al., 2010; Dominiak et al., 2019). Second, if our behavioural and neuronal findings were due to the mice moving their whiskers to maximise contact with the meshes, we would expect increased evoked neuronal responses to be associated with greater Perf, not just with greater Pref. This pattern was not observed. Of course, the mice might have employed different whisker movement strategies during epochs of high Pref and Perf, but this seems unlikely and is not a parsimonious explanation for our findings. Third, as noted in the Methods section of the paper, we deliberately positioned the meshes close to the base of the whiskers, limiting the impact of whisker movements on stimulus detectability and the incentive to make them.

      In contrast, Reviewer 2 questions the interpretation of our findings as evidence of spatial attention and suggests they might reflect working memory instead. Current research suggests attention and working memory are intimately related integrative brain functions. Indeed, some researchers have even proposed that working memory might be a form of internally directed attention (Awh & Jonides, 2001; Chun, 2011; Gazzaley & Nobre, 2012; Kiyonaga & Egner, 2013; or vice versa: Libedinsky & Fernandez, 2019). Consistent with the comments of Reviewer 2, more recent work seems to emphasise the coordination of attention and working memory (e.g. Joe & Kim, 2023; Zhu et al., 2026; for reviews see Huynh Cong & Kerzel, 2021; van Ede & Nobre, 2023), along with shared mechanisms (Kiyonaga et al., 2021; Panichello & Buschman, 2021), and nuanced dissociations (Liu et al., 2025). Attention is difficult to dissociate from working memory partly because there are multiple definitions (and/or types) of attention. We did not discuss the various definitions and/or forms of attention at length in our paper, but we will briefly discuss this in the revised manuscript.

      The “interesting debate” to which Reviewer 1 refers could also be described as vigorous, despite approximately three decades of research. This debate broadly relates to the degree to which attentional control is driven by exogenous (e.g. colour contrast) versus endogenous factors (e.g. the focus of spatial attention, see Fig.2 in Belopolsky et al., 2007; see also: Liesefeld & Mueller, 2020; Manini et al., 2021; Beffara et al., 2022), and the degree to which this is a function of experimental context. The review article by Luck et al. (2021) entitled “Progress toward resolving the attentional capture debate” provides a striking illustration of this debate, as do the twenty-two commentaries (and three commentary responses) associated with it. Admittedly, this debate largely revolves around human attention experiments, and human cognition may be more complex than mouse cognition. However, the complexity of human cognition may also be easier to study and appreciate because complex behavioural experiments can be explained to, understood, and performed by human participants with relative ease.

      (5) Comments relating to future directions and the significance of our findings

      The complexity of the attentional capture debate underscores the importance of developing accessible and scalable animal experiments which can be used to provide mechanistic insights. If the human attention literature is any indication, a diversity of rodent experimental paradigms will be necessary to thoroughly map the neuronal implementation of spatial attention. Returning to our paradigm, Reviewer 1 noted that valuable insights into the mechanisms of vibrissal spatial attention might be obtained from comparing the magnitude of attentional modulation we observed between putative regular and fast-spiking categories of units, and between units located in different cortical layers. We agree it is important to understand spatial attention with cell-type and circuit (including laminar) specificity. However, because we could not persuasively cluster our units based on waveform width, and because of the lack of histological data, segregating units on the basis of such variables is not feasible. Despite our assertion that our findings reflect the effects of covert attention (contra Reviewer 1), we agree that future experiments will be required to conclusively rule out overt attention. Noting the proximity of the meshes to the base of the whiskers in our paradigm, and the difficulty of tracking whiskers in this context, Botulinum toxin injections (as in Ramamurthy et al., 2025) might be a means of achieving this.

      The above notwithstanding, our findings provide multiple contributions to the literature on spatial attention (and perhaps working memory). We detected significant attentional gain modulation across a population of 1461 responsive units. While the gain modulation exhibited by the median unit was modest (albeit statistically significant), the top 25% of responsive units showed a ~12% response modulation (relative to firing rate range for each unit), and ~21% of responsive units were suppressed by the average vibrissal stimulus in the unattended state. Our experimental framework offers an accessible platform for future studies leveraging genetic and circuit-level interventions to dissect the cell-type specific mechanisms of spatial attention. Our work is timely, noting the recent focus of human research on the nexus of attention, selection history, and valence (e.g. Serences, 2008; Della Libera & Chelazzi, 2009; Della Libera et al., 2011; van den Berg et al., 2014; Kim & Anderson, 2019, 2023). Our work is also uniquely poised to stimulate new interdisciplinary research into the circuit mechanisms of value-driven attentional capture, with translational relevance to psychopathologies such as ADHD, addiction, and depression; where value-driven attentional capture is altered (for a review see Anderson, 2021).

      References

      Anderson, B. A. (2021). Relating value-driven attention to psychopathology. Curr Opin Psychol, 39, 48-54. https://doi.org/10.1016/j.copsyc.2020.07.010

      Anderson, B. A., Laurent, P. A., & Yantis, S. (2011). Value-driven attentional capture. Proceedings of the National Academy of Sciences of the United States of America, 108(25), 10367-10371. https://doi.org/10.1073/pnas.1104047108

      Awh, E., & Jonides, J. (2001). Overlapping mechanisms of attention and spatial working memory. Trends Cogn Sci, 5(3), 119-126. https://doi.org/10.1016/s1364-6613(00)01593-x

      Beffara, B., Hadj-Bouziane, F., Ben Hamed, S., Boehler, C. N., Chelazzi, L., Santandrea, E., & Macaluso, E. (2022). Dynamic causal interactions between occipital and parietal cortex explain how endogenous spatial attention and stimulus-driven salience jointly shape the distribution of processing priorities in 2D visual space. Neuroimage, 255. https://doi.org/10.1016/j.neuroimage.2022.119206

      Belopolsky, A. V., Zwaan, L., Theeuwes, J., & Kramer, A. F. (2007). The size of an attentional window modulates attentional capture by color singletons. Psychonomic Bulletin & Review, 14(5), 934-938. https://doi.org/10.3758/Bf03194124

      Chun, M. M. (2011). Visual working memory as visual attention sustained internally over time. Neuropsychologia, 49(6), 1407-1409. https://doi.org/10.1016/j.neuropsychologia.2011.01.029

      Della Libera, C., & Chelazzi, L. (2009). Learning to Attend and to Ignore Is a Matter of Gains and Losses. Psychological Science, 20(6), 778-784. https://doi.org/10.1111/j.1467-9280.2009.02360.x

      Della Libera, C., Perlato, A., & Chelazzi, L. (2011). Dissociable Effects of Reward on Attentional Learning: From Passive Associations to Active Monitoring. PLoS One, 6(4). https://doi.org/10.1371/journal.pone.0019460

      Diamond, M. E., & Arabzadeh, E. (2013). Whisker sensory system - from receptor to decision. Prog Neurobiol, 103, 28-40. https://doi.org/10.1016/j.pneurobio.2012.05.013

      Dominiak, S. E., Nashaat, M. A., Sehara, K., Oraby, H., Larkum, M. E., & Sachdev, R. N. S. (2019). Whisking Asymmetry Signals Motor Preparation and the Behavioral State of Mice. J Neurosci, 39(49), 9818-9830. https://doi.org/10.1523/JNEUROSCI.1809-19.2019

      Gazzaley, A., & Nobre, A. C. (2012). Top-down modulation: bridging selective attention and working memory. Trends Cogn Sci, 16(2), 129-135. https://doi.org/10.1016/j.tics.2011.11.014

      Girardi, G., Antonucci, G., & Nico, D. (2013). Cueing spatial attention through timing and probability. Cortex, 49(1), 211-221. https://doi.org/10.1016/j.cortex.2011.08.010

      Herrmann, K., Montaser-Kouhsari, L., Carrasco, M., & Heeger, D. J. (2010). When size matters: attention affects performance by contrast or response gain. Nat Neurosci, 13(12), 1554-1559. https://doi.org/10.1038/nn.2669

      Hu, F., & Dan, Y. (2022). An inferior-superior colliculus circuit controls auditory cue-directed visual spatial attention. Neuron, 110(1), 109-119 e103. https://doi.org/10.1016/j.neuron.2021.10.004

      Huynh Cong, S., & Kerzel, D. (2021). Allocation of resources in working memory: Theoretical and empirical implications for visual search. Psychon Bull Rev, 28(4), 1093-1111. https://doi.org/10.3758/s13423-021-01881-5

      Joe, J., & Kim, M. S. (2023). Spatial Attention in Visual Working Memory Strengthens Feature-Location Binding. Vision (Basel), 7(4). https://doi.org/10.3390/vision7040079

      Kanamori, T., & Mrsic-Flogel, T. D. (2022). Independent response modulation of visual cortical neurons by attentional and behavioral states. Neuron, 110(23), 3907-3918 e3906. https://doi.org/10.1016/j.neuron.2022.08.028

      Kim, H., & Anderson, B. A. (2019). Dissociable neural mechanisms underlie value-driven and selection-driven attentional capture. Brain Research, 1708, 109-115. https://doi.org/10.1016/j.brainres.2018.11.026

      Kim, H., & Anderson, B. A. (2023). Primary Rewards and Aversive Outcomes Have Comparable Effects on Attentional Bias. Behavioral Neuroscience, 137(2), 89-94. https://doi.org/10.1037/bne0000543

      Kiyonaga, A., & Egner, T. (2013). Working memory as internal attention: toward an integrative account of internal and external selection processes. Psychon Bull Rev, 20(2), 228-242. https://doi.org/10.3758/s13423-012-0359-y

      Kiyonaga, A., Powers, J. P., Chiu, Y. C., & Egner, T. (2021). Hemisphere-specific Parietal Contributions to the Interplay between Working Memory and Attention. J Cogn Neurosci, 33(8), 1428-1441. https://doi.org/10.1162/jocn_a_01740

      Kyriakatos, A., Sadashivaiah, V., Zhang, Y., Motta, A., Auffret, M., & Petersen, C. C. (2017). Voltage-sensitive dye imaging of mouse neocortex during a whisker detection task. Neurophotonics, 4(3), 031204. https://doi.org/10.1117/1.NPh.4.3.031204

      Lee, C. C. Y., Kheradpezhouh, E., Diamond, M. E., & Arabzadeh, E. (2020). State-Dependent Changes in Perception and Coding in the Mouse Somatosensory Cortex. Cell Rep, 32(13), 108197. https://doi.org/10.1016/j.celrep.2020.108197

      Libedinsky, C. D., & Fernandez, P. F. (2019). Graded Memory: A Cognitive Category to Replace Spatial Sustained Attention and Working Memory
 Yale J Biol Med, 92(1), 121-125. https://www.ncbi.nlm.nih.gov/pubmed/30923479

      Liesefeld, H. R., & Mueller, H. J. (2020). A theoretical attempt to revive the serial/parallel-search dichotomy. Attention Perception & Psychophysics, 82(1), 228-245. https://doi.org/10.3758/s13414-019-01819-z

      Liu, Y., Fu, Y., Tang, E., Wu, H., Han, J., Xie, M., Zhang, Y., Peng, B., Huang, J., Liu, H., Chen, H., & Qin, P. (2025). Neural dissociation of attention and working memory through inhibitory control. Nat Commun, 17(1), 22. https://doi.org/10.1038/s41467-025-66553-7

      Luck, S. J., Gaspelin, N., Folk, C. L., Remington, R. W., & Theeuwes, J. (2021). Progress toward resolving the attentional capture debate. Visual Cognition, 29(1), 1-21. https://doi.org/10.1080/13506285.2020.1848949

      Manini, G., Botta, F., Martin-Arevalo, E., Ferrari, V., & Lupianez, J. (2021). Attentional Capture From Inside vs. Outside the Attentional Focus. Frontiers in Psychology, 12. https://doi.org/10.3389/fpsyg.2021.758747

      O'Connor, D. H., Clack, N. G., Huber, D., Komiyama, T., Myers, E. W., & Svoboda, K. (2010). Vibrissa-based object localization in head-fixed mice. J Neurosci, 30(5), 1947-1967. https://doi.org/10.1523/JNEUROSCI.3762-09.2010

      Ollerenshaw, D. R., Bari, B. A., Millard, D. C., Orr, L. E., Wang, Q., & Stanley, G. B. (2012). Detection of tactile inputs in the rat vibrissa pathway. J Neurophysiol, 108(2), 479-490. https://doi.org/10.1152/jn.00004.2012

      Panichello, M. F., & Buschman, T. J. (2021). Shared mechanisms underlie the control of working memory and attention. Nature, 592(7855), 601-605. https://doi.org/10.1038/s41586-021-03390-w

      Peterson, S. A., & Gibson, T. N. (2011). Implicit attentional orienting in a target detection task with central cues. Conscious Cogn, 20(4), 1532-1547. https://doi.org/10.1016/j.concog.2011.07.004

      Ramamurthy, D. L., Rodriguez, L., Cen, C., Li, S., Chen, A., & Feldman, D. E. (2025). Reward history guides focal attention in whisker somatosensory cortex. Nat Commun, 16(1), 5580. https://doi.org/10.1038/s41467-025-60592-w

      Reynolds, J. H., Pasternak, T., & Desimone, R. (2000). Attention increases sensitivity of V4 neurons. Neuron, 26(3), 703-714. https://doi.org/10.1016/s0896-6273(00)81206-4

      Serences, J. T. (2008). Value-Based Modulations in Human Visual Cortex. Neuron, 60(6), 1169-1181. https://doi.org/10.1016/j.neuron.2008.10.051

      van den Berg, B., Krebs, R. M., Lorist, M. M., & Woldorff, M. G. (2014). Utilization of reward-prospect enhances preparatory attention and reduces stimulus conflict. Cognitive Affective & Behavioral Neuroscience, 14(2), 561-577. https://doi.org/10.3758/s13415-014-0281-z

      van Ede, F., & Nobre, A. C. (2023). Turning Attention Inside Out: How Working Memory Serves Behavior. Annu Rev Psychol, 74, 137-165. https://doi.org/10.1146/annurev-psych-021422-041757

      Vandevelde, J. R., Yang, J. W., Albrecht, S., Lam, H., Kaufmann, P., Luhmann, H. J., & Stuttgen, M. C. (2023). Layer- and cell-type-specific differences in neural activity in mouse barrel cortex during a whisker detection task. Cereb Cortex, 33(4), 1361-1382. https://doi.org/10.1093/cercor/bhac141

      Wang, L., & Krauzlis, R. J. (2018). Visual Selective Attention in Mice. Curr Biol, 28(5), 676-685 e674. https://doi.org/10.1016/j.cub.2018.01.038

      You, W. K., & Mysore, S. P. (2020). Endogenous and exogenous control of visuospatial selective attention in freely behaving mice. Nat Commun, 11(1), 1986. https://doi.org/10.1038/s41467-020-15909-2

      Zhu, P., Guan, C., Fu, Y., Shen, M., & Chen, H. (2026). Working memory encoding of attended information is adaptive to future relevance. J Exp Psychol Learn Mem Cogn. https://doi.org/10.1037/xlm0001582

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors present comprehensive experimental observations and a theoretical framework to explain the heterogeneous behaviour of sarcomeres in cardiomyocytes. They show that a stochastic component exists in their contractile activity, which may act as a feedback mechanism regulating physiological function.

      Strengths:

      Experiments and data analysis are robust and valid. The rigorous statistical analysis and unbiased methods enable the authors to draw well-supported conclusions that go beyond the existing literature. Their outcomes inform about cellular activity at the individual level and the authors explain how the transient dynamics of single sarcomeres are governed by a force-velocity relationship and lead to the complex contractile patterns. The similarity of the results to the study cited in [24] demonstrates the validity of the in vitro setup for answering these questions and the feasibility of such in-vitro systems to extend our knowledge of out-of-equilibrium dynamics in cardiac cells.

      Very interesting the suggestion that the interplay between intrinsic fluctuations and the dynamic instability are part of a feedback mechanism for maintaining structural and functional homeostasis.

      The addition of the theoretical model and the new text of the manuscript improves the clarity of the study.

      Reviewer #2 (Public review):

      Summary:

      Sarcomeres, the contractile units of skeletal and cardiac muscle, contract in a concerted fashion to power myofibril and thus muscle fiber contraction.

      Muscle fiber contraction depends on the stiffness of the elastic substrate of the cell, yet it is not known how this dependence emerges from the collective dynamics of sarcomeres. Here, the authors analyze contraction time series of individual sarcomeres using live imaging of fluorescently labeled cardiomyocytes cultured on elastic substrates of different stiffness. They find that a reduced collective contractility of muscle fibers on unphysiologically stiff substrates is partially explained by a lack of synchronization in the contraction of individual sarcomeres.

      This lack of synchronization is at least partially stochastic, consistent with the notion of a tug-of-war between sarcomeres on stiff sarcomeres. A particular irregularity of sarcomere contraction cycles is 'popping', the extension of sarcomers beyond their rest length. The statistics of 'popping' suggest that this is a purely random process.

      Strengths:

      This study thus marks an important shift of perspective from whole-cell analysis towards an understanding the collective dynamics of coupled, stochastic sarcomeres.

      Reviewer #3 (Public review):

      The manuscript of Haertter and coworkers studied the variation of the length of a single sarcomere and the response of microfibrils made by sarcomeres of cardiomyocytes on soft gel substrates of varying stiffness.

      The measurements at the level of a single sarcomere are an important new result of this manuscript. They are done by combining the labeling of the sarcomeres z line using genetic manipulation and a sophisticated tracking program using machine learning. This single sarcomere analysis shows strong heterogeneities of the sarcomeres that can show fast oscillations not synchronized with the average behavior of the cell and what the authors call popping eveents which are large amplitude oscillations. Another important result is the fact that cardiomyocyte contractility decreases with the substrate stiffness, although the properties of single sarcomeres do not seem to depend on substrate stiffness.

      The authors suggest that the cardiomyocyte cell behavior is dominated by sarcomere heterogeneity. They show that the heterogeneity between sarcomere is stochastic and that the contribution of static heterogeneity (such as composition differences between sarcomeres) is small.

      Strengths:

      All the results are, to my knowledge, new and original. The authors also made a theoretical model where each sarcomere is described by a Langevin equation based on a non-linear coupling between force and velocity of the sarcomeres. This model accounts well for the experimental results including the observation of what the authors call popping events.

      We thank you and the reviewers for the positive evaluation of our revised manuscript.

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Origin of the 3-Hz oscillation and required model extension. These oscillations are reproduced by our model, and their origin is already discussed in the manuscript (see lines 403–406).

      (2) Inclusion of all 5085 LOIs vs. the selected 2321. We have expanded the explanation of the LOI selection criteria in the manuscript and clarified that the main conclusions are not sensitive to this choice (lines 161-166)

      (3) Fig. 3G caption — popping rate. The caption has been updated to clarify the units and normalization. 

      (4) Fig. 4G — "Length x" vs. ΔL. Notation corrected for consistency.

      (5) Fig. 4G — gray data points. Confirmed: these represent the mean, and the caption has been updated accordingly.

      (6) Relation of k_l to the true substrate stiffness. We have added the following clarification: "The model evaluation compared the distributions of sarcomere length changes and velocities from simulations with representative experimental LOIs from substrates (5, 15, and 85 kPa, mapped to k_l = 0.5, 1.5 and 8.5 in our 1-D model; k_l is unitless, so only the ratios between values are meaningful — rescaling k_l leaves model output unchanged under correspondingly rescaled parameters) covering the full range of mechanical loads." (lines 365-369)

      (7) Could a simpler model fit the data? The cubic polynomial in Eq. (3) was deliberately chosen as a generalist ansatz rather than imposed: its coefficients were obtained by data-driven inference via Differential Evolution, and if lower-order terms within this family had sufficed, the higher-order coefficients would have been driven toward zero. The inferred nonmonotonic force–velocity relation has two extrema separated by an unstable negative-slope branch, which sets a lower bound on the polynomial order — a linear F–v is monotonic and a quadratic admits only a single extremum, so cubic is the minimum polynomial order capable of producing the observed shape. Furthermore, the qualitative phenomena we report — popping events, dynamic instability, and stochastic heterogeneity — cannot arise from any monotonic force–velocity relation, as discussed in the section on the non-monotonic instability. With 10 parameters covering complex contractile dynamics at the individual sarcomere and myofibril level across different substrate stiffnesses, the present model is parsimonious within the family of polynomial force–velocity ansätze; we have not exhaustively searched alternative non-polynomial functional families, but any such alternative would still need to reproduce the same non-monotonic shape that the data require.

      (8) Lines 497–507 in the Discussion. On reflection, we feel these lines provide useful context for the broader interpretation and would prefer to retain them.

      (9) Line 331 — motivation of Eq. (3). We have added citations to prior work motivating this form of the equation for the broader readership.

      (10) Line 427 — "scaled". Corrected.

      Reviewer #3 (Recommendations for the authors):

      We thank the reviewer for the recommendation of a theoretical appendix. The full model code, with the formulation and implementation documented in detail, is publicly available in our GitHub repository accompanying the paper, which we believe provides a complete reference for readers wishing to explore the model further. We therefore feel an additional appendix is not necessary within the scope of this revision.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We thank the Editors for the positive assessment on our manuscript. We also thank the Reviewers for their positive remarks and constructive comments. Based on the Reviewers’ feedback, we have conducted additional experiments and provided supporting data to address Reviewers’ comments. Particularly, we provided quantitative measurement for rotational polarity of ependymal cells in Agbl5<sup>M1/M1</sup> mutants and assessed the microtubule polarization. We quantified the intensity of apical actin network in ependymal cells to strength the role of CCP5 in organizing actin network. Using scanning electron microscopy, we demonstrated the affected polarity of trachea multicilia in Agbl5<sup>M1/M1</sup>. We co-immunostained ependymal cilia with GT335 and acetylated tubulin to address the effects on their length in cilia in the mutant. We assessed the presence and length of primary cilia in ependymal cell progenitors to identify their potential contribution to the defective polarity in Agbl5<sup>M1/M1</sup> ependymal cells. We feel that these revisions have much strengthened this MS.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Dad et al. explored the roles of cytosolic carboxypeptidase 5(CCP5)in the development of ependymal multicilia in the brain. CCP family are erasers of polyglutamylation of ciliary-axoneme microtubules. The authors generated a new mutant mouse of Agbl5 gene, which encodes CCP5, with deletion of its N-terminus and partial carboxypeptidase (CP) domain (named AGBL5M1/M1).

      Strengths:

      The mutant mice revealed lethal hydrocephalus due to degeneration of ependymal multicilia. Interestingly, this is in contrast with the phenotype of Agbl5 mutants with disruption solely in the CP domain of CCP5 (named AGBL5M2/M2) that did not develop hydrocephalus despite increased glutamylation levels in ependymal cilia as observed for AGBL5M1/M1 mutants. The study has been well-performed and the findings suggest a unique function of the N-domain of CCP5 in ependymal multicilia stability.

      Weaknesses:

      The content of this article is relatively descriptive and lacks molecular insights.

      We thank the Reviewer’s positive comments. To address the molecular insights of the dysregulated planar cell polarity (PCP) in Agbl5<sup>M1/M1</sup> ependyma, we have conducted additional experiments to assess the microtubule polarization in ependymal cells (Figure 7O-P). We quantified the intensity of actin networks around BB patches to better understand how it is affected in the ependyma of the mutants and contributes to the dispersion of BBs (Figure 4M-N), (Please see Recommendations for the authors).

      We also assessed trachea multicilia in Agbl5<sup>M1/M1</sup> mutants using SEM and found that the polarity of trachea multicilia was affected as well (Figure S2).

      Reviewer #2 (Public review):

      Summary:

      This study analyzed the consequences of Agbl5 mutation on ependymal cell development and function. The authors first characterize their mutant mouse line reporting a reduced lifespand and severe hydrocephalus. Next, they report a defect in ependymal cell cilia number and motility. They provide evidence for impaired basal body organisation and cilia glutamylation.

      Strengths:

      Description of a mutant mouse which implicates Cytosolic Carboxypeptidase 5 (the product of Agbl5 gene) for proper ependymal cells.

      Weaknesses:

      Description of phenotype is incomplete:

      We thank the Reviewer’s constructive comments. We have performed additional quantitative analysis of the phenotypes in Agbl5<sup>M1/M1</sup> that we feel strengthen this study.

      Figure 3G - the sequence from the movie is not really informative. Providing beating frequencies as quantification of the data would be more informative.

      We have provided the beating frequency as well as the mean vector length of cilia beating directions (that reflects the coordination of cilia) in Figure 3H and 3I respectively in the revised manuscript.

      Figure 3 - the quantification of actin network would strengthen the message.

      We agree with the Reviewers. We have quantified the total intensity of actin around BBs and the actin intensity normalized to signals of the BB marker (CEP164). The data have been provided in Figure 4M and 4N respectively. The quantitative analysis showed that both the total intensity of apical actin network and the intensity of F-actin per BB are reduced in Agbl5<sup>M1/M1</sup> ependymal cells compared to that in wild-type mice, suggesting that CCP5 is involved in organizing actin network around BB. This analysis certainly improves the clarity of this message.

      Lines 219 -220 - the authors conclude «Taken together, in Agbl5M1/M1 ependymal cells, the expression of genes promoting multiciliogenesis were not impaired but certain proteins associated with differentiated ependymal cells are not properly expressed». However, they do not assess gene but protein expression (IF). In addition, their quantification shows differences in the number of FoxJ1 positive cells which indeed is an impaired expression.

      We will clarify this statement and emphasize the number of FoxJ1-positive cells.

      Microtubules are involved in the local organization of ciliary basal bodies (see Werner et al., Vladar et al.,2011; Boutin et al., 2014). It would be interesting for the authors to check whether the subapical network of microtubules is glutamylated or not during ependymal cell differentiation and how this network is affected in their mutants.

      We thank the Reviewer’s constructive comments. We conducted an immunostaining on whole-mount lateral walls of lateral ventricles for GT335 and Centrin1, the position of the latter being used to localize the subapical layer. While the GT335 signal in multicilia is increased in Agbl5<sup>M1/M1</sup> ependyma (Figure S8E), its signals underneath BBs are not much different between the mutant and wild-type (Please see Figure S8C, D, G, H).

      Showing the data mentioned in the discussion on Cep110 would be a nice addition to the paper.

      These data have been provided in Supplementary Figure S9.

      Line 354: "The latter serves as a component of tissue polarity that is required for asymmetric PCP protein localization in each cell (Boutin et al., 2014; Vladar et al., 2012)." The cited reference did not demonstrate that this microtubule network is required for asymmetric PCP localization.

      We thank the Reviewer for critical reading. The cited reference (Bountin et al., 2014) has been removed.

      Reviewer #3 (Public review):

      Summary:

      The authors developed a new Agbl5 KO allele, extending the deletion to the N-terminus of CCP5 to explore its function in mouse ependymal cells.

      Strengths:

      They show that the KO mice exhibit severe hydrocephalus due to disorganized and mislocated basal bodies. Additionally, they present evidence of both impaired beating coordination and a reduction in ciliary beating.

      Weaknesses:

      The manuscript is well-written but lacks specific interpretations of the results presented. Further experiments are needed to be fully convincing.

      We thank the Reviewer’s comments. We have performed further analysis and conducted additional experiments to strengthen this study.

      (1) We have quantified the intensity of actin staining around BB patches and its intensity relative to the number of BBs to assess to which extent the actin networks in Agbl5<sup>M1/M1</sup> ependymal cells are affected (please refer to the above response to the comments of Reviewer 2#). The results were shown in Figure 4M-N.

      (2) We Co-stained tdTomato with an ependymal cell-specific markers to strengthen the expression of Agbl5 in ependymal cells (please see Figure 6C-E).

      (3) We have conducted co-immunostaining of GT335 and Ac-Tub and compared the length of their signals in ependymal multicilia between WT and Agbl5<sup>M1/M1</sup> mice (please see Figure 6O, P, R, S).

      (4) We quantified the area of ependymal cells in the wild-type and Agbl5<sup>M1/M1</sup> mice. Indeed, the area of ependymal cells is increased in the mutants. However, the primary cilia are present in the ependymal cell progenitors of Agbl5<sup>M1/M1</sup> mice and have similar length with that in the wild-type (Please see Figure 7M, N and our response to this point below).

      (5) We performed additional analysis to address the affected rotational polarity in the Agbl5<sup>M1/M1</sup> mutant mice (please see Figure 3I, Figure 7E).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) The authors showed that the actin networks were severely affected, leading to impaired stability of basal bodies and that the intensity and length of acetylated tubulin signal in the multicilia were dramatically reduced in AGBL5M1/M1mutant mice (Figures 3 and 5). Data also suggested the dysregulation of planar cell polarity. Are expression and localization of other planar cell polarity proteins such as tyrosinated tubulin and Fzd6 affected in mutant mice?

      We thank the Reviewer’s recommendations. We have assessed the expression of tyrosinated tubulins and found they are similarly polarized in ependymal cells from wild-type and Agbl5<sup>M1/M1</sup> mice. The results are presented in Figure 7O, P in the revised MS. We also tried to assess the expression of Fzd6. However, with the antibody we tested, Fzd6 signals were not convincing. Therefore, we prefer to not showing the results and drawing a conclusion on it.

      (2) The phenotype of multiciliated cells in tracheas should also be examined in mutant mice. It is important to elucidate whether AGBL5 commonly functions in multiciliated cells of other organs.

      We thank the Reviewer’s suggestion. We have assessed the multicilia in the tracheas of P30 mice using scanning electron microscopy. Indeed, unlike the multicilia in wild-type mice that orientate to the same direction, those in the tracheas of Agbl5<sup>M1/M1</sup> mice often radiate to different directions in individual cells (Figure S2). Therefore, Agbl5 appears commonly involved in the alignment of multicilia.

      (3) According to Figure 1B, AGBL5 is highly expressed in the brain. Which cells in the brain express it besides ependymal cells?

      Based on the localization of tdTomato tracer engineered in Agbl5 mutant alleles (Figure 5B), Agbl5 is broadly expressed in the brain, including most if not all neurons, but its expression is much weaker in the subventricular zone (Please see Figure 5B). We clarified this in the revised MS.

      (4) From a mechanistic point of view, it is necessary to identify binding proteins with the N-domain of AGBL5 and perform functional analyses.

      We agree with the Reviewer. We feel that identification of the binding partners of CCP5 N-domain and functional analysis may be more suitable to go along with other mechanistic analysis on the function of CCP5 in ependymal cell polarities in our future study.

      Reviewer #2 (Recommendations for the authors):

      (1) Movie 3: The authors could comment on beating direction that seems impaired at the cell scale here, analysis of rotational polarity would be a plus.

      We thank the reviewer’s recommendation. We have analyzed the beating directions of cilia in individual cells and presented their consistency in each cell using mean vector length. These results indeed demonstrated defective rotational polarity in the cell level in Agbl5<sup>M1/M1</sup> mice (please refer to Figure 3I). We also analyzed the beating directions of ependymal multicilia in earlier stage in tissue level (Figure 7E). The mean vector length of cilia beating direction in Agbl5<sup>M1/M1</sup> mice is significantly reduced compared to that in wild-type, suggesting an aberrant rotational polarity in the tissue level in the mutant (Figure 7E).

      (2) Line 166 : ref to Werner et al., 2011 is not correct (no ependymal cells in that paper).

      We thank the reviewer’s critical reading. This reference has been removed.

      (3) Figure S4: B and D look similar picture to me same for C and F.

      We apologize for using the wrong images in this Figure. It has been corrected (Revised Figure S5).

      (4) Line 328: "Therefore, CCP5 apparently contributes to the establishment of both translational and tissue polarities in ependymal cells." Should be rephrased since translational polarity is also a tissue-level parameter which is the coordinated positioning of the ciliary patch. Cf Mirzadeh et al., 2010; Boutin et al., 2014.

      We thank the Reviewer’s comments. The sentence has been rephrased. This concept has been clarified where else needed in the revised manuscript. 

      (5) Line 348: "Planar cell polarity (PCP) pathway is essential for the establishment of rotational and tissue polarities in ependymal cells" Rotational polarity also has a tissular component (ie coordination of beating direction across tissue which is reflected by coordination of basal body polarities across tissue).

      We thank the Reviewer’s comments. We have clarified this point in the revised MS.

      (6) Incomplete bibliography citation (ie Walentek et al. without date).

      We thank the Reviewer’s critical reading. This bibliography citation has been fixed.

      Reviewer #3 (Recommendations for the authors):

      (1) Figure 3: The authors assert that the mutant's apical actin networks are significantly disrupted. However, the cell shown in Figure 3Q-R exhibits less compact centrioles than the controls, which could account for the reduction in phalloidin staining. Because centriole dispersion is variable in the mutant, quantifying actin staining in representative cells would be necessary to support such a statement.

      We thank the Reviewer’s comments. To address this concern, we have quantified the total intensity of actin network around BBs as well as the intensity of F-actin signals normalized to the level of immunosignals of BBs ((revised Figure 4M, N) please also refer to our response to Reviewer 1#). The results indicated the intensity of actin signal per BB is reduced in the mutant compared to that of wild-type mice. We feel that this analysis strengthened our statement.

      (2) Figures S3 and 4A-B show that the authors examine tdT expression to show that Agbl5 is expressed in ependymal cells but not in the SVZ. However, the tdT signal intensity is very low, and cells are very dense in this brain region. Double staining with specific markers of ependymal and/or SVZ cells would help convince readers that tdT is not expressed in SVZ cells.

      We agree with the Reviewer that the intensity of tdT signal is low, but broadly detectable in brain. Compared with its expression in ependymal cells, that in SVZ is much lower if any (Figure 4B’). To further confirm the identity of tdT-positive cells along the surface of ventricles, we have co-stained the brain sections of Agbl5<sup>WT/M1</sup> mice for tdT and S100b, a marker of mature ependymal cells (Figure 5C-E). The signal of tdt is colocalized with that of S100b and is much lower in cell layers next to S100b-positive cells.

      (3) Figure 4C-D and S4: The authors demonstrate that the number of FoxJ1+ cells per section increases at P7 (4C-E), while the number of S100β+ cells per mm decreases. Quantifications should be carried out in a similar manner to ensure comparability (number of positive cells per mm). Additionally, it remains unclear how to interpret these results, as S100β and FoxJ1 are two markers of differentiated cells, yet they exhibit opposite trends compared to controls. Is this a direct or indirect effect of Agbl5 mutation? The increase in the number of FoxJ1+ cells is particularly surprising given that the number of GT335 multicilia per mm remains unchanged (Figure 5).

      We agree with the Reviewer that quantifications should be carried out in a similar manner. In the revised MS, the quantification of Foxj1-positive cells is presented in number per mm (Figure 5I). To be noted, the expression of Foxj1 was assessed at P7 when ependymal cells are differentiating. while the expression of S100β was assessed at P17 when ependymal cells are supposed to be fully mature. Although S100b is used as a marker of mature ependymal cells, given its unclear function, we removed the results of S100b-positiving cell counting to avoid confusion in the revised manuscript.

      (4) Figure 5: In this figure, the authors analyze the labeling obtained with GT335, Acetylated Tubulin, and Arl13b antibodies. They show that the area of the cilium labeled by GT335 has increased, while the area labeled by the Acetylated Tubulin antibody has decreased in the knockout (KO) compared to the control. However, the length of the cilia observed through labeling with the Arl13b antibody remains unchanged. These observations are intriguing, but the low-magnification images in Figure 4 do not allow for the differences in ciliary axoneme labeling to be seen. Double GT335/AcTub labeling and higher magnifications are necessary for improved visualization of the differences in labeling along the axonemes.

      We thank the Reviewer comments. We have co-stained the cilia with GT335 and Ac-Tub antibodies, re-quantified cilia length labeled with respective antibodies and provided high magnification images. Please see the revised Figure 6O,P,R,S.

      (5) Figure 6: An analysis of ciliary beats using a high-speed camera shows no difference in ciliary beat frequency between the control and KO groups. At least, 3 animals should be analyzed. According to Figure 5, these findings indicate that the decrease in ciliary acetylation and the increase in ciliary glutamylation do not affect the beat frequency; instead, they disrupt the orientation of the beats. While these results are intriguing, they require further confirmation. Analyzing ciliary beats with a high-speed camera is informative, but at least three animals per genotype should be examined to ensure rigor. Furthermore, if the coordination of ciliary beats is impaired within the cells, this should be validated by double-labeling centrioles and basal feet to demonstrate that the orientation of cilia within the cells is abnormal.

      We thank the Reviewer’s comments. Sections shown in Figure 5 (currently Figure 6) are from P7 mice, while the ciliary beating analysis shown in Figure 6 (currently Figure 7) is from P15 mice. As the PTM changes in cilia were also observed in Agbl5<sup>M2/M2</sup>, we don’t think this is the cause that disrupts the orientation of the beats. The rotational polarity of Agbl5<sup>M1/M1</sup> ependymal cells is affected. Please refer to the analysis in Figure 3I and Figure 7E in the revised manuscript.

      (6) Figure 6F-G: β-Catenin labeling reveals cells of varying sizes in the KO. This phenotype is typical of ciliary mutants that lack primary cilia (Mirzadeh et al., 2010). Hence, it is essential to examine the mutation's impact on the presence, length, and positioning of the primary cilium in ependymal cell progenitors.

      We thank the Reviewer’s constructive comments. We assessed the area of ependymal cells labeled with β-Catenin. Indeed, the ependymal cells in the mutant showed larger area than that of wild-type. The ratio of the area of BB patch over that of cell surface is reduced (please see Figure 7O, P in the revised manuscript). However, primary cilia are present in ependymal cell progenitors in the mutant and exhibit comparable length with those in the wild-type (Figure S8). Due to some technique problems, we were unable to get convincing results from whole-mount ventricle walls for the primary cilium positioning at this time. We speculate that the localization of certain sensory proteins in primary cilia or the positioning of primary cilia might be affected in Agbl5<sup>M1/M1</sup> mice. We discussed this possibility and will certainly systemically assess this intriguing aspect in our future investigation.

      (7) Given the regular beating frequency in the KO at P15, how do the authors explain the complete absence of ciliary beating in the adult? How many animals were analyzed? One would expect ciliary beating to remain unaffected as it was at P15 unless the cilia structure was specifically altered at the adult stage. Is that the case?

      We thank the Reviewer’s critical questions. We do think that the ciliary structure of Agbl5<sup>M1/M1</sup> ependymal cells is likely altered during aging. Given that only Agbl5<sup>M1/M1</sup> but not Agbl5<sup>M2/M2</sup> mice develop hydrocephalus, we speculate the N-domain of CCP5 may contribute to the integrity of ependymal multicilia. We have added this in the Discussion section. For each genotype, 2 mice were analyzed.

      (8) Line 264 of the manuscript: replace intercellular with intracellular.

      It has been revised.

      (9) Indicate the number of animals analyzed in each experiment

      It has been included in figure legends.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Gruskin and colleagues use twin data from a movie-watching fMRI paradigm to show how genetic control of cortical function intersects with the processing of naturalistic audiovisual stimuli. They use hyperalignment to dissect heritability into the components that can be explained by local differences in cortical-functional topography and those that cannot. They show that heritability is strongest at slower-evolving neural time scales and is more evident in functional connectivity estimates than in response time series.

      Strengths:

      This is a very thorough paper that tackles this question from several different angles. I very much appreciate the use of hyperalignment to factor out topographic differences, and I found the relationship between heritability and neural time scales very interesting. The writing is clear, and the results are compelling.

      We thank Reviewer 1 for their kind words and enthusiastic support of our manuscript.

      Weaknesses:

      The only "weaknesses" I identified were some points where I think the methods, interpretation, or visualization could be clarified.

      (1) On page 16, the authors compare heritability in functional connectivity (FC) and response time series, and find that the heritability effect is larger in FC. In general, I agree with your diagnosis that this is in large part due to the fact that FC captures the covariance structure across parcels, whereas response time series only diverge in terms of univariate time-point-by-time-point differences. Another important factor here is that (within-subject) FC can be driven by intrinsic fluctuations that occur with idiosyncratic timing across subjects and are unrelated to the stimulus (whereas time-locked metrics like ISC and timeseries differences cannot, by definition). This makes me wonder how this connectivity result would change if the authors used inter-subject functional connectivity (ISFC) analysis to specifically isolate the stimulus-driven components of functional connectivity (Simony et al., 2016). This, to me, would provide a closer comparison to the ISC and response time series results, and could allow the authors to quantify how much of the heritability in FC is intrinsic versus stimulus-driven. I'm not asking that the authors actually perform this analysis, as I don't think it's critical for the message of the manuscript, but it could be an interesting future direction. As the authors discuss on page 17, I also suspect there's something fundamentally shared between response time series and connectivity as they relate to functional topography (Busch et al., 2021) that drives part of the heritability effect.

      We agree that investigating the heritability of ISFC (or stimulus-driven functional connectivity) would make for a very interesting future direction. Ultimately, we chose to analyze FC (vs. ISFC) profiles to allow for direct comparison with the sizable existing literature on the heritability of FC (such as in our Movie vs. Rest FC analysis) and decided to refrain from analyzing ISFC data in order to keep the present manuscript focused. ISFC analysis of this dataset will be a focus of future work.

      (2) The observation that regions with intermediate ISC have the largest differences between MZ, DZ, and UR is very interesting, but it's kind of hard to see in Figure 1B. Is there any other way to plot this that might make the effect more obvious? For example, I could imagine three scatter plots where the x- and y-axes are, e.g., MZ ISC and UR ISC, and each data point is a parcel. In this kind of plot, I would expect to see the middle values lifted visibly off the diagonal/unity line toward MZ. The authors could even color the data points according to networks, like in Figure 3C. (They also might not need to scale the ISC axis all the way to r = 1, which would make the differences more visible.)

      We thank R1 for this helpful suggestion- we originally set the y-axis limits to r = 1 in order to facilitate comparison between ISC (Fig. 1B) and FC profile (Fig. 6B) similarity, but we agree that this renders the group differences harder to discern and have updated the plot accordingly (along with thicker lines to enhance readability). We prefer to keep the line plots in the main body as they allow for direct comparison of all three groups on the same plot, but we have included the scatter plot version in Fig. S2 for those who are interested.

      (3) On page 9, if I understand correctly, the authors regress the vector of ISC values across parcels out of the vector of heritability values across parcels, and then plot the residual heritability values. Do they center the heritability values (or include some kind of intercept) in the process? I'm trying to understand why the heritability values go from all positive (Figure 2A) to roughly balanced between positive and negative (Figure 2B). Important question for me: How should we interpret negative values in this plot? Can the authors explain this explicitly in the text? (I also wonder if there's a more intuitive way to control for ISC. For example, instead of regressing out ISC at the parcel/map level, could they go into a single parcel and then regress the subject-level pairwise ISC values out when computing the heritability score?).

      We indeed included an intercept in this model using MATLAB’s fitlm function. This means that the model estimates the best-fitting line of the following form: heritability<sub>i</sub>=β0+β1ISC<sub>i</sub> +ε<sub>i</sub>. We agree that the interpretation of these ε<sub>i</sub> values and alternative approaches to controlling for ISC should be clarified. As such, we have added the following passages to the text:

      Methods: “Because the heritability of ISC is constrained by the degree of synchronization in a given area, we also sought to identify areas in which BOLD time courses were more/less heritable than would be expected based on ISC alone by fitting a linear model of the form heritability<sub>i</sub>=β0+β1ISC<sub>i</sub>+ε<sub>i</sub> and plotting the residuals. Regarding alternative approaches to controlling for ISC, although the heritability model introduced by Ge et al. allows for the inclusion of covariates defined at the subject level (e.g., age), it does not allow for covariates that are defined at the dyad level (e.g., pairwise ISC).”

      Results: “Here, negative values in the residual map indicate parcels where heritability is lower than expected based on ISC, while positive values indicate higher-than expected heritability.”

      (4) On page 4 (line 155), the authors say "we shuffled dyad labels"- is this equivalent to shuffling rows and columns of the pairwise subject-by-subject matrix combined across groups? I'm trying to make sure their approach here is consistent with recommendations by Chen et al., 2016. Is this the same kind of shuffling used for the kinship matrix mentioned in line 189?

      Briefly, shuffling the kinship matrix involved permuting the rows and columns of the matrix in the same manner (also known as the quadratic assignment procedure), whereas shuffling the dyad labels involved random permutations of the three group labels (MZ, DZ, unrelated), which could not be done through matrix operations as the age- and gender matching precluded the use of a complete similarity matrix. However, given concerns raised by Reviewer 2, we have removed our significance claims from this (and similar) sections, which we discuss in more detail in response to Reviewer 2’s weakness A.

      (5) I found panel A in Figure 4 to be a little bit misleading because their parcel-wise approach to hyperalignment won't actually resolve topographic idiosyncrasies across a large cortical distance like what's depicted in the illustration (at the scale of the parcels they are performing hyperalignment within). Maybe just move the green and purple brain areas a bit closer to each other so they could feasibly be "aligned" within a large parcel. Worth keeping in mind when writing that hyperalignment is also not actually going to yield a one-to-one mapping of functionally homologous voxels across individuals: it's effectively going to model any given voxel time series as a linear combination of time series across other voxels in the parcel.

      We agree that our efforts to present a simplified depiction of hyperalignment may mislead less familiar readers and have amended Fig. 4A according to this suggestion. We have also added text to the methods section (below) to clarify that the outputs of hyperalignment are time series that reflect linear combinations of other voxels’ time series from that parcel.

      “This approach independently transforms each subject's data within discrete anatomical parcels into the common space, yielding functionally aligned vertex time series that are calculated as weighted linear combinations of the original time series from all other vertices within that same parcel for that subject.”

      (6) I believe the subjects watched all different movies across the two days, however, for a moment I was wondering "are Day 1 and Day 2 repetitions of the same movies?" Given that Day 1 and Day 2 are an organizational feature of several figures, it might be worth making this very explicit in the Methods and reminding the reader in the Results section.

      We agree that this would be helpful and have added the following text to the relevant sections:

      “All clips were only viewed once by each subject, with the exception of the brief montage which was included at the end of each of the four runs for test-retest purposes.”

      “To characterize the heritability of brain responses to complex stimuli, we used 7T fMRI data from 178 HCP Young Adult subjects acquired across two days (using two largely non-overlapping sets of movie stimuli, see Methods)…”

      References:

      Busch, E. L., Slipski, L., Feilong, M., Guntupalli, J. S., di Oleggio Castello, M. V., Huckins, J. F., Nastase, S. A., Gobbini, M. I., Wager, T. D., & Haxby, J. V. (2021). Hybrid hyperalignment: a single high-dimensional model of shared information embedded in cortical patterns of response and functional connectivity. NeuroImage, 233, 117975. https://doi.org/10.1016/j.neuroimage.2021.117975

      Chen, G., Shin, Y. W., Taylor, P. A., Glen, D. R., Reynolds, R. C., Israel, R. B., & Cox, R. W. (2016). Untangling the relatedness among correlations, part I: nonparametric approaches to inter-subject correlation analysis at the group level. NeuroImage, 142, 248259. https://doi.org/10.1016/j.neuroimage.2016.05.023

      Simony, E., Honey, C. J., Chen, J., Lositsky, O., Yeshurun, Y., Wiesel, A., & Hasson, U. (2016). Dynamic reconfiguration of the default mode network during narrative comprehension. Nature Communications, 7, 12141. https://doi.org/10.1038/ncomms12141

      Reviewer #2 (Public review):

      Summary:

      The authors attempt to estimate the heritability of brain activity evoked from a naturalistic fMRI paradigm. No new data were collected; the authors analyzed the publicly available and well-known data from the Human Connectome Project. The paper has 3 main pieces, as described in the Abstract:

      (1) Heritability of movie-evoked brain activity and connectivity patterns across the cortex.

      (2) Decomposition of this heritability into genetic similarity in "where" vs. "how" sensory information is processed.

      (3) Heritability of brain activity patterns, as partially explained by the heritability of neural timescales.

      Strengths:

      The authors investigate a very relevant topic that concerns how heritable patterns of brain activity among individuals subjected to the same kind of naturalistic stimulation are. Notably, the authors complement their analysis of movie-watching data with resting-state data.

      Weaknesses:

      The paper has numerous problems, most of which stem from the statistical analyses. I also note the lack of mapping between the subsections within the Methods section and the subsections within the Results section. We can only assess results after understanding and confirming the methods are valid; here, however, Methods and Results, as written, are not aligned, so we can't always be sure which results are coming from which analysis.

      (A) Intersubject correlation (ISC) (section that starts from line 143): "We used nonparametric permutation testing to quantify average differences in ISC for each parcel in the Schaefer 400 atlas for each day of data collection across three groups: MZ dyads, DZ dyads, and unrelated (UR) dyads, where all UR dyads were matched for gender and age in years." ... "some participants contributed to ISC values for multiple dyads (thus violating independence assumptions)"

      This is an indirect attempt to demonstrate heritability. And it's also incorrect since, as the authors themselves point out, some subjects contribute to more than one dyad.

      Permutation tests don't quantify "average differences", they provide a measure of evidence about whether differences observed are sufficient to reject a hypothesis of no difference.

      Matching subjects is also incorrect as it artificially alters the sample; covarying for age and sex, as done in standard analyses of heritability, would have been appropriate.

      It isn't clear why the authors went through the trouble of implementing their own nonparametric test if HCP recommends using PALM, which already contains the validated and documented methods for permutation tests developed precisely for HCP data.

      The results from this analysis, in their current form, are likely incorrect.

      We appreciate that permutation tests do not quantify average differences and intended to write “We used non-parametric permutation testing to quantify [the significance of] average differences…”. Our intention with this analysis was not to demonstrate heritability, but rather to quantify group differences in ISC in a manner that is interpretable for readers who are unfamiliar with h<sup>2</sup> (e.g., “identical twins’ BOLD time courses were 59% more similar than those from pairs of unrelated individuals”) and motivate the formal heritability analysis used later in the paper. Indeed, all of the heritability analyses in this paper leveraged a validated multidimensional heritability method first introduced by Ge et al. (2016) and used by many other investigators since then. Furthermore, we covaried for age and sex at the subject level in all our heritability analyses, and always tested the significance of these heritability values using a validated permutation procedure (the quadratic assignment procedure; Hubert & Schultz, 1976) that respects the non-independence of dyadic data.

      Regarding the shuffling procedure used for Figure 1, while PALM is the standard for univariate, subject-level GLMs in the HCP pipeline and can accommodate nested designs (i.e., subjects within families), it is not designed to handle the unique relational dependencies of dyadic ISC analysis (i.e., the same subject contributing to multiple dyads). Although the element-wise resampling approach was the most appropriate approach available, it is known to inflate the false positive rate (Chen et al., 2016; doi:10.1016/j.neuroimage.2016.05.023); given that this analysis was simply meant to motivate our later hypothesis testing heritability analyses, we have removed significance claims from this section of the manuscript. Still, we emphasize that this has no bearing on the validity of our conclusions which were supported by our formal heritability analyses; throughout our paper we have correctly used the appropriate methods to back the stated claims.

      (B) Functional connectivity (FC) (section that starts from line 159): Here the authors compute two 400x400 FC matrix for each subject, one for rest, one for movie-watching, then correlate the correlations within each dyad, then compared the average correlation of correlations for MZ, DZ, and UR. In addition to the same problems as the previous analysis, here it is not clear what is meant by "averaging correlations [...] within a network combination". What is a "network combination"? Further, to average correlations, they need to be r-to-z transformed first. As with the above, the results from this analysis in its current form are likely incorrect.

      We regret that R2 had difficulty understanding our analysis and have added the following text to the relevant Methods section to clarify our approach:

      “For example, there are 16 parcels in the Kong et al. Auditory network and 17 parcels in the Language network, so the FC profile for a given subject’s Auditory-Language network combination consists of the (16 * 17 =) 272 correlation coefficients between all unique pairs of one parcel from each network.”

      As we stated in the previous Methods paragraph, “All Pearson r values in this and all other analyses were Fisher z-transformed before averaging (and converted back to Pearson r for visualization)”. Thus, contrary to the reviewer’s assertion, these analyses were performed correctly. Once again, we emphasize that this analysis was not intended to demonstrate heritability, but rather to describe group differences in FC in familiar units.

      (C) ISC and FC profile heritability analyses (section that starts from line 175): Here, the authors use first a valid method remarkably similar to the old Haseman-Elston approach to compute heritability, complemented by a permutation test. That is fine. But then they proceed with two novel, ill-described, and likely invalid methods to (1) "compare the heritability of movie and rest FC profiles" and (2) to "determine the sample size necessary for stable multidimensional heritability results". For (1), they permute, seemingly under the alternative, rest and movie-watching timeseries, and (2), by dropping subjects and estimating changes in the distribution.

      The (1) might be correct, but there are items that are not clearly described, so the reader cannot be sure of what was done. What are the "153 unique network combinations"? Why do the authors separate by day here, whereas the previous analyses concatenated both days? Were the correlations r-to-z transformed before averaging?

      The (2) is also not well described, and in any case, power can be computed analytically; it isn't clear why the authors needed to resort to this ad hoc approach, the validity of which is unknown. If the issue is the possibility that the multidimensional phenotypic correlation matrix is rank-deficient, it suffices that there are more independent measurements per subject than the number of subjects.

      Regarding (1), we have clarified in section 2.6 that the 153 unique network combinations reflect each unique pair of 17 Kong networks. All of our analyses, including this one, were performed separately for each day of data collection, as we state throughout the paper and visualize in our figures (although we acknowledge that, on some occasions, we [conservatively] performed FDR-correction on a combined set of p-values, as discussed in our response to K). Given that the null hypothesis for this analysis is that rest FC and movie FC are equally heritable, we are not sure why permuting rest and movie FC matrices would be invalid. All Pearson r values were z-transformed before averaging, as we stated in our paper.

      Regarding (2), we included this analysis in response to editorial concerns that our heritability analyses were not sufficiently powered, and we chose this approach because it serves as a simple way to demonstrate the stability of our results at various sample sizes whose validity is self-evident. Furthermore, this sort of subsampling approach has been used many times before in our field (e.g., Marek et al., 2022) and others (e.g., Manyara et al., 2024) to demonstrate the sample-size dependence and stability of statistical effects. We have added text explaining this to the relevant Methods section (2.6).

      (D) Frequency-dependent ISC heritability analysis (from line 216): Here, the authors decompose the timeseries into frequency bands, then repeat earlier analyses, thus bringing here the same earlier problems and questions of non-exchangability in the permutations given the dyads pattern, r-z transforms, and sex/age covariates.

      We did not use dyadic permutation testing for any of the frequency-dependent ISC analyses; rather, we used the jackknife SEMs to compare heritability across frequency bands and have added an explicit description of this to section 2.7. We have addressed the r-z transform and covariate concerns in previous comments.

      (E) FC strength heritability analysis (from line 236): Here, the authors use the univariate FC to compute heritability using valid and well-established methods as implemented in SOLAR. There is no "linkage" being done here (thus, the statement in line 238 is incorrect in this application. SOLAR already produces SEs, so it's unclear why the authors went out of their way to obtain jackknife estimates. If the issue is non-normality, I note that the assumption of normality is present already at the stage in which parameters themselves are estimated, not just the standard errors; for non-normal data, a rank-based inversenormal transformation could have been used. Moreover, typically, r-to-z transformed values tend to be fairly normally distributed. So, while the heritabilities might be correct, the standard errors may not be (the authors don't demonstrate that their jackknife SE estimator is valid). The comparison of h2 between dyads raises the same questions about permutations, age/sex covariates, and r-z transforms as above.

      We used jackknife SEs for these analyses to maintain consistency with the multidimensional heritability package used here, which only outputs jackknife SEs. We note that this jackknife approach (and the corresponding multidimensional heritability analysis) was detailed in prior work (Anderson et al., 2021), and that the leave-one-family-out jackknife has a long history of being used to estimate SEs in heritability studies, especially when working with smaller samples (Knapp et al., 1989). We are also not sure what “the comparison of h2 between dyads” means- heritability cannot be compared “between” dyads; rather, it is defined across dyads.

      (F) Hyperalignment (from line 245): It isn't clear at this point in the manuscript in what way hyperalignment would help to decompose heritability in "where vs. how" (from the Abstract). That information and references are only described much later, from around line 459. The description itself provides no references, and one cannot even try to reproduce what is described here in the Methods section. Regardless, it isn't entirely clear why this analysis was done: by matching functional areas, all heritabilities are going to be reduced because there will be less variance between subjects. Perhaps studying the parameters that drive the alignment (akin to what is done in tensor-based and deformation-based morphometry) could have been more informative. Plus, the alignment process itself may introduce errors, which could also reduce heritability. This could be an alternative explanation for the reduced heritability after hyperalignment and should be discussed. An investigation of hyperaligment parameters, their heritability, and their co-heritability with the BOLD-phenotypes can inform on this.

      To help set up our hyperalignment analyses, we have added text to the introduction explaining how hyperalignment would help to decompose heritability. The description in the Methods section included a reference to Bazeille et al., 2021, in which the hyperalignment method used here is discussed in detail. Still, we have added citations to additional papers (also cited in the Bazeille et al. paper, and elsewhere in our paper) in case that might be helpful. We note that it is not the case that all heritabilities were reduced by hyperalignment- as can be seen in Figs. 4D, 8A, and S15, hyperalignment did increase heritability in some voxels and network combinations. This would be expected under the alternative (albeit unlikely) hypothesis that functional topographies are not heritable, such that topographic variation between related individuals would obscure similarities in their (heritable) topography-independent brain responses. Recognizing that this alternative is unlikely, we believe the main novelty of this analysis comes from the magnitude of the hyperalignment effect (up to 40% of brain-wide heritability) and its spatial pattern (e.g., larger heritability decreases in visual vs. auditory cortex, the opposite of our NT result).

      We agree that we would see lower post-hyperalignment heritability if the alignment process itself introduced errors/noise, but this would be deeply surprising as hyperalignment increases ISC by design (and errors/noise could only decrease ISC). To demonstrate this, we have added Figure S7 which shows that (as expected) ISC across all voxels and subject pairs increases after hyperalignment (and that this increase is larger when hyperalignment is performed in larger parcels). Given that hyperalignment increased ISC, and that it is blind to twin status, we are unsure how it could have introduced errors that would have confounded this result.

      (G) Relationships between parcel area and heritability (from line 270): As under F), how much the results are distorted likely depends on the accuracy of the alignment, and the error variance (vs heritable variance) introduced by this.

      We agree that alignment accuracy could potentially impact parcel-level differences in how much heritability changes following hyperalignment, and we included the frequency dependent h<sup>2</sup><sub>residuals</sub> (controlling for differences in ISC) in Fig. 3 for this reason, as more accurate hyperalignment should result in greater increases in ISC, raising the heritability ceiling. We note that we observe similar relationships between parcel rank and frequency dependent changes in these residualized maps, suggesting that our parcel-level differences are not simply the result of better alignment in more sensory parcels.

      (H) Neural timescale analyses (from line 280): Here, a valid phenotype (NT) is assessed with statistical methods with the same limitations as those previously (exchangability of dyads, age/sex covariates, and r-z transforms). NT values are combined across space and used as covariates in "some multivariate analyses". As a reader, I really wanted to see the results related to NT, something as simple as its heritability, but these aren't clearly shown, only differences between types of dyads.

      We have addressed the exchangeability, covariates, and r-z transform comments above (in A). As we explained for our FC strength analyses, we are underpowered to evaluate the heritability of unidimensional traits (like the heritability of NT magnitude), and the heritability of a closely-related measure (BOLD turnover magnitude) has already been established in a larger sample of HCP subjects (https://doi.org/10.1152/jn.00402.2022). Still, we agree that more results related to the heritability of NTs would be of interest to our readers. As such, we have added an analysis in section 3.4 quantifying the heritability of multivariate NT topographies and used SOLAR to quantify the heritability of NT magnitudes, with the disclaimer that this and similar analyses are underpowered (hence the large difference in day 1 and day 2 heritability effect sizes). We also removed significance claims for the dyadic NT similarity analysis.

      (I) Significance testing for autocorrelated brain maps and FC matrices (from line 310): Here, the authors suddenly bring up something entirely different: reliability of heritability maps, and then never return to the topic of reliability again. As a reader, I find this confusing. In any case, analyses with BrainSMASH with well-behaved, normally distributed data are ok. Whether their data is well behaved or whether they ensured that the data would be well behaved so that BrainSMASH is valid is not described. As to why Spearman correlations are needed here, Mantel tests, or whether the 1000 "surrogate" maps are valid realizations of the data under the null, remains undemonstrated.

      We brought up reliability in this section because we show the reliability of our results across the two days of data collection several times in the paper. R2 is correct to point out that BrainSMASH was validated using normally distributed brain maps, and although some of our brain maps contain normally distributed values, others are right skewed (due largely to the fact that many voxels/parcels exhibit low ISC while visual/auditory areas have very high ISC). In preparing our original manuscript, we visualized BrainSMASH’s variogram outputs for one of the most skewed inputs (vertex-wise BOLD time course heritability) and found that the autocorrelation structures of the empirical and null maps were well-matched. We did not include this in the original manuscript as it is not commonplace in the field to report the variograms, see Author response image 1. Furthermore, our use of Spearman (vs. Pearson) correlations renders these distributional differences less relevant, as the Spearman correlation transforms all inputs to a uniform distribution. To empirically check that these distributional differences do not bias our results, we retested the significance of all brain map associations using the spin test (10.1016/j.neuroimage.2018.05.070), an alternative method that does not assume normally distributed inputs, and obtained identical p-values for all analyses (P<.001 in all cases).

      Author response image 1.

      (J) Global signal was removed, and the authors do not acknowledge that this could be a limitation in their analyses, nor offer a side analysis in which the global signal is preserved.

      Although we agree that GSR is a contentious preprocessing step for certain analyses, it has explicitly been shown to increase ISC signal-to-noise without compromising FC fingerprints (Graff et al., 10.1016/j.dcn.2022.101087), and it is uncommon to perform ISC analyses with and without GSR. Still, we have added additional text to our Methods section explaining our rationale for using GSR and that this could affect our results. We also re-ran our main analysis (BOLD time course heritability) with and without GSR and found that GSR had little impact on our results; we have included this in our manuscript as Fig. S4.

      Specifically, we see that GSR resulted in a slight increase in heritability (average Day 1 h<sup>2</sup> with/without GSR = .064/.060; Day 2: .068/.061) and almost no effect on the spatial pattern of our results (With GSR/without GSR Spearman ρ = .99, P<sub>brainSMASH</sub> < .001 on both Day 1 and Day 2).

      (K) FDR is used to control the error rate, but in many cases, as it's applied to multiple sets of p-values, the amount of false discoveries is only controlled across all tests, but not within each set. The number of errors within any set remains unknown.

      We agree that the FDR usage in our original manuscript was inconsistent, in that for two analyses we FDR-corrected p-values from the two days of data collection together (instead of correcting p-values from each day separately and reporting voxels/parcels/etc. that were significant at q<.05 on both days, as in the rest of our analyses). We note that both approaches are more conservative than reporting significant results at q<.05 separately; regardless, to maintain consistency we have updated all analyses such that FDR correction is always performed separately for each day of data collection.

      (L) Generally, when studying the heritability of a trait, the trait must be defined first. Here, multiple traits are investigated, but are never rigorously defined. Worse, the trait being analyzed changes at every turn.

      Here, we analyze the heritability of movie-evoked BOLD time courses (Figures 1-5) as well as FC profiles (Figures 6-8). We defined FC profiles in our Introduction as an individual’s pattern of pairwise FC strengths (and further detailed how we quantified FC profiles in the relevant Methods section), and believe that “BOLD time course” is a well understood phrase in the field and does not need to be further defined. We also used hyperalignment to decompose the heritability of these traits into topography-dependent and independent portions, and (new to this version) also explicitly quantify the heritability of neural timescales, which we defined as the AUC of the ACF until the first negative ACF value in both the relevant Results and Methods sections.

      To make this clearer, we have modified the last paragraph of our Introduction to begin with:

      In the present work, we address these questions by analyzing 7T fMRI recordings of a twin sample acquired by the Human Connectome Project (Van Essen et al., 2013) to quantify the heritability of two distinct high-dimensional traits—stimulus-evoked BOLD time courses and functional connectivity profiles—across the cortex.

      Reviewer #3 (Public review):

      Strengths:

      It's sort of novel to study the heritability of movie-watching fMRI data. The methodology the authors used in the paper is also supportive of their findings. Figures are nicely organized and plotted. They finally found that sensory processing in the human brain is under genetic control over stable aspects of brain function (here referring to neural timescale and resting state connectivity).

      Weaknesses:

      What I am worried about most is the sample size and interpretation of heritability.

      (1) Figure 1. I assumed that the authors just calculated the ISC within each group (MZ, DZ, and UR). Of course, you can get different variations between each group. Therefore, there is heritability. Why not calculate ISC across the whole sample, then separate MZ, DZ, and UR?

      We believe that this question is getting at the difference between pairwise ISC (i.e., correlating one BOLD time course from one subject with that from another subject) and leave-one-subject-out ISC (i.e., correlating one BOLD time course from one subject with the corresponding average time course across all other subjects). We chose to use the pairwise ISC method because it allows us to capitalize on the information contained in the n<sup>2</sup> pairwise ISC matrix (whereas the other approach averages out meaningful information to yield a n<sup>1</sup> ISC matrix) and leverage a more sophisticated multidimensional heritability approach. Also, the leave-one-subject-out approach introduces additional issues re: handling family-level data (e.g., should we include a subject’s twin in the leave-one-subject-out average? If so, how should we handle subjects who don’t have a twin in the dataset, as averaging data from different numbers of subjects will lead to different ISC magnitudes? etc.).

      (2) Heritability scores in the paper are sort of small. If the sample size is small, please consider p-values, which will tell more about the trustworthiness of your heritability.

      We report p-values for heritability throughout our paper (e.g., stating that BOLD time courses are significantly heritable in 99% of parcels in Figure 2), and we believe that the reliability of our spatial maps across days of data collection (also quantified with p-values) further demonstrates the trustworthiness of our results. Finally, as we demonstrate in Figure S5, our sample size is more than sufficient to reliably detect small effects.

      (3) I don't understand the high-frequency signals in fMRI data. It's always regarded as noise, the band 1 here in particular.

      In addition to driving shared neuronal responses (which are captured in BOLD signal oscillations <.1 Hz or so), movies also elicit shared cardiac, respiratory, and motion responses across participants at higher frequencies. Although we used a relatively conservative denoising approach here, we believe some of these non-neuronal signals are still present in our data; alternatively, it is also possible that these signals reflect “fast” BOLD responses at >.15 Hz (as discussed in 10.1016/j.neuroimage.2021.118658). In any case, the fact that information in this frequency band is considerably less heritable than information in slower frequency bands supports the idea that this band is noisier and suggests that our heritability results are driven by canonical neuronal activity-related BOLD signals.

      (4) The statement "we show that the heritability of brain activity patterns can be partially explained by the heritability of the neural timescale" should come from Figure 5. However, after controlling for NT, the heritability decreased max. 0.025 in temporal areas. I am not sure this change supports the statement. If the visual cortex is outlined, and combining ISC changes in the visual cortex, I think this would somehow be answered. Instead of delta h2, adding a new model h2 would be obvious to the readers.

      Although the decrease of 0.025 is small, we note that this constitutes around ~50% of BOLD time course heritability in some voxels (seen in comparison to Fig. 4C), and the spatial pattern of this result is quite consistent across days of data collection, indicating its reliability. Furthermore, the whole-brain distributions of results shown in Fig. 5B are clearly skewed towards negative values, indicating that controlling for NT partially reduces (or “explains”) BOLD time course heritability. Still, we agree that showing raw h<sup>2</sup> values in addition to the difference maps would be helpful for some readers and have added a corresponding supplementary figure (S12) which shows these.

      (5) Figures 7 and 8, when getting the difference of heritability, please also consider the standard errors of the heritability estimates. Then you can compare across networks/regions.

      We did consider adding standard errors for these heritability estimates, but found that visualizing standard errors for each of the 153 unique network combinations in our heatmaps rendered the visualizations difficult to parse, and given that our hypotheses concerned global (e.g., hyperaligned vs. MSM-aligned) or network-level (e.g., sensory vs. associative) patterns, we focused on calculating standard errors/p-values for these analyses (although we note that dyad-level standard errors can be found in Fig. 6B, where they are clearly marginal compared to the group effects).

      (6) I think movie VS resting state is a really important result in this paper. However, there is almost no discussion. Discussing this part would be more beneficial for understanding the genetic control over the neuron arousal and excitation circuits.

      We agree that this result was relatively under-explored in our Discussion section and have added additional text (lines 851-855) to connect this result to recent work on arousal-dependent uniqueness of FC.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Do the authors have any ideas why we see this hotspot of heritability in pMTG/LOTC? It really jumps out in Figure 1A and Figure 2. The more posterior sensory MT+ area seems to drop when regressing out ISC in Figure 2B, but this pMTG area stays hot. Is there anything special about this kind of multimodal biological motion/action observation / social perception area (Pitcher & Ungerleider, 2021)? I don't think this is necessary to discuss in the manuscript, but I'm curious if the authors have any speculation.

      We are not certain as to why BOLD time courses in this parcel are particularly heritable- although this area is associated with biological motion, that particular function tends to be more right lateralized, and here we see nominally higher heritability in the left hemisphere. Per a Neurosynth review (and consistent with the left lateralization), we believe this may have more to do with speech processing, but a more definitive answer will require further investigation.

      (2) Page 3, line 127: "More information on these clips"-it might be worth saying a little bit more here just to make sure people understand that these are audiovisual clips, they include language, they're long enough to convey meaningful social and narrative information, etc.

      We agree and have added additional details on the clip composition to the relevant methods paragraph.

      (3) Figure 1 caption: can you add a sentence reminding readers what's going on with Day 1 and Day 2?

      We thank R1 for this suggestion and have added a sentence to this effect at this location.

      (4) Page 9, line 379: "although these more associative parcels do not encode a substantial amount of stimulus-specific information"-is this really true? I suspect these association areas still have decent ISCs, even if there are many processing stages downstream of the raw stimulus.

      Although these parcels are not the most synchronized by the stimulus, we agree that it is unfair (and vague) to say that they do not encode a substantial amount of stimulus-specific information. We have edited this sentence to make a more specific claim and highlight the relatively lower ISC in these parcels vs. more unimodal sensory areas.

      (5) Page 9, line 417: Can you unpack a bit more what you mean by "supra-BOLD frequency band"?

      Here, we refer to the fact that BOLD signals resulting from neuronal firing events have frequencies below ~.15 Hz (Josephs and Henson, 1999). We have added additional text and the Josephs and Henson citation to this line to further unpack this point.

      (6) Page 18, line 695: This discussion of how attention and gaze might partly shape response time series reminded me of recent work by Borovska & de Haas (2024)-might be worth citing.

      We are grateful to R1 for alerting us to this very relevant work and have included a reference to it in our discussion.

      (7) Page 19, line 755: I'm not sure I'd describe the hyperalignment results here as a "deleterious effects [on] heritability"-my reading was that hyperalignment allows you to say something more specific about heritability of function by allowing you to effectively factor out heritability effects that reduce to individual differences cortical topography; this seems like a good thing!

      We agree that “deleterious” was a poor word choice given its negative connotation, and have edited this sentence to read:

      “With this in mind, future studies investigating genetic correlations between brain function and behavioral variables may benefit from hyperalignment, as it can factor out individual-specific cortical topography and thus yield more precise estimates of functional heritability.”

      (8) I would love to see a ventral view in some of these plots! Not asking you to recreate the figures, but the ventral temporal cortex is an area of interest for many folks in the movie fMRI space (e.g., Haxby et al., 2011).

      We agree that ventral views would be of interest to some readers and have added the corresponding maps for our main results in supplementary figures S3 and S9.

      References:

      Borovska, P., & de Haas, B. (2024). Individual gaze shapes diverging neural representations. Proceedings of the National Academy of Sciences, 121(36), e2405602121. https://doi.org/10.1073/pnas.2405602121

      Haxby, J. V., Guntupalli, J. S., Connolly, A. C., Halchenko, Y. O., Conroy, B. R., Gobbini, M. I., Hanke, M., & Ramadge, P. J. (2011). A common, high-dimensional model of the representational space in human ventral temporal cortex. Neuron, 72(2), 404416. https://doi.org/10.1016/j.neuron.2011.08.026

      Pitcher, D., & Ungerleider, L. G. (2021). Evidence for a third visual pathway specialized for social perception. Trends in Cognitive Sciences, 25(2), 100-110. https://doi.org/10.1016/j.tics.2020.11.006

      Reviewer #2 (Recommendations for the authors):

      (1) To address the common core analytical problems listed under A), B), C), D), E), and basically throughout the methods:

      (a) Conduct permutations with exchangability restrictions to account for the pattern of dyad-relationships as e.g. implemented in PALM.

      (b) Control for age and sex covariates as covariates (e.g. as in SOLAR), rather than by matching.

      (c) Perform r-to-z transforms when conducting further analyses on correlations that assume normality.

      (d) For all analyses that assume normal distributions, e.g. in SOLAR and BrainSMASH, check that this is the case.

      We have explained how PALM is not suited for the study of effects that are defined at the dyad level (A), that we controlled for age and sex covariates in all our formal heritability analyses in our original submission (B), that we always performed r-to-z transforms when indicated in our original submission (C), and that our spatial permutation results don’t hinge on distributional differences (D).

      (2) Replace SEs derived from kacknife approach with those from SOLAR, or provide a comparison and motivation and/or demonstrate that SEs are correct.

      A more thorough explanation of the block jackknife procedure can be found in prior work introducing the multidimensional heritability method used here (Anderson et al., 2021).

      (3) Given problem (F & G):

      (a) Consider studying the parameters that drive the hyperalignment. They can be included as covariates in heritability analyses, and/or their heritability is of interest to understand the reasons for the heritability reduction post-hyperaligment.

      We agree that this would be interesting but the specific parameters that drive hyperalignment are beyond the scope of this study.

      (b) Include the alternative explanation of hyperalignment-induced noise in the discussion.

      We have added a figure showing that hyperalignment does not increase noise in ISC and explained here why “hyperalignment-induced noise” does not constitute a reasonable alternative explanation for our results.

      (4) Add heritability results for NT phenotypes.

      We have added heritability analyses for NT topography and (global) NT magnitude, as detailed above.

      (5) Motivate global signal removal, and acknowledge this process typically alters results substantially.

      We have added an explanation of our rationale for using GSR and shown in this response that it does not in fact substantially alter the results.

      (6) Rephrase and/or clarify the following:

      (a) "permutations quantify average differences" (under A).

      (b) "network combinations" and related analyses (under B & C).

      (c) why some analyses are separated per visit/day and others not (C).

      (d) methods and reasons for sample size estimation (C).

      We have rephrased or clarified all of the above.

      Reviewer #3 (Recommendations for the authors):

      (1) Participants should be recleared. I know HCP 7T data has 184 subjects. How can the authors have 176 twins and 690 unrelated subjects?

      As we reported in our Methods section, 178 subjects had complete movie-watching datasets, and 176 subjects had complete movie-watching and resting-state datasets. Of the 178 subjects with complete movie-watching data, we identified 690 age- and sex-matched dyads.

      (2) Figure 1. I don't find Figure S1A in Figure S1.

      We thank R3 for catching this error- we have amended this reference to read Fig. S1.

      (3) I could also suggest putting Figure 1 and Figure 2 together.

      We thank R3 for this suggestion- ultimately, we prefer to keep these figures separate to reinforce the difference between our dyadic similarity and formal heritability analyses.

    1. Author response:

      The following is the authors’ response to the original reviews.

      We are most grateful to both reviewers for providing valuable feedback on our manuscript.

      Reviewer 1 had solely favorable comments, with no suggestions for revision.

      Reviewer 2 pointed out that experiment evaluating the effect of CP4 on pVHL half-life (originally included as Figure 3c) was difficult to evaluate because of CP4’s effect on pVHL abundance prior to cycloheximide treatment. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      Reviewer 2 also pointed out that experiment evaluating the effect of CP4.29 on HIF-2α half-life (originally included as Figure 4g) was not very compelling. We agree with this assessment, and we opted to remove this experiment from the revised manuscript since it was not central to our overarching conclusions.

      We agree with Reviewer 2’s suggestion that additional experiments could further solidify that C4.29 downregulates HIF2 in a purely “on-target” manner, however we prefer to reserve such studies for the future.

      Reviewer 2 also made several valuable suggestions for the text itself (awkward wordings / citations / clearer figure legends). We appreciate this feedback and have updated the text accordingly.

    1. Author response:

      We would like to express our gratitude for the thorough evaluation of our manuscript by the editors and reviewers. We are grateful for the overall positive assessment. The suggestions for improvement are reasonable, and we are certain that addressing these points will improve the clarity, accessibility, and scientific integrity of the study. Thus, we plan to conduct a revision of the manuscript, addressing all the points raised. The most important planned adjustments are outlined below.

      (1) Improving the accessibility of the probabilistic modeling framework

      Reviewer 1 kindly stated that our Bayesian modeling framework for testing for species differences 'sets a new standard for our field.' As a new standard, however, the method should be explained in a more accessible way. Hence, we plan to provide additional explanations for the statistical workflow, e.g., by providing comprehensible visuals, to make the workflow easier to understand and easier to apply.

      (2) Statistical validation of qualitative claims

      We acknowledge that a statistical validation of qualitative claims regarding the relationship between seed and tongue movements and between upper and lower beak movements would considerably strengthen the validity of our findings. We thank Reviewer 2 for bringing permutation tests to our attention for quantifying the correlation between time series. Since permutation tests involving index-shuffling of one of the data sets are generally not valid for time-series data [1, 2], we'll consider a variant of a trial-swapping permutation test, such as a permute-match test [3]. Alternatively, the truncated time shift (TTS) test [2] might be an option, as also this method is valid for auto-correlated time series data. At this point, we can't tell yet which method we'll use for the revised manuscript. We need more time to assess the requirements of each method and evaluate which test is most appropriate to answer our specific research questions and best fits our kind of data.

      (3) Adjustments in the discussion

      Following the suggestion by Reviewer 1, we'll refine our discussion on the effects of skull size differences, putting more emphasis on the implications of potential effects for feeding kinematics in small species.

      Furthermore, as suggested by Reviewer 2, we'll soften our discussion on potential functions of lingual papillae in seed processing, as the current literature lacks experimental evidence for the claimed mechanistic roles.

      References

      (1) Yuan, A. E., & Shou, W. (2022). Data-driven causal analysis of observational biological time series. Elife, 11, e72518.

      (2) Yuan, A. E., & Shou, W. (2024). A rigorous and versatile statistical test for correlations between stationary time series. PLoS biology, 22(8), e3002758.

      (3) Yuan, A. E., & Shou, W. (2025). Permute-match tests: Detecting significant correlations between time series despite nonstationarity and limited replicates. eLife, 14.

    1. Author response:

      We thank the reviewers for their time and attention which will significantly improve the paper. Further, we are grateful for their appreciation of our goals and work. In sum, the reviewers point to our overstated discussion of experimental evidence which we will tone down, some slightly confusing points of argumentation which we will clarify, and some discussion points on the role of normative theories that we will add text to address. We believe this will improve the paper significantly and hope you agree!

      Major Concern: Experimental Support for Path-Integration is not as strong as suggested

      The major point raised by all reviewers (reviewer 1 comment 1, reviewer 2 comment 1, reviewer 3’s only weakness) was that our presentation of the experimental perturbation evidence for path-integration is stronger than the reality. On reflection, we agree with this evaluation. We thank the reviewers for raising it; we will moderate our writing and include the sensible caveats raised. In sum, we still think that the convergence of evidence points to path-integration: first, disruptions to grid cells lead to path-integration problems, though these perturbations admittedly aren’t perfectly precise; second, normative theories of path-integration lead to grid cells and predict grid cell behaviour; third, mechanistic models of path-integration match grid cell behaviour and predict connectivity subsequently measured in entorhinal cortex. However, the evidence is not as all-encompassing as we suggested.

      That said, we’d like to further comment on one point. It is argued (reviewer 1, comment 1) that there are other theories of grid cell function, and that we discuss these theories. We discuss efficient-coding only models of grid cells and emphasise strongly why we reject them. We also briefly discuss oscillatory-interference models of path-integration and our reasons for not pursuing them further. As such, the reviewer is correct that our reading of literature strongly points us towards path-integration rather than other theories. We will slightly change the framing of the paper to make it clear that we are making a case. However, we are not aware of other theories the reviewer might be referring to. If the reviewer can point us to the other suggested theories that we do not address we would be happy to evaluate and include them.

      We now turn to the remaining comments, and how we plan to address them.

      Reviewer 1, Comment 2 – There could be multiple roles for grid cells

      The reviewer is indeed right that grid cells might perform multiple functions. This could just mean that the same computational motif (e.g. path-integration) is reused across different computations though that introduces no changes to the required normative theory. A stronger claim would be that grid cells perform both path-integration and some other function. This, according to a normative perspective, would most likely change how grid cells were optimally structured. We use the fact that large parts of the grid cell code can be captured with only path-integration as an argument against additional roles for grid cells. That said, there exist properties of grid cells not well-captured by path-integration which could well be smoking guns for additional roles of grid cells. The review already discusses both discrepancies between grid cells in three and two dimensions, and inhomogeneities in the grid in complex environments, and we will add two more (heading direction and peak-to-peak/angular variability, discussed below) that we are grateful to the reviewers for raising, and we discuss each of these in detail below.

      That said, whether these are necessarily arguments against purely path-integration or a reflection of interesting mappings of the core path-integration mechanism to the measurements we make remains to be seen. We would argue that both 3D grid cells (as explained below: there appear to be 2D slices in which grid cells behave as you’d expect) and spatial inhomogeneities (as explained in the paper: mappings of torus to world can introduce warping) can be explained without reference to additional computational roles of grid cells, which remain to us the most parsimonious explanation. We discuss next the slight update to path-integration only that the heading direction story suggest. But in sum, our view is that these discrepancies are likely not fatal for our path-integration-centric view of grid cells, but may well suggest some very interesting clarifications.

      Reviewer 1, Comment 4 – The system has two heading signals: true & internal, why?

      The reviewer is right to point to the puzzle over true vs. purely internal heading direction and which drives grid cells. We believe recent work from Abraham Vollan has effectively solved this puzzle: there appear to be two parallel circuits, one theta-modulated and following internal heading direction, another theta-unmodulated and aligning more with true heading direction. We will make sure to include discussion of this exciting work in our revised submission. This serves as a good example of an update we concede to the most austere version of the path-integration only view. Rather, it seems there are two parallel path-integrators working with different heading signals. The reasons for this remain unclear, but seem to be related to attention and planning (Vollan et al. 2026).

      Reviewer 2, Comment 3: Real Grid Cells have peak-to-peak variability & Angular variability

      The reviewer is right to point to the discrepancy in peak-to-peak firing rate and angles within a module that we did not adequately address. First, it is Sorscher’s RNN models, not nonnegative PCA that can generate a distribution of grid angles (Redman et al. 2025), which suggests that path-integration and such variability are compatible. We emphasise this point because the non-path-integration results from nonnegative PCA produce grid cells oriented at 30 degree offsets, something not measured even when you’re careful as in Redman et al. 2025. Thus, this becomes an interesting target for future work: perhaps using theories of path-integration up to an error threshold (rather than perfect) such angular diversity would be recovered. We will include this in our discussion. Further, we will include discussion of peak-to-peak variability that, as yet, has no obvious role.

      Reviewer 2, Comment 1: grid cells are inhomogeneous in 3D or complex environments, doesn’t that break the theory?

      Disrupted grid coding in extended or 3D environments indeed deserve more discussion, which we will add. In particular, we will add recent evidence that grid cells in 3D can be understood via the correct sequence of 2D projections(Qi & Yartsev, 2026). These two phenomena seem, to us, consistent with a path-integration only view of grid cells, as discussed above, and we hope to make this position clearer.

      Reviewer 2, Comment 5: Couldn’t there be other reasons for multiple modules?

      We have suggested a consistent normative framework in which multiple modules are explained through their role in non-linear coding. We think this elegant, and the most parsimonious current theory. We could, of course, be wrong. The discrepancies pointed to above might be good clues to follow to work out what else these modules might be doing, but currently these alternative explanations seem not to exist. We will text to clarify this.

      Reviewer 1, Comment 3: The review confuses computational and parameter parts of normative theory

      We disagree with the reviewer’s dichotomisation of normative theory. We view a normative theory as the complete procedure that produces the predictions. Almost all such theories have parameters and hence fitting a theory to data comprises both elements (a) [computational role] and (b) [specific parameters] identified by the reviewer. Occasionally theories have no parameters in the traditional sense, e.g. Rebecca et al.; instead they have heavy assumptions that play an equivalent role. It is true that, as the reviewer says, Sorscher et al.’s work was criticised for producing grid cells only for specific parameter values. We never found this as damning as Schaeffer et al. argued: simply it says that that theory is only correct within the given parameter range. Rather, arbitrating between models, parameters, or assumptions seems the same basic process: see what they predict and keep working with models while they remain useful ways to understand measured phenomena. If a model with very specific parameter values remains useful, that seems okay. In fact, we argued extensively why we think the nonnegative PCA model is not a useful model, but this was for completely different reasons. To us this story just reinforces the importance of hygiene in normative research: perform parameter sweeps and clarify how they constrain the claims you are making, carefully arbitrate what models can capture. Indeed, that is the whole goal of this review. We might be misunderstanding and, if so, we welcome correction.

      Reviewer 2, Comment 4: Normative Models of Cells Beyond Grid Cells

      The reviewer is right that extending these models to other cell types is an interesting area for further work, and that other cell types do seem to be involved in aspects of navigational computations both in RNNs and the brain. We will include a discussion to this effect in the revised manuscript. That said, we think the modularity of grid cells and their tight-linking to path-integration calculations should also be appreciated as a win!

      Reviewer 2, Comment 2: Multi-modularity is not cleanly explained

      We thank the reviewer for the comments, we agree. We will clarify the story regarding multiple modules, and will explain the equation further.

      Reviewer 1, Comment 5: the early introduction of phase-shifted Grid Cells seem the perfect place to normatively argue for Path-integration!

      We agree with the reviewer that this point can be made both normatively (‘oh look! If I try to do this optimally, I get translations!’) or, as we did early in the paper, mechanistically (‘oh look! With these cells I can do this!’). Indeed, a large part of the point of our paper is that path-integration is what is required to normatively derive phase-shifted grid modules, something discussed by Rebecca et al., our earlier work, and RNN studies, and appreciated for two decades. The earlier part of the paper does not discuss these papers as that section is aimed at giving intuition for the solution (mechanism). Later sections then heavily discuss the normative angle. We hope that division of labour makes sense.

      Finally, we will refine our summary of Rebecca et al. The reviewer is right that neurons don’t have to be discrete, we apologise for that error, but our understanding is that the only meaningful role of a neuron in Rebecca et al.’s work is the region in which is active, effectively making every neuron a binary unit, which seems dubious. We will clarify that by “predict velocity from each current and next encoding” we mean that the normative constraint they enforce is axiom 1: sequential activity of sets of neurons i then j can be uniquely interpreted as a trajectory, i.e. a step or velocity. Their work is elegant, and we will try to do more justice to it in the revision.

      To conclude, we thank the reviewers for their extensive comments, and look forward to releasing a version that addresses their concerns.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Weaknesses:

      The pipeline is very complete, but also complex. Workflows (optimal artifact removal, best curation for data from a particular brain area or species) will vary according to experiment. Therefore, a discussion of the adaptability of the pipeline in the “Limitations” section would be helpful for readers.

      We added a dedicated paragraph in the Discussion section under “Limitations” focusing explicitly on the adaptability and flexibility of the pipeline. Furthermore, we took this feedback as an opportunity to make the pipeline itself significantly more modular and customizable with the most recent release (v1.2.0: https://aind-ephys-pipeline.readthedocs.io/en/latest/releases/1.2.0.html).

      Reviewer #1 (Recommendations for the authors):

      (1) In the description of the Phase-shift correction (Line 166-167): The current text reads “As a result, different groups of channels are sampled asynchronously.” A better description would be: “Sample times for different groups of channels are offset in time by a known amount.”

      We replaced the phrase in the manuscript text with the suggested formulation.

      (2) Figure 5 and description of the benchmarking overview (Line 326-336): How were spike trains (times) selected for the injected ground truth units? What was the range of firing rates?

      All injected spike trains were generated as independent Poisson processes featuring a mean firing rate of 15 Hz. We have now incorporated this explicitly into the main text to clarify the ground-truth injection process.

      (3) Figure 6, panel b: Are the gray points in the raster the original spikes in the test recording? From the pattern, it looks like there are 8 recovered ground truth units. Were the other 2 undetected by either sorter?

      That is correct; the two remaining units were undetected by both sorters. To clear up any confusion, we updated the caption for Figure 6 to state: “Note that spikes undetected by any of the sorter are not shown in the plot.”

      (4) Figure 7, panel c: Are all units returned from KS included in these distributions? (i.e., regardless of the KS refractory metric calculated by the sorter) - it would be useful to add that detail to the caption. It would also be helpful for panel C to include a total unit count from the two sorters... Also, since there are multiple ways to calculate the refractory period contamination, it would be good to state the calculation used here.

      Because we rely directly on the hybrid ground-truth for accurate validation, we included all raw units returned by Kilosort for this specific analysis. We have explicitly added a note detailing this to the caption. Panel C does report the total raw unit count returned by the two sorters (N = 3046 for KS2.5; N = 3652 for KS4).

      Additionally, to clarify the evaluation procedure, we appended the following statement to the main text: “For all results, we perform spike train comparisons and compute performance metrics as defined in (Buccino et al. 2020), using all units returned by the spike sorter (without any sorterspecific curation).”

      (5) Comments about the pipeline:

      The paper clearly demonstrates the immense utility of the pipeline in the authors’ work. I did some testing to try to understand its adaptability to workflows at my institution.

      I tested the pipeline on our local cluster running LSF. I’ve worked on a similar pipeline using Nextflow to automate ephys analysis with the same sorters. Questions that came up for me that would be usefully addressed in the ’Limitations’ section:

      (i) Is the pipeline meant to be run only in total? In particular, is it possible to start with preprocesseddata? (aind-ephys-preprocessing/code/params.json does not appear to include any means to turn off filtering, for example). Is the pipeline meant to be run only in total? In particular, is it possible to start with preprocessed data? (aind-ephys-preprocessing/code/params.json does not appear to include any means to turn off filtering, for example).

      To accommodate users who wish to run only parts of the workflow or use external preprocessing setups, we have refactored the codebase to support a custom preprocessing pipeline option. This makes it possible to turn off standard filtering or inject custom workflows.

      (ii) For debugging purposes, is there a means to go from preprocessing or sorting to result collection,so that interim results can be interpreted even when some steps of the pipeline aren’t working?

      The pipeline is designed to be a spike sorting pipeline, so the spike sorting step cannot be skipped. However, we have rewritten the post-sorting architecture to make it highly lightweight and fault-tolerant. The postprocessing step now only requires the random spikes and templates computation and downstream steps have been update to accomodate this lightweight option. As an example, if no quality metrics are computed, the curation step will be skipped. The visualization and QC steps also required updates to be tolerant to missing extensions. This required coordinate updates across several components:

      Postprocessing: PR #12

      Curation: PR #13

      Visualization: PR #21

      Quality Control: PR #20

      (iii) If these options to skip processes and output data ’partway’ are available, it would be great toadd that to the documentation.

      We have fully updated our online documentation for v1.2.0 (release notes: https://aind-ephys-pipeline.readthedocs.io/en/latest/releases/1.2.0.html), introducing a brandnew “Customization” guide page that comprehensively explains how to construct and provide custom preprocessing and postprocessing strategies, as well as how to integrate a new spike sorter in the pipeline: https://aind-ephys-pipeline.readthedocs.io/en/latest/customization.html

      Reviewer #2 (Public review):

      Summary:

      This work presents a reproducible, scalable workflow for spike sorting that leverages parallelization to handle large neural recording datasets. The authors introduce both a processing pipeline and a benchmarking framework that can run across different computing environments (workstations, HPC clusters, cloud). Key findings include demonstrating that Kilosort4 outperforms Kilosort2.5 and that 7× lossy compression has minimal impact on spike sorting performance while substantially reducing storage costs.

      Strengths:

      (1) Extremely high-quality figures with clear captions that effectively communicate complex workflow information.

      (2) Very detailed, well-written methods section providing thorough documentation.

      (3) Strong focus on reproducibility, scalability, modularity, and portability using established technologies (Nextflow, SpikeInterface, Code Ocean).

      (4) Pipeline publicly available on GitHub with documentation.

      (5) Clear cost analysis showing ~$5/hour for AWS processing with transparent breakdown.

      (6) Good overview of previous spike sorting benchmarking attempts in the introduction.

      (7) Practical value for the community by lowering barriers to processing large datasets.

      Weaknesses:

      No significant weaknesses were identified, although it is noted that the limitations section of the discussion could be expanded.

      We thank the reviewer for their constructive feedback on our manuscript.

      Reviewer #2 (Recommendations for the authors):

      The authors could discuss why 2.25 bps is the “lowest supported” level and whether more aggressive compression could be achieved with custom approaches, potentially exploring where performance breakdown occurs.

      The 2.25 bits-per-sample (bps) limit is an inherent constraint of the WavPack lossy compression library itself. While more aggressive, domain-specific, or custom compression schemes could be explored, we focused on WavPack due to its native support in modern neurophysiology ecosystems and its excellent performance in our prior simulated benchmarks (Buccino et al. 2023). We agree that using this hybrid benchmarking framework to explore alternative compression configurations is a highly valuable avenue for future work. We have added the following text to the Discussion: “The benchmarking pipeline will continue to develop as an open evaluation framework, enabling transparent and reproducible comparisons of spike sorting and preprocessing methods across the community. As one example, the work on lossy compression could be extended with additional codecs and parameter settings, exploiting our ability to read out spike sorting degradation directly from the hybrid ground truth spike times.”

      (2) The limitations section would benefit from expansion to include: (i) discussion of how simulated data limitations may affect generalization of benchmarking results to real neural data, and (ii) clarification of the effort required to add new spike sorters, including configuration complexities for coordinating Nextflow processes beyond simple SpikeInterface integration.

      We have expanded the Discussion section to address both items:

      (i) We added a paragraph detailing the specific limitations of hybrid ground-truth datasets (e.g., how idealized template injection might miss extreme multi-unit overlapping dynamics or nonstationary noise properties found in real tissue).

      (ii) We added a structural overview section clarifying the workflow complexity, detailing exactly what steps are required to map a new spike sorter into a Nextflow execution processes beyond its baseline addition to Spike Interface.

      (3) The authors should clarify the terminology of “hypothetical experiment” in the introduction to improve reader comprehension.

      We have removed the word hypothetical from the introduction to ground the explanation more directly.

      (4) The cost analysis could be improved by making it clearer whether “runtime” refers to wall-clock vs. total parallel compute time.

      We mean wall-clock time. While total parallel compute time aggregated across cloud workers remains roughly identical to the overall sequential execution on a lone cloud instance, cluster parallelization slashes the wall-clock time drastically. We have updated the text to explicitly state that reported runtimes represent wall-clock time.

      (5) The authors could address the Nextflow Java dependency limitation by discussing containerized execution options (Docker/Singularity) as a solution, while noting relevant HPC system restrictions.

      We have updated the text to mention the official pre-built Nextflow container images as an elegant workaround for environments where local Java installations are blocked or restricted: “However, one option to bypass installation issues is to run the main pipeline script in container images packaged with Nextflow (https://hub.docker.com/r/nextflow/nextflow).”

      (6) Figure 8 analysis would be strengthened by explicitly noting that compression effects are more substantial for lower-accuracy units, suggesting better preservation of higher SNR units.

      We appreciate this insight. To evaluate this systematically, we generated a new supplementary figure (Figure S3) which shows sorting performance during lossy compression as a function of the Signal-to-Noise Ratio (SNR) of ground truth units. The plot demonstrates that for Neuropixels 2.0 recordings, the slight drop in sorting accuracy is indeed heavily concentrated among low-SNR units. We have integrated this observation into the Results section.

      Reviewer #3 (Public review):

      (1) Could the authors please expand on the statement on line 274, that processing their test dataset serially “on a single GPU-capable cloud workstation... would take approximately 75 hours and cost over 90 USD.” How were these values calculated? I was a bit surprised that this is a ¿4-fold slowdown from their pipeline, but only increases the cost by 1.35x... More context on why this is, and maybe some context on what a g4dn.4xlarge is compared to the other instances, might help.

      We have expanded the cost analysis section in the manuscript methods to explain these figures explicitly. The serial run relies on a single continuous, higher-tier GPU workstation instance (g4dn.4xlarge) running uninterrupted for 75 hours.

      Our distributed pipeline, by contrast, dynamically provisions CPU-only instances to process chunked preprocessing steps concurrently, then spins up short-lived GPU spot instances only when Kilosort executes. While this parallel execution compresses the overall wall-clock time by over 4-fold, the cost is only moderately reduced because the CPU-only instances with many parallel processing cores are only slightly less expensive than GPU instances.

      (2) One of the most commonly used preprocessing pipelines for Neuropixels data is the CatGT/ecephys pipeline from the developers of SpikeGLX at Janelia. It may be worth commenting very briefly... on how the preprocessing steps available in this pipeline compare to the steps available in CatGT. For example, is “destriping” similar to the “-gfix” option in catGT to remove high-amplitude artifacts?

      We have added a section drawing direct comparisons to CatGT preprocessing workflows. We explicitly clarify that our phase-shift correction performs the exact same function as CatGT’s Tshift. We also point out that while our current version lacks a direct equivalent to CatGT’s saturation removal feature (-gfix), this capability is scheduled for incorporation in our upcoming pipeline release.

      (3) Why are there duplicate units (line 194), and how often is this an issue? I understand that this is likely more of a spike sorter issue than an issue with this pipeline, but 1-2 sentences elaborating why might be helpful for readers.

      Duplicate units are primarily an artifact of template-matching sorting routines (such as Kilosort), which can occasionally split a single biological neuron into multiple overlapping spatial templates or over-extract templates in highly active channel regions. We have added two clarifying sentences explaining this phenomenon in the text: “Next, duplicated units, that can arise when using template-matching methods if different templates are consistently fit to the same spikes, are removed based on the fraction of overlapping spikes.”

      Customizability of cluster curation parameters It seems from the parameter files on GitHub that the cluster curation parameters are customizable - correct? If so, it may be worth explicitly saying so in the curation section of the text... A presence ratio of >0.8 could be particularly problematic for some recordings (e.g. state transitions, behavior specific cells).

      (4) Yes, they are completely customizable. We agree that a rigid presence ratio cutoff of 0.8 would erroneously discard highly valid units that are modulated by specific behavioral states, or are active only during sleep vs. wake cycles. We have explicitly added text in the Curation section clarifying that all quality metric thresholds can be modified by the user: “Units are tagged as passing a default_qc when they satisfy the following criteria based on quality metrics thresholds. Thresholds can be user defined, and these are the default”.

      (5) The axis labels in Figures 3d-e are too small to see, and Figure 3d would benefit from a brief description of what is shown.

      We have updated the figures with enlarged, high-visibility axis labels and expanded the caption of Figure 3d to clearly describe the visualization.

      Figure 4 labels (“neural” vs “passing QC”) (6) What is the difference between “neural” and “passing QC” in Figure 4?

      We have updated the figure caption for Figure 4 to include an explicit cross-reference to the Curation methodology section, which defines the strict quantitative boundary between raw neural classification and formal automated QC passage.

      (7) I understand the current paper is focused on spike data... but I am curious about the NP2.0 probes that save data in wideband. Does the lossy compression negatively affect the LFP data? Is software filtering applied for the spike band before or after compression?

      Compression is applied to the raw streams prior to any secondary downstream software processing. For Neuropixels 1.0, compression is executed strictly on the action potential (AP) stream. For Neuropixels 2.0, compression operates directly on the unified wide-band data stream.

      Software filtering to separate bands is conducted post-decompression, as captured in our baseline workflow definitions (e.g., WavPack compression → decompression → preprocessing → Kilosort4). To clarify this, we added the following text: “In all cases, compression was applied before any preprocessing took place. For Neuropixels 1.0, we compressed the AP stream only. For Neuropixels 2.0, we compressed the full wide-band data.”

      Because LFP signals possess inherently smooth continuous dynamics across both space and time, they are much more amenable to lossless or near-lossless compression. Thus, the minor losses introduced by lossy compression are overwhelmingly localized to high-frequency spike band features, leaving LFP components virtually unaffected.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      The superiority of the optimized system might simply be due to insufficient T7 RNA polymerase in the initial lysate.

      We performed a T7 RNA polymerase titration (0–1600 ng/µL) in the initial system to test this hypothesis. Standard CFPS protocols typically utilize T7 RNA polymerase at ~90–100 ng/µL<sup>1</sup>. To fully characterize the concentration-dependent effect and determine the exact saturation threshold of T7 RNA polymerase in our system, we tested an extended range from 0 to 1600 ng/µL. As shown in the revised Figure S3B, the initial system's output reaches a plateau at ~800 ng/µL—a concentration nearly ten times higher than standard protocols. Increasing the concentration further (up to 1600 ng/µL) led to a decline in yield, likely due to inhibitory effects of excess enzyme or buffer components. Even under these T7-saturated conditions, our optimized system achieved ~45-fold higher NLuc output compared to the maximum possible output of the initial system. Notably, when the lysate concentration is increased to 70%, the productivity gap reaches nearly 80-fold, further demonstrating the extraordinary efficiency of our platform.

      As revised in the Discussion, this improvement confirms that the performance gain is not a result of a mere increase in T7 concentration. Instead, it represents a systemic synergy where our streamlined buffer and the optimized metabolic environment of the fast lysate together alleviate the transcriptional bottlenecks inherent in traditional platforms.

      Reviewer #2 (Public review):

      Performance or efficiency claims... needs to be supported by comparisons with typical cell free expression systems.

      We agree that robust benchmarking is essential for validating our claims of high efficiency. Our comparative evaluation was conducted across three levels:

      (1) Literature-based benchmarking: As detailed in Figures 3C, 4A-D, S3A-B, S4, and S5C, we extensively compared our system against the "initial" (35-component) and "PEPbased" platforms, which are established benchmarks widely utilized in CFPS literature. These diverse comparisons consistently demonstrate the superior performance and robustness of our optimized system across various conditions.

      (2) Commercial benchmarking: To provide independent verification, we performed a head-to-head comparison with a high-end commercial E. coli CFPS kit (PePExpress, Shanghai Epizyme, EC010L). As shown in the comparative data provided in this response (See author response image 1), our system exhibited remarkable rapid-expression capability, significantly outperforming the commercial kit in both speed and absolute yield. Our platform reached near-maximum yield within 2 hours, demonstrating a significant efficiency advantage over the commercial alternative.

      (3) Robustness and translational quality: The comparison was extended to challenging targets beyond standard reporters. As shown in Figures 4E-H, the successful synthesis of active BsaI restriction enzyme (a cytotoxic protein) and the functional assembly of vimentin (an aggregation-prone protein) demonstrate that our optimized system maintains superior translational quality and robustness compared to typical platforms that often struggle with such complex targets. By outperforming established academic benchmarks and a leading commercial platform in both yield and the ability to handle challenging proteins, our results provide compelling evidence that the simplified 7component system is highly efficient. In the revised Conclusion, we have explicitly contextualized "efficiency" as the integration of high protein productivity, reduced reaction complexity, and accelerated preparation speed.

      Author response image 1.

      Comparative evaluation of sfGFP yields between our _e_CFPS system (70% lysate) and a commercial kit (PePExpress) over an 8-hour time course.

      Summary of revisions: T7 titration data have been added to Supplementary Figure S3B in the revised manuscript. To provide the additional benchmarking evidence requested, commercial comparison data (PePExpress kit) are provided in Author response image 1, while the main manuscript remains focused on the mechanistic synergy and streamlined architecture of the system.

      We hope that these substantial new data and the corresponding revisions satisfy the reviewers' queries.

      References:

      (1) Kigawa, T. et al. Cell-free production and stable-isotope labeling of milligram quantities of proteins. FEBS Lett. 442, 15–19 (1999).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Pecak et al have deciphered the conformational dynamics of a heterodimeric model ABC transporter, TmrAB, a functional homolog of the human antigen transporter TAP, using single-molecule Forster resonance energy and fluorophores attached to residues at either nucleotide binding domains or periplasmic gate. The analysis not only differentiated ATP-free and bound states but also enabled the real-time monitoring of protein conformational changes, precisely dissecting transport cycles and resolving transient intermediates. This study is absolutely significant in providing and establishing a general pipeline delineating the conformational dynamics in heterodimeric ABC transporters.

      We thank the reviewer for this accurate and thoughtful summary of our work and its broader significance. We agree that the combination of single-molecule FRET with orthogonal validation approaches enables mechanistic resolution of conformational states and transitions that are not accessible by ensemble measurements. In particular, this framework allows direct discrimination of ATP-free and ATP-bound conformations, real-time tracking of transport cycle progression, and identification of transient intermediates in the heterodimeric ABC transporter TmrAB. We further agree that these capabilities support a generalizable strategy for dissecting conformation dynamics in related ABC transporters.

      Strengths:

      The scientific study is very well documented for experimental design, results, and conclusions supported by the experimental data. The authors have determined the conformational dynamics of TmrAB across different ATP concentrations, including physiological ones, and resolved an outward open state and other conformational states consistent with previous cryoEM and DEER studies.

      Weaknesses:

      The scientific study needs a bit of in-depth analysis with respect to consistency in K<sub>d</sub> and its implications on the mechanism.

      The apparent K<sub>d,ATP</sub> values were determined using two complementary approaches that report on different aspects of the system. Ensemble FRET measurements yielded values of 51 ± 38 µM (TmrAB<sup>NBD</sup>), 68 ± 25 µM (TmrAB<sup>PG</sup>), and 95 ± 26 µM (TmrAB<sup>PG_EQ</sup>), which are in good agreement with previously reported biochemical estimates (~100 µM for TmrAB<sup>EQ</sup>) (Stefan et al, 2020). The slightly elevated value observed for the E→Q variant may reflect modest perturbation of nucleotide handling in this slow-turnover background. Notably, the close agreement between labeled and unlabeled variants indicates that fluorophore attachment does not measurably affect ATP binding.

      In contrast, smFRET-derived K<sub>d,ATP</sub> values (13 ± 1 µM for TmrAB<sup>NBD</sup> and 2 ± 1 µM for TmrAB<sup>PG</sup>) are systematically lower. This difference likely arises from the difficulty of deconvoluting overlapping FRET populations at sub-K<sub>d,ATP</sub> concentrations, particularly for TmrAB<sup>PG</sup>, where state assignment is less well separated. Despite this quantitative offset, both approaches consistently indicate ATP saturation well below physiological concentrations and therefore support the same mechanistic conclusion that ATP binding drives conformational switching in TmrAB.

      Reviewer #2 (Public review):

      In their manuscript entitled 'ATP-driven conformational dynamics reveal hidden intermediates in a heterodimeric ABC transporter', Pečak et al. use elegant single-molecule FRET experiments in detergent to investigate the heterodimeric ABC transporter TmrAB. By combining simulations of the transporter's accessible volume with elegant trapping strategies, the authors identify an unresolved outward-facing open state and conclude that it is usually obscured by a rapidly interconverting ATP-bound ensemble. Overall, the study demonstrates that smFRET can resolve the short-lived intermediate states of TmrAB and potentially other ABC transporters that are obscured in ensemble measurements.

      It is a very interesting study that highlights the power of combining high-resolution structural information with spectroscopic approaches. I have three major points and a few minor criticisms.

      We thank the reviewer for the thoughtful and constructive evaluation of our manuscript and for highlighting the strength of combining structural and single-molecule approaches. We have addressed all major and minor points in detail below and revised the manuscript where appropriate to clarify limitations, justify analysis choices, and improve transparency.

      Major points:

      (1) The main weakness is that the authors base their conclusions on a very limited set of FRET pairs. While TmrAB has been extensively studied in terms of its structure, the authors should at least acknowledge this limitation more clearly.

      We agree that our conclusions are based on a limited number of FRET reporter pairs, and we now explicitly state this limitation in the revised manuscript. The chosen labeling positions were selected to probe two functionally critical regions—the nucleotide-binding domains and the periplasmic gate—based on prior structural and spectroscopic evidence. While this represents sparse sampling of the full conformational space, it is consistent with typical smFRET studies of membrane transporters, where experimental constraints generally limit the number of simultaneously accessible labeling positions (Asher et al, 2021; Asher et al, 2022; Levring et al, 2023; Wang et al, 2020).

      Importantly, both independent reporter variants yield consistent ATP-dependent population shifts, supporting the robustness of the observed trends. We further clarify that additional labeling sites could, in principle, resolve finer structural sub-states; however, given the already limited population separation in the current variants, such extensions would likely provide diminishing returns in state resolvability under the present experimental conditions. This trade-off is now explicitly discussed.

      (2) Most smFRET distributions were fitted with one, two, or three Gaussians. However, in several cases, additional populations with noticeable amplitudes appear to be present (e.g., Figure 3c at 0.1 mM and 3 mM ATP; Figure 4a, apo; Figure 4c, 0.3 mM R9L). Could the authors clarify why these populations were not included in the analysis?

      We thank the reviewer for this careful observation. Low-amplitude sub-populations are occasionally detected in individual histograms; however, they were not included in the quantitative model because they do not meet criteria for reproducibility, amplitude robustness, or structural assignability. Specifically, these features vary between replicates, contribute minimally to total population, and cannot be mapped to structurally or biochemically defined states based on available cryo-EM (Hofmann et al, 2019), DEER/PELDOR (Barth et al, 2018; Barth et al, 2020), or accessible-volume simulations.

      Similar minor subpopulations have been reported in smFRET studies and often attributed to photophysical or labeling heterogeneity effects (Asher et al, 2022; Husada et al, 2018). To avoid over-parameterization, we therefore restricted analysis to reproducible, structurally supported states. This rationale is now clarified in the revised manuscript.

      (3) Figure 3c (3 mM ATP): Is it truly possible to distinguish the two states in this distribution?

      We agree that state separation in the TmrAB<sup>PG</sup> variant is limited (ΔE = 0.11), and we now explicitly acknowledge this constraint in the manuscript. To improve robustness under these conditions, we used a constrained fitting strategy in which the apo-state distribution was fixed from nucleotide-free measurement, reducing parameter degeneracy during fitting of ATP-bound datasets.

      While single-molecule trajectory-based approaches such as Hidden Markov Modeling would be ideal for resolving dynamic interconversion, this was not feasible due to the low fraction of dynamic traces at the available temporal resolution. We therefore rely on population-level analysis, which remains consistent across replicates and reporter variants.

      Notably, independent measurements from two reporter positions (TmrAB<sup>NBD</sup> and TmrAB<sup>PG</sup>) yield similar ATP-bound population fractions at saturating ATP concentrations (~77% vs. ~80%), supporting the robustness of the inferred state distribution despite partial overlap.

      We have revised the manuscript to more clearly articulate methodological limitations, strengthen the justification of our analytical approaches, and improve the clarity of data presentation. These revisions enhance the transparency and robustness of the study and address the reviewer’s concerns.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Here are a few comments that can help to improve the study.

      (1) Line 115: The authors have checked the purity and monodispersity of the protein sample using SDS-Gel and size exclusion chromatography; however, additional characterization using negative stain electron microscopy, which clearly shows the monodispersity, will be useful.

      We agree that negative stain EM can provide an additional assessment of sample homogeneity. Given the extensive prior structural characterization (Hofmann et al, 2019; Nocker et al, 2026; Nöll et al, 2017) and the SEC profiles presented here, we believe that additional negative stain EM would unlikely provide substantial new information regarding sample homogeneity. We have clarified this point in the manuscript by explicitly referencing the relevant cryo-EM studies.

      (2) Line 116: The authors have mentioned that the enzymatic activity of TmrAB was retained after purification. Although smFRET results showing conformational dynamics of TmrAB confirm its ATPase activity, a comment on the effect of labelling on ATPase activity will be useful.

      We appreciate this important point. Previous studies on spin-labeled TmrAB<sup>NBD</sup> demonstrated transport activity comparable to wild-type TmrAB, indicating that cysteine substitution and label conjugation do not substantially perturb this variant (Barth et al, 2018). In addition, AV simulations showed that fluorophores at the TmrAB<sup>NBD</sup> labeling positions do not interfere with ATP- or substrate-binding sites, supporting the conclusion that FRET labeling does not affect ATP binding, hydrolysis, or transport. For TmrAB<sup>PG</sup>, however, equivalent transport data were not available, and AV simulations suggested interference of fluorophores with periplasmic gate dynamics. We therefore directly compared the transport activity of LD555/LD655-labeled TmrAB<sup>PG</sup> and unlabeled wild-type TmrAB using a single-liposome transport assay with the fluorescein-labeled peptide C4F (RRYC<sup>F</sup>KSTEL) (<sup>F</sup>, fluorescein; Fig. 1– Fig. S3a). Both variants showed indistinguishable transport activity, demonstrating that fluorophore conjugation at the periplasmic gate preserves transport function.

      (3) Line 117 and Figure S1c. Please add the reference for consistency of ATPase activity with previous studies on TmrAB.

      We have added a reference to previous biochemical studies reporting comparable ATPase activity and kinetic parameters for TmrAB to support the consistency of our measurements.

      (4) Line 119: It mentions that "Cysteine-maleimide labeling of detergent-solubilized TmrAB achieved site-specific labeling efficiencies exceeding 90%". The legend of Figure S1d mentions about labeling efficiency in the range of 40-50%. A clarification will be helpful for the reader. Also, calculations can be extended to the ratio of LD555 and LD655 labels on the molecule, which can be considered in analyzing results.

      We apologize for the lack of clarity. The reported >90% labeling efficiency refers to the site-specific cysteine labeling efficiency per accessible site, as determined by dye incorporation. In contrast, the 40–50% values shown in Fig.1–Fig. S1d reflect the per-site efficiency for donor-lonely and acceptor-only populations respectively, which together account for the >90% overall labeling efficiency. We have revised the main text and figure legend to clearly distinguish between per-cysteine labeling efficiency and the fraction of correctly double-labeled molecules. We also clarify that only complexes with appropriate donor– acceptor stoichiometry were included in the smFRET analysis.

      (5) Figure 1: Line 627: This line mentions "For all simulations, TmrA is shown in blue with LD655 (orange) and TmrB in yellow with LD555 (green)." Is it (which label on which subunit) known for the experimental setup?

      We thank the reviewer for pointing out this potential source of confusion. In the experimental system, fluorophore attachment occurs stochastically. Therefore, the assignment of donor and acceptor dyes to specific subunits is random. The representation shown in Figure 1 reflects one possible configuration for visualization purposes only. We have clarified this explicitly in the figure legend to avoid misinterpretation.

      (6) Figure S1-2a. Tau value can be better represented in a graph for visual readers instead of in the form of a table, and a dotted line with the threshold (~1 ns) will give a better representation of no change. Values can be included in the graph as well.

      We appreciate this helpful suggestion. We have revised Figure S1-2a to include a graphical representation of fluorescence life times, including a reference line around ~1 ns to facilitate visual comparison. Numerical values are retained alongside the plot for completeness.

      (7) Figure 2a: Each component of the assembly has been pointed with an arrow, which can mix two components and confuse readers. It would be good to make a legend column on the left or right and depict or indicate each component of the assembly clearly.

      We have changed the labeling in Figure 2a to improve clarity by separating the components and introducing a clearer legend layout, ensuring that each element of the assembly is unambiguously labeled.

      (8) The physiological concentration of ATP can range up to 5-10 mM. A comment on choosing the ATP concentration specifically to be 3 mM would be useful for the readers.

      We appreciate this suggestion. While intracellular ATP concentrations can reach up to 5–10 mM, values around 3 mM are commonly used as physiologically relevant conditions in in vitro biochemical and biophysical studies. We selected 3 mM ATP as a representative near physiological concentration that ensures saturation of ATP-dependent conformational transitions while remaining comparable to previous studies on TmrAB (Hofmann et al, 2019; Nocker et al, 2026; Nöll et al, 2017; Stefan et al, 2020). We have clarified this rationale in the manuscript.

      (9) Figure 2c is not cited in the text.

      We thank the reviewer for noting this oversight. Figure 2c is now explicitly cited in the main text.

      (10) Results in Figure 2 and 3 have been analyzed using 2 and 3 Gaussian distributions, respectively. It would be good to explain the rationale for it.

      We appreciate that this important point was brought to our attention. The number of Gaussian components was determined based on the minimal model required to describe reproducible and structurally supported populations. For ATP titration experiments (Figure 2 and Figure 3), two populations (apo and ATP-bound) were sufficient and consistent across replicates. In contrast, three populations were required under trapping conditions (Figure 4), where an additional state (OFF<sup>open</sup>) becomes kinetically stabilized and clearly resolved. We have clarified this rationale in the manuscript.

      (11) Figure 3b: data points do not seem to be saturated with respect to ATP concentration. It needs more points beyond 3 mM. Different K<sub>d</sub> at different sites in the structure could represent differential local dynamics over the structure.

      Previous structural studies demonstrated that 1 mM ATP is sufficient to saturate both nucleotide-binding sites under trapping conditions (Hofmann et al, 2019), indicating that the concentration range used here is adequate. Consistent with this, both ensemble and smFRET measurements approach saturation by 3 mM ATP, a near-physiological condition commonly used in biochemical studies. While additional data points above 3 mM could further define the plateau, they are unlikely to alter the mechanistic conclusion. We have clarified this point in the manuscript.

      (12) Figure 3 and Figure 1 - S1 have two different Kd values with respect to ATP concentration; both of these graphs measure conformational changes using smFRET. A comment specifying these Kd values based on single molecule verses ensemble measurement from will be helpful for readers.

      We appreciate this important point and have clarified it in the manuscript and the response to Reviewer #1 above. The K<sub>d,ATP</sub> values in Fig. 1–Fig. S1 are derived from ensemble FRET measurements, whereas those in Fig. 3 are obtained from smFRET population analysis. This difference likely arises from the difficulty of deconvoluting overlapping FRET populations at sub-K<sub>d,ATP</sub> concentrations, particularly for TmrAB<sup>PG</sup>, where state assignment is less well separated. Despite this quantitative offset, both approaches consistently indicate ATP saturation well below physiological concentrations and therefore support the same mechanistic conclusion that ATP binding drives conformational switching in TmrAB. We now explicitly distinguish these methods and their interpretation in the manuscript.

      (13) Figure 4: Slow-turnover TmrAB mutant has been employed in cysteine mutant on the PG opening side, but not towards the NBD side. Either experimental data or a comment on not pursuing it would be helpful for the reader. Similarly, experiments in the presence of peptide and in the absence of ATP, which can help to understand the role of substrate in conformational dynamics in the absence of ATP, are not pursued in this study. Along similar lines, experiments with wild type, in the presence of MgADP +/- substrate, are not shown in this study.

      We thank the reviewer for these insightful suggestions. The slow-turnover variant was specifically applied to the periplasmic gate reporter (TmrAB<sup>PG</sup>) because this construct provides direct sensitivity to outward-facing conformations, which are central to resolving the OF<sup>open</sup> state. In contrast, the NBD reporter primarily monitors nucleotide-binding domain (NBD) dimerization and is less suitable for distinguishing periplasmic conformational differences.

      Experiments in the absence of ATP but in the presence of peptide, as well as MgADP ± substrate, would indeed be valuable for further dissecting substrate effects. However, these conditions are beyond the scope of the current study, which focuses on ATP-driven conformational dynamics and the identification of kinetically hidden intermediates. We have added a statement in the Discussion to acknowledge these possibilities as directions for future work.

      (14) Figure 4, peptide concentration has been varied in the right panel. The result can also be presented as the % of OFopen and OFoccluded state with increasing concentration of peptide.

      We thank the reviewer for this suggestion. While such a plot would indeed be informative and could improve our understanding of substrate binding and substrate-induced trans-inhibition, the current dataset does not contain sufficient data points to construct a reliable concentration-dependent curve, particularly given that peptide saturation was not reached in our experiments. The characterization of substrate binding is further complicated by the presence of two distinct substrate-binding sites one in the outward-facing and one in the inward-facing state with likely completely different K<sub>d</sub> values and would require a more complex binding model. We have therefore decided against including this plot in the current manuscript. We do acknowledge, however, that future smFRET studies with improved temporal resolution are particularly well suited to investigating substrate binding to TmrAB and its effects on conformational equilibrium, and we have noted this in the Discussion.

      Reviewer #2 (Recommendations for the authors):

      (1) In all figures, can you please label the transporter schematics with the conformational states they represent?

      We thank the reviewer for this suggestion. All transporter schematics in the main and supplementary figures have been updated to include clear labels indicating the corresponding conformational states, thereby improving clarity and consistency.

      (2) As a suggestion, it may improve clarity to include the labelling positions (residue numbers) directly in Figure 1a and b, even though they are provided in the legend.

      We appreciate this suggestion. Residue numbers corresponding to labeling positions have now been added directly to Figure 1a and b to improve readability and facilitate interpretation.

      (3) Lines 183-188: This is a key point. It would be helpful to include a reference line for the expected state (0.63). Interestingly, this value coincides with the shoulder observed in Fig. 3c (0.1 mM ATP). Is there an explanation for this (see also point 2)?

      We thank the reviewer for highlighting this point. We considered adding a reference line at 0.63 to the plot; however, we decided against it. While a subpopulation does appear at ~0.63 —consistent with the expected FRET efficiency of the OF<sup>open</sup> conformation—it is only present in a single condition (0.1 mM ATP) and is not observed across other ATP concentrations for this TmrAB variant. It more likely reflects a minor non-reproducible subpopulation or photophysical artefact, in line with our response to Point 2 of the public review (Reviewer #2).

      (4) The final section of the Results section seems like an afterthought, especially since the heading suggests a broader scope.

      We appreciate this comment. We have revised the final section of the Results to improve its structure and ensure that the scope indicated by the heading is fully reflected in the content. This section now more clearly integrates kinetic and thermodynamic aspects of the transport cycle.

      References

      Asher WB, Geggier P, Holsey MD, Gilmore GT, Pa; AK, Meszaros J, Terry DS, Mathiasen S, Kaliszewski MJ, McCauley MD, Govindaraju A, Zhou Z, Harikumar KG, Jaqaman K, Miller LJ, Smith AW, Blanchard SC, Javitch JA (2021) Single-molecule FRET imaging of GPCR dimers in living cells. Nat Methods 18: 397–405. doi:10.1038/s41592-021-01081-y

      Asher WB, Terry DS, Gregorio GGA, Kahsai AW, Borgia A, Xie B, Modak A, Zhu Y, Jang W, Govindaraju A, Huang LY, Inoue A, Lambert NA, Gurevich VV, Shi L, Lefkowitz RJ, Blanchard SC, Javitch JA (2022) GPCR-mediated beta-arrestin activation deconvoluted with single-molecule precision. Cell 185: 1661– 1675 e1616. doi:10.1016/j.cell.2022.03.042

      Barth K, Hank S, Spindler PE, Prisner TF, Tampé R, Joseph B (2018) Conformational coupling and transinhibition in the human antigen transporter ortholog TmrAB resolved with dipolar EPR spectroscopy. J Am Chem Soc 140: 4527–4533. doi:10.1021/jacs.7b12409

      Barth K, Rudolph M, Diederichs T, Prisner TF, Tampé R, Joseph B (2020) Thermodynamic basis for conformational coupling in an ATP-binding cassette exporter. J Phys Chem LeJ 11: 7946–7953. doi:10.1021/acs.jpclett.0c01876

      Hofmann S, Januliene D, Mehdipour AR, Thomas C, Stefan E, Brüchert S, Kuhn BT, Geertsma ER, Hummer G, Tampé R, Moeller A (2019) Conformation space of a heterodimeric ABC exporter under turnover conditions. Nature 571: 580–583. doi:10.1038/s41586-019-1391-0

      Husada F, Bountra K, Tassis K, de Boer M, Romano M, Rebuffat S, Beis K, Cordes T (2018) Conformational dynamics of the ABC transporter McjD seen by single-molecule FRET. EMBO J 37: e100056. doi:10.15252/embj.2018100056

      Levring J, Terry DS, Kilic Z, Fitzgerald G, Blanchard SC, Chen J (2023) CFTR function, pathology and pharmacology at single-molecule resolution. Nature 616: 606–614. doi:10.1038/s41586-023-05854-7

      Nocker C, Pečak M, Nocker T, Fahim A, Sušac L, Tampé R (2026) Single-molecule dynamics reveal ATP binding alone powers substrate translocation by an ABC transporter. Nat Commun 17 doi:10.1038/s41467-026-70021-1

      Nöll A, Thomas C, Herbring V, Zollmann T, Barth K, Mehdipour AR, Tomasiak TM, Bruchert S, Joseph B, Abele R, Olieric V, Wang M, Diederichs K, Hummer G, Stroud RM, Pos KM, Tampé R (2017) Crystal structure and mechanistic basis of a functional homolog of the antigen transporter TAP. Proc Natl Acad Sci U S A 114: E438–E447. doi:10.1073/pnas.1620009114

      Stefan E, Hofmann S, Tampé R (2020) A single power stroke by ATP binding drives substrate translocation in a heterodimeric ABC transporter. eLife 9: e55943. doi:10.7554/eLife.55943

      Wang L, Johnson ZL, Wasserman MR, Levring J, Chen J, Liu S (2020) Characterization of the kinetic cycle of an ABC transporter by single-molecule and cryo-EM analyses. eLife 9: e56451. doi:10.7554/eLife.56451

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      WIPI1 is a PROPPIN family protein that has been implicated in Retromer-mediated membrane fission events. Although the cargos that it has been tested to be important for are diverse, one of the cargos that is unaffected is Beta1-Integrin. This leads the authors to assess another PROPPIN family protein - WIPI2, which is a homolog of WIPI1. KD using siRNA is effective and had no consequences on LAMP1, EGFR trafficking or GLUT1 trafficking. Integrin-B1, however, had a large and significant defect in its recycling from the endosome, with a clear endosomal colocalisation. Complementation experiments with WT WIPI2 recovered the phenotype, but various mutant WIPI2 complements resulted in elongated tubules, and there was also a dominant negative effect of the mutant. Integrin is a classic retreiver cargo, so the authors rationalise that WIPI2 may be playing a role with retreiver that WIPI1 plays with retromer. To assess this, they perform a set of immunoprecipitations. SNX17, the retreiver-associated sorting nexin, co-IPs with WIPI2 in a VPS26C-dependent manner. VPS26C but not VPS26 co-IPs with WIPI2, and the reciprocal with WIPI1. These interactions were not present for the FSSS mutation of WIPI2. WIPI2 localises to Rab11 endosomes mainly, as does retriever. Mutations of WIPI2 not only affected WIPI2 localisation, but also VPS35L mutations, indicating that there is a functional relationship between the two.

      On the whole, I find the manuscript compelling. The manuscript is very clearly written, the results are convincing and well performed. The flow of experiments is logical, and although not comprehensive in the subsequent mechanistic understanding, the fundamental findings are important and convincing. My comments below are, on the whole, minor and are intended to support the communication of the findings to the field.

      We are happy that the reviewer has received our work quite positively.

      (1) The IP interaction data were convincing; however, for me and some others, an interaction is only convincing when performed in vitro, and understood at a structural level. I do not suggest the authors do that in this case; however, I think, at a minimum, some sensible moderation of claims would be useful here.

      Indeed, quantitative in vitro data on the affinities would be a nice addition. However, we have significant trouble to recombinantly express and purify well-behaved WIPI2 in sufficient quantities for such studies. We keep working in this direction but are not there yet.

      We have now inserted a phrase into the discussion section highlighting this limitation: "Our immunoprecipitation assays cannot distinguish and more detailed structural and interaction studies with pure compounds will be necessary to elucidate the nature of this interaction". We nevertheless think that the the isoform specificity of the IPs, the effect of the point mutations in WIPI2 on these interactions, and the functional effects in vivo lend signficant support to the notion of a complex even if there is no proof of direct binding of WIPI2 to Retriever.

      (2) I found the final localisation data and its interpretation confusing. My interpretation of that data would not be that the retreiver is relocalised, but rather that there is less of both recruited to the membrane and the remaining localisation distribution is shifted. In addition, I am not quite sure of the model here - is the idea that WIPI2 recruits retreiver, if that is the case, I find it hard to resolve with its role as a mediator of fission. Clarity would be appreciated here.

      We are not quite sure what "final" localisation data the reviewer refers to, but we guess it is Fig. 9. This figure primarily provides in vivo evidence supporting the connection between Retriever and WIPI2. It does this by showing that the S67 substitution shifts both proteins. In WIPI2 wildtype cells, WIPI2 and VPS35L strongly colocalize in Rab11 compartments. S67 substitutions in WIPI2 abolish this localisation; WIPI2 shifts mainly to Rab5 compartments, where VPS35L shows only a moderate increase, and to Rab7 compartments, where VPS35L shows no increase at all.

      We do not understand the reviewer's interpretation that less Retriever would be recruited to the membranes in the S67 variants. VPS35L remains completely associated with punctate, presumably membrane-bounded structures also in the mutants, providing no evidence for a detachment from the membrane. The same is observed in a WIPI2 knockdown. Therefore, we did not claim that WIPI2 is the main factor recruiting Retriever to the membrane, for which our experiments yield no hints. This does not exclude that the interaction of WIPI2 could strengthen membrane recruitment, or that two pools of Retriever exist, one interacting with Snx17 and another interacting with WIPI2, and that both link to each other in a coat. We did not dwell on this in the discussion because our experiments cannot distinguish these possibilities and were not conceived to analyse membrane recruitment of Retriever.

      (3) I am concerned that the repeats being compared for statistical analysis are not biological repeats but technical repeats (cells in the same experiment). I should think the idea of the statistical comparison is to show experimental reproducibility and variability across biological repeats. Therefore, I would expect an appropriate number of biological repeats (3 or more minimum), to be the data compared in the statistical analysis and graphs. I think it is appropriate to average the technical repeats from each biological repeat. I find these to be useful resources https://doi.org/10.1083/jcb.202401074, https://doi.org/10.1083/jcb.200611141

      The repeats being compared are biological repeats from independent experiments. This is described in Methods, where the reviewer may not have seen it. In order to make the independent experiments more evident in the figures, we have now colour coded the individual cell measurements from the three independent experiments. This allows to visualize both the individual data points, the average from each experiment and the variability across the independent experiments.

      Reviewer #2 (Public review):

      Summary:

      The manuscript from De Leo and Mayer presents evidence that the PROPPIN protein, WIPI2, associates with the Retriever complex, and is required for the proper transport of the SNX17-Retriever cargo, beta1-integrin. This finding fits with prior papers from the Mayer lab, which showed that a related PROPPIN, WIPI1, is required for the transport of some SNX27-Retromer cargo, including GLUT1. The retromer and retriever complexes are architecturally similar. Importantly, they act at the same endosomes, and each transports cargo from endosomes to the plasma membrane. Thus, the possibility that each also requires a structurally related PROPPIN is of interest. However, the manuscript is incomplete, and the main claims are only partially supported.

      Strengths:

      The topic that PROPPIN proteins are important for the function of the Retromer and Retriever complexes expands our view of the trafficking complex.

      Weaknesses:

      Many important controls are missing. Several points that are made in the manuscript are only supported through a single approach.

      We made a serious effort and implemented many suggestions of this reviewer, but orthogonal approaches are not always available or accessible.

      Reviewer #3 (Public review):

      Summary:

      The manuscript of Mayer and colleagues analyzes the function of WIPI proteins in mammalian cells. The authors previously identified CROP as a complex consisting of WIPI1 and the retromer complex, primarily in yeast cells. In mammalian cells, both WIPI1 and WIPI2 exist, whereas retromer has a homologous complex termed retriever. They now find that WIPI2 can form a complex with retriever subunits. They named this complex CROP2. Their data further indicate that CROP2 and CROP1 have distinct substrate specificities as knockdown of CROP2 subunits affects beta1 integrin sorting, whereas knockdown of CROP1 affects EGFR and GLUT1. They further identify a similar sequence (FSSS) in both WIPI1 and WIPI2, which is required for their specific binding to retromer and retriever.

      Strengths:

      CROP1 and CROP2 seem to use similar features for their formation, and have different substrates, which is convincingly shown.

      Weaknesses:

      The analysis lacks information that this is a complex as claimed. It can be deduced from the interaction analysis, but was not shown.

      It is of course desirable to obtain a detailed structural and in vitro characterisation of this interaction, which we have not provided because we currently do not have sufficient amounts of well-behaved source material for this. We nevertheless think that the interaction we show, which is strictly isoform-specific and dependent on single amino acid substitutions in a motif that in CROP1 is necessary for the interaction its recombinant subunits, supports that CROP2 is a similar a complex. We don't show a direct interaction but also don't claim in the manuscript that the interaction between WIPI2 and Retriever is direct and independent of additional factors.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you will see, the reviewers generally value the contribution to the field, but they feel that some claims require additional experimental support.

      (1) I have summarized the major points below.

      (a) Both reviewers 1 and 2 agree that the quality of localization data presented in Figure 9 and S5-S7, and the interpretation of the data, could be improved. See comment 2 from reviewer 1 and comments 23, 24 and 25 from reviewer 2. They not only suggest ways to improve the presentation of the data, but additionally suggest improving the staining of the Rab11 marker and additionally explain the lack of co-localization between VPS35 and Rab5, which has been reported in the literature.

      This impression was due to the fact that some figures showed projections of image stacks, which was not indicated clearly in the figure legend. We have changed this and now show single image planes throughout all figures.

      (b) Both reviewers 1 and 3 note that the evidence supporting a functional WIPI2-Retriever complex in vivo is currently weak. We agree that additional biochemical data demonstrating the presence of the CROP1 and CROP2 complexes in vivo would strengthen the central message of the paper and elevate it to a more fundamental discovery.

      We understood that the reviewers did not ask for further in vivo evidence but would welcome structural characterisation of the complex and quantitative binding data in vitro with purified proteins. Structural characterisation is out of scope of our study and in vitro binding studies have remained hampered by the fact that WIPI2 is hard to express and purify and not well behaved in vitro.

      (c) All reviewers agree that the authors should carefully repeat their statistical analysis to account for the number of biological replicates. Reviewer 1 suggests publications that the authors could refer to.

      The reviewers have probably overlooked the respective description in the methods section, where it had been stated that we analysed biological replicates from independent experiments. In graphs showing measurements from individual cells we now make this evident through colour coded dots, in which each colour represents data points stemming from an independent experiment. This makes it evident that the variance from experiment to experiment is low. The means (n = 3) were generally compared using a two-tailed unpaired t-test.

      (d) Reviewer 2 additionally has various minor points that would greatly improve the readability and presentation of the work, and we recommend addressing (comments 1, 2, 3, 4, 12, 15, 17, 20, 27, 28, 29). All reviewers, in general, provide great minor suggestions. It would be great if the CROP1 and 2 complexes could be clearly introduced in each figure. We also agree that the WIPI2 CT labelling is confused and should be changed to "control" or similar.

      Many of the points raised by this reviewer were actually quite minor or questions of personal preference, not major problems as stated in the review. Nevertheless, we found a number of useful suggestions in this review and have addressed these points as detailed in the response to reviewer 2.

      (2) In addition to the major shared concerns laid out in the points above, reviewer 2 has some further minor suggestions:

      (a) Comment 6. Could the author explain the discrepancies between the example blot shown in Figure 1D and the quantification (1E).

      The two have actually been quite consistent. The reviewer might have mistaken the marker lane as the 0 min reference value to arrive at this impression. We have now removed the marker lane to avoid this.

      (b) Comment 9 - could the authors clarify how surface labelling experiments were carried out?

      This had been clearly described in the methods section, where this reviewer has probably not seen it.

      (c) Comment 11 - The reviewer suggests normalizing the surface levels of markers to the cell area and not per cell. This is a reasonable suggestion.

      The analysis had already been performed as proposed. This had been clearly described in the methods section, which the reviewer may not have looked at.

      (d) Comment 19 "In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures." The authors could try a Rab4 or Rab11 overexpression plasmid to show whether these are elongated recycling tubules.

      This has now been added.

      Reviewer #1 (Recommendations for the authors):

      Minor comments:

      (1) The figures are not colourblind friendly, and should be changed to be so. Additionally, single colour images should be grayscale.

      That was a good learning opportunity. We adapted the colour schemes of the images to make them more colourblind friendly, now using magenta, green, and white for the overlaps. In doing so we have relied on published recommendations, but we have not found a colourblind colleague to check the efficacy of this change.

      (2) WIPI2^CT labels are confusing, as people may think they are a mutant. I suggest changing to "control" or similar.

      These have been changed.

      (3) "The effect was comparable to that of a knockdown of SNX17 (Figure 3 A, B)." On page 6. Based on this sentence, I was expecting to see a comparison to SNX17 KD, but it was not there as far as I can tell.

      This statement referred to a publication by P.Cullen and collaborators. We have changed the wording and inserted the (missing) reference to make this clear.

      Reviewer #2 (Recommendations for the authors):

      The manuscript is modest. In addition, many of the claims should be better supported by the addition of orthogonal data. Moreover, the quality of some of the data presented needs to be improved. Overall, the manuscript requires better descriptions of the methods. In many figures, it was not clear how the experiments were performed.

      The experimental descriptions that the reviewer refers to had been provided in the Methods section, where this reviewer may have overlooked them.

      The paper should also be better organized. Some less important findings are in the main figures, whereas some critical results are in the supplemental figures. In addition, there were multiple issues with the readability of the paper, and the authors should consider using a professional editor to make the paper easier to read.

      We had given the paper to colleagues who found it clear, and also Reviewer 1 has underlined its clarity. Nevertheless, we have re-phrased the manuscript in some parts to optimise it.

      One of the main claims in the paper is that the FSSS motif of WIPI2, as well as a conserved amphipathic helix, is critical for WIPI2 function in the CROP2 complex. It is notable that these are the same regions that are also critical for the role of WIPI2 in autophagy (Gubas et al., 2024 PMID: 39152217). The authors should include this information in the manuscript and cite the paper.

      Indeed. We mention this now in the introduction of the revised version.

      Additional Major Issues:

      While some of the issues raised below are actually minor and/or matters of personal preference, several comments led us to improve and correct the figures and we thank this reviewer for the constructive suggestions.

      (1) In Figure 1, it appears from the representative images that WIPI2 KD cells have higher levels of EGFR (Figure 1A and 1B). Is this correct?

      To some degree. This increase is not systematic. A moderate increase has been observed only in 2 experiments out of 4. Therefore, we did not investigate this.

      (2) Also in Figure 1, the colocalization is difficult to see. The authors should add the separate channels in addition to the merged images. Since the point is supposed to be that there is no impact on EGFR, all of this data could go into the supplement.

      We had considered this already for the original version but dismissed the idea. The overlap is quantified in Fig. 1C, which provides the relevant values from four experiments. Fig. 1A/B provide only sample pictures, which also permit to see overlap (yellow) 0 and 5 min after the induction of degradation, which vanishes at later timepoints. Separating the channels would quadruple the space that this figure occupies, which would not be practical and not change the point to be made.

      (3) The scale bars for each panel differ from each other. To better assess the data, the exact same magnification should be shown for each panel.

      Corrected

      (4) Figure 1C is confusing. The authors should explain which lines correspond to EEA1 and LAMP1.

      Corrected

      (5) In Figure 1D, the authors show different blots for control and WIPI2 KD. Could the authors compare WIPI2 and EGFR in the same blot? Without a comparison on the same blot, it is impossible to know whether the starting levels of EGFR are the same. Moreover, the quantitation in Figure 1E sets the value for each cell line to 100%. Instead, the starting levels in each cell line should be compared. The authors should use the amount of EGFR at zero time in the control cells to define 100%, and then indicate the relative initial EGFR levels in the WIPI2KD cells.

      A new blot is shown now and the quantification has been performed as proposed.

      (6) The quantification in Figure 1E does not match the representative blot shown in Figure 1D. According to the graph, the rate of degradation of EGFR is similar in both cell lines. But the representative blot shows that there are large differences.

      We do not understand this comment. The representative blot shows similar kinetics for both. Perhaps the reviewer got confused by the fact that a marker lane was still present on the left blot and not labelled as such. The new version of the figure corrects this.

      (7) The blot showing the WIP2 knockdown in Figure 1D has a lot of background. However, the blot of the WIPI2 knockdown in Figure S1 looks very good. The authors should make sure that they load enough sample and use a good antibody for the experiments in Figure 1.

      The new blot that we added in response to comment 5 corrects this.

      (8) In Figure 2 and Figure 3A, the cells are too confluent. This is an issue because the cells might not be metabolically active. In addition, the signal is saturated. The authors should make sure that all of the data is collected on cells that are not too confluent.

      The confluency of the culture cannot be judged from single frames, which were selected to show several cells. We had controlled confluency and underlined in the Methods section that “For microscopy, the cells were plated on 18-mm-diameter glass coverslips on 24-well plates and grown for 2 or 3 days according to the protocol of DNA or siRNA transfection by reaching a confluency of 70-80%”. The reviewer may not have seen this.

      (9) One main issue with these figures, especially the non-permeablized cells, is that it is impossible to assess how much of the signal is on the cell surface. The authors should provide the methods that they used to prevent inadvertent permeabilization of the cells. Were these experiments performed at 4 degrees? The authors should include a control of an antibody to a protein that is not found on the cell surface.

      There is an internal control in that the non-permeabilised WIPI2KD cells, which have been treated with the same antibody, show no much less staining than the control cells (Fig. 3A). In WIPI2KD cells, integrin becomes accessible for antibody staining only upon detergent permeabilization. This demonstrates that our procedure does not lead to significant inadvertent permeabilization of the cells.

      (10) The authors should perform surface biotinylation assays as an orthogonal approach to determine GLUT1 levels and beta1-integrin levels at the cell surface, respectively.

      There is a strong, qualitative difference in the surface labelling of beta1-integrin that is not observed for GLUT1. Given that, it is not obvious to us what additional argument would be provided by surface biotinylation or subfractionation experiments.

      (11) In quantifying surface levels of GLUT1 or beta1-integrin by microscopy, the authors should normalize to the cell area, rather than per cell.

      The reviewer has probably not seen that the Methods section states that the cell area has been used for normalisation.

      (12) In Figure 3, the nuclear DAPI stain in the KD cells is much less bright than in the control cells. The authors should make sure to choose representative images.

      The nuclear DAPI signal has been visible in all cells. Depending on the position of the nucleus, is shape and dimension in the z-direction, individual nuclei can show different degrees of staining. The images shown are representative. We have adjusted the settings now to make the nuclei in the WIPI2KD cells easier to spot.

      (13) For the immunofluorescence studies, the authors should be using single z planes rather than maximum projection.

      Images have been exchanged by single planes.

      (14) For the experiments in Figure 3, the authors should check the total levels of EEA1 and LAMP1 by western blot to test whether WIPI2 KD affects the levels of these proteins. If these organelle marker proteins are impacted, this could impact the colocalization measurements shown in Figures 3C and D.

      We have measured the total fluorescence intensity of EEA1 and LAMP1 in the images. It shows no significant difference between control and WIPI2 knockdown cells (new Fig. 3F, H).

      (15) In Figure 4A, the helical representation is rotated in the WIPI2-Sloop; the orientation of the residues that are not mutated should stay the same.

      Yes. Done.

      (16) In Figure 4B and 4C, cells that were not transfected with WIPI2 WT or WIPI2 Sloop should be shown.

      Since the transfection efficiency is limited, the fields contain both non-transfected (lacking green fluorescence) and transfected cells (showing green fluorescence). We have now marked transfected cells with an asterisk.

      (17) The cells in the lower panel of 4B have an unusual morphology and are much more round. The authors should choose cells that are representative of each experimental condition.

      We now provide another field.

      (18) In Figure 4C, it looks like the magnification of the top panels is different from the bottom panels. The same magnification for all the panels should be shown (and the size of the scale bars should be the same.

      Corrected

      (19) In Figure S4, the authors observe tubular structures. The authors should perform immunofluorescence with endosomal markers such as EEA1, LAMP1 and Retromer to determine the nature of the tubulovesicular structures.

      We have done this (new Fig. S4). Rab4 is on tubules. Rab5 on the structures from which the tubules emanate.

      (20) In Figure 5A, the top scale bar is missing.

      Corrected.

      (21) In Figure 5B, the confluency is too high.

      See our response above. A single field does not permit to judge this. Confluency was controlled for all cultures. The cultures were not confluent.

      (22) The IP studies shown in Figures 6, 7 and 8, should be accompanied by colocalization studies.

      Colocalization measurments have now been integrated into the manuscript (Figs. S5, S6). They are consistent with the IP data.

      (23) Figure 9 was very confusing and should be broken up into multiple figures. Data showing that localization did not change in any of the cell lines can be put in figures that are distinct from figures that show that localization changed in the various mutants. Figures that show no change can go in the supplement.

      Since every panel of Fig. 9 shows a statistically significant difference we left the figure unchanged.

      (23) Representative figures should be shown in the same figure as the corresponding graph. In addition, the order of the colocalization data shown in the graphs and figures should match the order described in the text.

      We consider the graphs of Fig. 9 as the relevant information. Representative images are just illustration. Integrating them with the graphs would make it necessary to split everything up into multiple figures, making it harder to compare the different combinations. Therefore, we left the figures unchanged.

      (24) In Figure S7, the Rab11 signal looks continuous, which makes the colocalization analysis meaningless. The authors should determine how to take images that can be evaluated. On a more minor note, the zoomed panels should be labeled as well.

      This is a result of having shown a projections of multiple planes. The images have now been replaced by single plane images. Zoomed panels have been labelled and the scale bar added.

      (25) The low colocalization of VPS35L with Rab5 is surprising, as SNX17 has been previously shown to co-localize with early endosomes positive for EEA1. This result may have occurred due to overexpression because the authors chose to utilize plasmids that express a tagged protein. There are antibodies to each of the endogenous proteins, and this is what should be used for this set of experiments.

      This comment made us control the analysis performed for these images, which by mistake had been performed on z-projections rather than on single planes. This distorted the values. The re-analysed data shows a higher colocalisation with Rab5, but it remains inferior to colocalisation with Rab11.

      (26) The authors should determine whether β1-integrin colocalizes with WIPI2 in endosomal compartments.

      This was done. WIPI2 colocalizes with beta-integrin on EEA1-and SNX17-positive strcutures but not positive for LAMP1 (Fig. 3E/F).

      Minor points

      (27) In one of the panels in Figure 1A, "30 min" is duplicated.

      Removed

      (28) In Figures 5C and 5D, the y-axis should indicate that this is surface β1integrin.

      Changed and added “surface”

      (29) In Figure 9 there is a typo in panel A. It is VPS35L and not VPS35.

      Corrected

      Reviewer #3 (Recommendations for the authors):

      This is an overall convincing study, which shows that the two complexes, CROP1 and CROP2 function at different membranes and serve different substrates. While I agree with their localization analysis, I have one key issue. The authors claim that each of the two forms a complex and base this on their specific pull-down and western blot analyses.

      I find it important that they show that both indeed form stable complexes in vivo, using pull-down and mass spectrometry approaches. They have all the necessary tools in hand and could use WIPI1 and WIPI2 to demonstrate the existence of the two complexes. The FSSS mutants of each are good controls for such an analysis.

      The manuscript actually presents the demanded in vivo experiments. Figs. 6 to 8 show pull-downs of WIPI1 and WIPI2 from cells, including also the FSSS mutant. While we haven't analysed this interaction by mass spectrometry, the Western blot analysis confirms the analysis. Cooperation of these proteins is further supported by the in vivo phenotypes, where the S67A substitution in WIPI2 produces a similar phenotype on integrin beta1 localisation as inactivation of Retriever.

      A second aspect is the general presentation. The paper would be a lot more accessible if the subunits of each complex (CROP1 and CROP2) were also introduced in the figures of each part. For readers, a final model is helpful to put the data into context and show where each complex operates in the cell.

      We have introduced a scheme of the respective complexes, including the names of the compunds, in Figs. 6 and 7 to avoid confusion.

      Finally, it is not clear how the statistics compare to repeats in their data. This should be clarified.

      This had been described in methods. Statistics has always been done on biological replicates stemming from independent experiments. We have added a cartoon (Fig. 10) depicting the trafficking pathways affected by CROP1 and CROP2.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Comments from Reviewing Editor:

      I want to share that both reviewers appreciated that this revision has appropriately addressed many of the concerns they raised. However, reviewers concurred that additional wet-lab experiments which validated the findings would have made the work much more impactful; and their concerns about the quality of chromatin accessibility data appear not to be fully resolved. Might I suggest a textual revision that specifically points out these caveats, if you are not able to provide additional data? This would then proceed to VOR without additional need to review. Thanks much for your patience while I assessed the manuscript claims and reviewer opinions.

      The changes were very minor (2 sentences in the Discussion and a small section in the Supplementary Notes). It would be great if we could proceed to the VOR stage.

    1. Author response:

      We appreciate the time and attention to our manuscript and the feedback from the reviewers, who were overall supportive of the work. Both reviewers validated the technical approach we used to differentiate the wild-type (WT) and knockout (KO) neurons noting: “The combination of sparse Cre delivery with channel rhodopsin-mediated optotagging in Npas4 fl/fl:Ai32 mice is technically elegant” and “the rigorous optogenetic tagging strategy used to distinguish KO from WT neurons in vivo makes the single-cell comparisons much more convincing.” Furthermore, they note the consistency of the reported results, stating: “The reported phenotype is internally consistent and converges on a coherent story”.

      Both reviewers also pointed out several concerns or points of improvement for the manuscript. Below, we first offer several scientific and methodological clarifications that we believe resolve a number of the reviewers' concerns. We then outline which remaining points we plan to address through revision, and which fall outside the scope of the current study.

      Scientific Clarifications:

      Request for a standard housing control. Both of the reviewers brought up the long-term enrichment paradigm (EE) we opted to use for this study and expressed interest in seeing data from standard housed (SE) animals. This is an approach the lab has taken in its slice physiology work [1-3], where comparing EE and SE conditions has revealed important differences between cellular phenotype. However, the in vivo experiments described here differ in a key way: obtaining these recordings requires extensive handling, training, and daily transport between the vivarium, home cage, and behavior room. These experimental steps themselves constitute the kind of novel, salient experience known to induce NPAS4, making a true SE comparison unattainable within this paradigm. In our experiment, mice were housed in EE as a supplemental, well-established strategy to induce NPAS4 in CA1 pyramidal neurons but we believe the behavior alone would be sufficient. We will describe this more clearly in the text of the manuscript.

      Consistent with this view, place fields recorded from wild-type mice in other studies using SE but undergoing comparable handling and training procedures, are similar in size, spatial information, and stability to the WT place fields we reported here [4,5]. As part of our revisions, we will consider statistical comparisons between our WT neurons and those reported in other studies to quantitatively assess whether a difference exists.

      More broadly, we note that the existing literature on NPAS4 induction does not, to our knowledge, establish a baseline level of NPAS4 expression in CA1 pyramidal neurons in the complete absence of behavioral experience. Reports of NPAS4 expression in CA1 have generally relied on animals exposed to some form of salient or novel experience [3,6,7], consistent with our framework that NPAS4 induction reflects behaviorally-driven activity rather than a constitutive baseline.

      Expression profile of NPAS4. Reviewer #2 brought up a concern about the extent of the NPAS4 expression, referring to the IHC results in Figure 1A stating: “Even under EE, only a few percent of CA1 pyramidal neurons express detectable NPAS4 at any given moment (Figure 1A), yet the AAV strategy deletes the gene in 30 to 60 percent of pyramidal neurons. In effect, the majority of cells classified as KO in this study would not have been expressing the protein under the relevant conditions.” We wish to clarify two points here. First, in the experimental paradigm used to obtain the IHC results, mice were exposed to enrichment for only 90 minutes while in the in vivo physiology paradigm, mice were housed in an enriched environment (with frequent toy changes to ensure novelty) for weeks. Thus, NPAS4 is almost certainly expressed in a much larger percentage of WT neurons in mice that were kept in chronic enrichment and used for the in vivo studies. Second, while the NPAS4 protein is only expressed in cells for several hours following neuronal activity, it initiates an inhibitory synapse phenotype that persists long-term. Thus, even though a small percentage of neurons are NPAS4+ in the IHC results, it is likely that a much larger percentage of them have expressed NPAS4 in the past and now show the inhibitory synapse phenotype. Evidence for this comes from the slice physiology results in Figure 1C (and see similar results from adolescents [1-3]) in which animals were housed in enrichment long-term and differences between inhibition persisted in nearly every WT/KO comparison.

      We also recognize the related possibility that NPAS4 expression may not be uniform across the pyramidal cell population, but may instead concentrate in particular functional subtypes, such as cells with higher firing rates or stronger spatial tuning. As part of our revisions, we plan to test this directly by stratifying the KO population by firing rate and relating it to the magnitude of the observed phenotype. Taken together, we believe that while only a small fraction of CA1 pyramidal neurons are NPAS4+ at any given moment, a much larger fraction have experienced NPAS4 induction and the accompanying synaptic reorganization over the timescale of chronic enrichment making the WT/KO comparison in this study substantially less diluted than the IHC snapshot alone would suggest.

      Timeline of NPAS4 expression and synaptic reorganization. Reviewer #1 pointed out that this study only examines the effects of NPAS4-deletion on longer timescales (weeks to months after the virus expression and subsequent knockout) stating “[the study] is less definitive about the immediate causal sequence by which NPAS4 induction alters inhibition and reshapes spatial and temporal coding”. The reviewer is correct, the temporal relationship between NPAS4 expression, changes in synaptic inhibition, and changes in neuronal firing are important outstanding questions in the field. Currently, we lack molecular tools that would enable us to clearly test these relationships but with our existing, albeit limited information, we have the following working model.

      When an animal is placed into a new context, a subset of CA1 pyramidal neurons will fire action potentials in a spatially refined manner. This activity will drive NPAS4 expression in those neurons, resulting in protein expression that persists for a couple of hours before the protein is degraded.

      Following expression, NPAS4 will bind to various sites in the genome and initiate a genetic program which results in changes in inhibition recruiting CCK basket cell synapses to the soma and destabilizing CCK dendritic synapses. The exact mechanism behind this reorganization of inhibition is unknown, but the phenotype likely emerges over the course of several hours following NPAS4 expression and persists for days following the stimulus that induced NPAS4.

      While our chronic knockout approach does not allow us to resolve the precise timing of events in this sequence, it does allow us to ask a distinct and complementary question: what is the long-term consequence for a neuron that has never been able to execute this program? Our results demonstrate that NPAS4-deficient neurons which cannot initiate NPAS4-dependent inhibitory reorganization regardless of their activity history show systematic degradation in spatial and temporal coding precision. This establishes that the NPAS4-dependent inhibitory phenotype has lasting and functionally meaningful consequences for in vivo information encoding, a question that shorter-timescale or acute manipulations would not be well-positioned to address. Resolving the immediate causal sequence between NPAS4 induction, synaptic reorganization, and changes in firing will be an important goal for future work as new molecular tools become available.

      Behaviors that drive NPAS4 expression. Reviewer #2 pointed out that “NPAS4 is also induced by contextual fear conditioning and other paradigms which would predict context-specific effects rather than a uniform refinement function.” They are correct NPAS4 is expressed in response to different behavioral paradigms, including fear conditioning and environmental enrichment. However, the subregion in which NPAS4 is induced depends critically on the behavioral paradigm. When mice are exposed to contextual fear conditioning, NPAS4 expression is robust in CA3 and the dentate gyrus but negligible in CA1 [6]. This is consistent with the known activity patterns of these subregions: CA3 neurons are strongly recruited during contextually-dependent associative learning, while CA1 neurons are more reliably driven by exposure to novelty and respond in a spatially-refined manner. Consistent with this, studies using fear conditioning have focused on behavioral discrimination and synaptic changes in CA3 and granule cells [6]. To our knowledge no study has examined the relationship between fear conditioning, NPAS4, and CA1 pyramidal neuron function. Whether behavioral paradigms beyond environmental enrichment and spatial navigation can induce NPAS4 in CA1, and what consequences that might have for pyramidal neuron firing, are interesting questions for future work.

      We also wish to address the conceptual framing underlying this concern. In CA1, we do not believe that “context-specific effects” are separable from a “uniform refinement function.” CA1 pyramidal neurons respond in a context-dependent manner. When a mouse is placed onto a linear track, there is a subset of neurons that will increase their activity over the course of that exposure. But within this subset, individual neurons will also show spatially-refined responses firing action potentials as the animal runs through the corresponding place field. The spatial precision NPAS4 confers is always nested within context-dependent mechanisms NPAS4 refines whatever representation a neuron is already computing, rather than overriding the context-dependency of that representation. We therefore do not view these as competing frameworks.

      The role of NPAS4 in shaping CCK synapses. Reviewer #2 made the point that “the CCK to pyramidal cell connectivity that the authors invoke as the mechanistic anchor is also dense in standard housing, so the absence of detectable NPAS4 in SE conditions raises the further conceptual problem of how NPAS4-negative neurons would normally be innervated by CCK+ basket cells in the first place.” We wish to clarify that NPAS4 is not necessary for the formation of CCK synapses onto CA1 pyramidal neurons there are likely a number of NPAS4-independent mechanisms that regulate this synaptic connectivity (for example, see [8]). Rather, we place NPAS4 in the role of an activity-dependent modulator that acts on top of this baseline connectivity: when NPAS4 is expressed in response to neuronal activity, it shifts the balance of CCK inhibitory input along the somatodendritic axis, increasing somatic and decreasing dendritic CCK synaptic strength [1,2]. The question is therefore not how CCK synapses are established in the absence of NPAS4, but rather how experience-dependent activity uses NPAS4 to fine-tune the distribution of those synapses and it is this fine-tuning that our study links to the precision of in vivo spatial and temporal coding.

      Methodological Clarifications:

      Clarification on how stability analysis was performed. Reviewer #2 requested additional analysis for the stability results: “A control analysis using a fixed reference window around the original peak, rather than re-identifying the peak each epoch, would help distinguish a genuine plasticity-like shift from instability driven by noise.” We wish to clarify that this is precisely the methodology that was used in the manuscript. For the stability analysis shown in Figures 4C-E, the activity was aligned to the peak activity in epoch 1 such that 0 always represents the location of the peak in epoch 1. This approach allows us to identify how that activity differs in subsequent epochs, namely whether it has shifted relative to the activity in epoch 1. We will make this more clear in the results and methods sections.

      Request for Ai32 control. Reviewer #2 made the point that “The comparison throughout the manuscript pits Cre+ ChR2+ neurons (NPAS4 KO) against neighboring non-transduced neurons (WT). This is internally elegant, but leaves open the possibility that part of the phenotype arises from chronic ChR2 expression or constitutive Cre activity rather than from NPAS4 loss, especially given that most of the readouts are subtle.” We agree this would be the ideal control and regret that it is no longer experimentally feasible, as the laboratory in which these experiments were conducted is no longer operating. However, we believe several features of the existing dataset make a ChR2 or Cre artifact unlikely. First, the effects of chronic ChR2 expression are not known to produce the specific pattern of phenotypes we observe in particular the redistribution of somatic versus dendritic inhibition, which is recapitulated independently in acute slice recordings from animals that did not undergo optotagging procedures (Figure 1C). Second, the phenotype we report is internally coherent across multiple independent metrics: place field size, stability, signal-to-noise ratio, theta coupling, and phase precession all shift in the same direction, in a manner consistent with a specific change in inhibitory synaptic balance rather than a nonspecific effect of transgene expression. Third, the sparse nature of the Cre expression means that KO and WT neurons share the same local network, same LFP, and same behavioral context any network-level effect of Cre or ChR2 would be expected to affect both populations similarly. We will add a discussion of these points to the manuscript.

      PSTH clarification (unit of opto-response). To quantify the opto-response, we treated each light-on + light-off period (a total of 2 seconds) as the one trial. We aligned the trials by the light-on period, binned the spikes by 1 msec bins, and then summed the responses across trials to produce a histogram. From this histogram we found the maximum response during light off (e.g. the 1 msec bin with the greatest response which should be reported as number of spikes). We subtracted this from the maximum response during light on. Thus, the unit of opto-response should be spike counts. We will clarify this in the text and figures.

      Use of male mice. Reviewer #1 rightfully pointed out that this study only used male mice. In this study, we only used mice that were larger than 20 grams to ensure the mice could carry the weight of the implanted drives while performing the behavior. As this genetic line of mice is on the smaller size, only male mice were above this weight threshold. Importantly, slice work conducted in the Blood good lab has not identified sex differences in NPAS4 phenotypes [3,9]. Future studies would benefit from the use of both male and female mice. We will state this more explicitly in the text and expand on the potential implications of excluding female mice from our study.

      Future planned changes to manuscript:

      As the reviewers suggested, we intend to add the following analyses and make the following changes to the manuscript:

      Stratify key analyses (stability, theta coupling, phase precession) by FR to determine whether there is a dependency on the firing rate of cells.

      Apply hierarchical bootstrapping and add per-animal color-coding to supplementary figures to assess animal-level variability and protect against pseudoreplication.

      Add a circular-linear phase-position correlation analysis as an additional quantification of phase precession strength, complementing the existing slope-based analysis.

      Improve discussion around the temporal phenotype being downstream of the spatial one.

      Tighten mechanistic framing in the Discussion to more clearly distinguish what is demonstrated in this study from what is inferred from prior work, and to acknowledge the contributions of other inhibitory cell types.

      Minor changes and figure clarifications as noted by reviewers.

      Outside of the scope of this study or unable to be performed:

      There were several recommendations or points that the reviewers brought up that we do not have the resources to address. Nevertheless, we appreciate the reviewers noting these.

      SE control (as discussed above)

      Ai32 control (as discussed above)

      Behavioral consequences of NPAS4 knockout and the effects on learning and memory • Ripple analysis

      Drift observed in E4 and what this might look like over larger timescales

      Comparison between male and female mice to determine whether there are sex-dependence differences

      In conclusion, the reviewers recognized this as a well-designed and internally consistent study. We believe that many of the critiques including the request for a standard housing control, questions regarding the extent of NPAS4 expression across the pyramidal cell population, and points about the timeline of NPAS4 expression and synaptic reorganization are addressed by the clarifications provided in this response. We agree with many of the suggested analytical and textual changes and look forward to incorporating those into the revised manuscript.

      References:

      (1) Heinz, D. A., Cui, W., Cooper, K. L. & Bloodgood, B. L. Experience-induced NPAS4 reduces dendritic inhibition from CCK+ inhibitory neurons and enhances plasticity. J. Neurophysiol. 134, 361–371 (2025).

      (2) Hartzell, A. L. et al. NPAS4 recruits CCK basket cell synapses and enhances cannabinoid-sensitive inhibition in the mouse hippocampus. Elife 7, (2018).

      (3) Bloodgood, B. L., Sharma, N., Browne, H. A., Trepman, A. Z. & Greenberg, M. E. The activity dependent transcription factor NPAS4 regulates domain-specific inhibition. Nature 503, 121–125 (2013).

      (4) Sharif, F., Tayebi, B., Buzsáki, G., Royer, S. & Fernandez-Ruiz, A. Subcircuits of deep and superficial CA1 place cells support efficient spatial coding across heterogeneous environments. Neuron 109, 363–376.e6 (2021).

      (5) Quirk, C. R. et al. Precisely timed theta oscillations are selectively required during the encoding phase of memory. Nat. Neurosci. 24, 1614–1627 (2021).

      (6) Ramamoorthi, K. et al. Npas4 regulates a transcriptional program in CA3 required for contextual memory formation. Science 334, 1669–1675 (2011).

      (7) Chiaruttini, N. et al. ABBA+BraiAn, an integrated suite for whole-brain mapping, reveals brain-wide differences in immediate-early genes induction upon learning. Cell Rep. 44, 115876 (2025).

      (8) Früh, S. et al. Neuronal Dystroglycan Is Necessary for Formation and Maintenance of Functional CCK-Positive Basket Cell Terminals on Pyramidal Cells. J. Neurosci. 36, 10296–10313 (2016).

      (9) Lin, Y. et al. Activity-dependent regulation of inhibitory synapse development by Npas4. Nature 455, 1198–1204 (2008).

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this paper, the authors use a doxycycline-inducible DLD1 cell line expressing a Clover-tagged RNA-binding-defective TDP-43 2KQ mutant that forms nuclear "anisosomes" (TDP-43 shell with HSP70 core) to carry out a small-molecule screen using the LOPAC 1280 library to identify compounds that reduce anisosome number or shift their morphology and dynamics. They also conducted a genome-wide siRNA screen to identify genetic modifiers of anisosome formation and dynamics. From these screens, the authors identify pathways in RNA splicing, translation, proteostasis (proteasome and HSP90), and nuclear transport, including XPO1. They then focus on XPO1 as their primary hit. Pharmacological inhibition of XPO1 using KPT-276, Verdinexor, and Leptomycin B reduces anisosome number while enlarging remaining condensates, which retain liquid-like behavior by FRAP and fusion assays. XPO1 overexpression causes fewer, enlarged TDP-43 puncta, including cytoplasmic puncta, with little or no FRAP recovery, interpreted as gel or solid-like aggregates. Anisosome induction reduces detectable nucleoplasmic XPO1 staining. Finally, the authors examine a homozygous TDP-43 K181E iPSC-derived forebrain organoid model, showing increased cytosolic pTDP-43 in K181E/K181E organoids compared to wild-type controls. Chronic low-dose KPT-276 reduces cytoplasmic pTDP-43 without changing total TDP-43 levels. Bulk RNA-seq shows only a modest fraction of dysregulated genes in K181E/K181E organoids are rescued by KPT-276. They conclude that nuclear export, via XPO1, is a key regulator of TDP-43 liquid-to-solid phase transitions and that cytoplasmic aggregation per se may contribute only modestly to TDP-43 proteinopathy, with RNA-processing defects being dominant.

      We thank the reviewer for carefully summarizing our study.

      The study presents well-executed chemical and genome-wide siRNA screens in a DLD1 TDP-43 2KQ anisosome model and follows up on nuclear transport, particularly XPO1, as a modulator of TDP-43 phase behavior and cytoplasmic aggregation. The screens are impressive in scale, and the microscopy and fluorescence recovery after photobleaching (FRAP) work is technically strong. However, the central mechanistic and disease-relevance claims are not yet sufficiently supported. There are major concerns about the heavy reliance on non-physiological, RNA-binding-defective, and acetylation-mimetic TDP-43 (2KQ) and a homozygous TDP-43 K181E organoid model. An underdeveloped and partly contradictory mechanistic link exists between XPO1 and TDP-43 phase transitions in the context of prior work showing TDP-43 is not a canonical XPO1 cargo. The paper also appears to overinterpret organoid data to conclude that cytoplasmic TDP-43 aggregation plays only a minor role in pathology, based largely on pTDP-43 antibody staining with limited sensitivity and relatively modest rescue readouts. A deeper mechanistic analysis and additional, more physiological validation are needed for this to reach the level of rigor and impact implied by the title and abstract. The work feels screen-rich but conceptually underdeveloped, with key claims outpacing the data. A major revision with substantial new data and tempering of conclusions is warranted. I outline several problematic areas below:

      (1) The central mechanistic discoveries are derived almost entirely from a DLD1 colon cancer cell line overexpressing an RNA-binding-defective, acetylation-mimetic TDP-43 2KQ mutant and homozygous TDP-43 K181E iPSC-derived organoids. Both systems are far from physiological. The 2KQ mutation is a synthetic double lysine-to-glutamine mutant originally designed to mimic acetylation and disrupt RNA binding. In this study, essentially all cell-based mechanistic data on phase behavior, screens, and XPO1 effects rely on 2KQ. Yet there is no quantification of how much endogenous TDP-43 is acetylated in degenerating human neurons, nor whether a 2KQ-like acetylation state is ever achieved in vivo. It is not established that the phase behavior of 2KQ recapitulates the physiological or pathological phase behavior of wild-type TDP-43 or genuine disease-linked mutants, which may retain partial RNA binding and different post-translational modification patterns. As a result, it is difficult to know whether the modifiers identified here regulate a highly artificial 2KQ condensate or physiologically relevant TDP-43 condensates. To address this concern, the paper would benefit from quantifying endogenous TDP-43 acetylation at the relevant lysines in control and ALS/FTD patient tissue or more disease-proximal models such as heterozygous TARDBP mutant iPSC neurons, which would justify the focus on an acetyl-mimetic mutant. Key phenomena, including XPO1 dependence of phase behavior, effects of proteasome and HSP90 inhibition, and effects of splicing and translation inhibitors, should be tested for wild-type TDP-43 expressed at near-physiological levels and for one or more bona fide ALS/FTD-linked TARDBP mutants that are not acetyl mimetics. At a minimum, the authors should show that endogenous TDP-43 in neuronally differentiated cells exhibits qualitatively similar responses to XPO1 modulation, rather than exclusively relying on DLD1 2KQ overexpression.

      Acetylation of endogenous TDP-43 was reported by several studies. Although it occurs at low levels under normal conditions, TDP-43 acetylation is upregulated under stress conditions (e.g. oxidative stress and proteotoxic stress) (PMID: 25556531; PMID: 28724966). Importantly, Cohen et al. reported the identification of acetylated TDP-43 in ALS patient spinal cord (PMID: 25556531), while Yu et al. showed that endogenous wildtype TDP-43 undergoes demixing when neurons were treated with either a deacetylase inhibitor or proteasome inhibitor (PMID: 33335017). These studies also show that acetylated TDP-43 is defective in RNA binding and more prone to aggregation. Furthermore, ectopic expression of acetylated TDP-43 mimetics in cells and mice induces cellular defects similar to those observed in disease models (PMID: 28724966). Thus, our findings, based on previously established TDP-43 mimetics, should provide valuable information regarding the phase regulation of a disease-relevant TDP-43 mutant. We have included more background information to justify the use of TDP-43 acetylation mimetics in the introduction.

      (2) The organoid model is based on a homozygous K181E knock-in line. However, in patients, TARDBP mutations are overwhelmingly heterozygous. Homozygosity is thus a severe, arguably non-physiological sensitized background that may exaggerate nuclear RNA mis-splicing and phase defects and alter the relative contribution of cytoplasmic aggregation versus nuclear loss-of-function. In addition, it is not fully clear from this manuscript whether the structures in K181E organoids are bona fide anisosomes as defined in Yu et al. 2021, characterized by HSP70-enriched central liquid cores with TDP-43 shells and similar FRAP and fusion behavior to anisosomes in the DLD1 model. At present, the organoid section is framed as validation of "anisosome-bearing organoids," but the figures in this manuscript mainly show pTDP-43 puncta and total TDP-43 immunostaining, without detailed structural or biophysical characterization. The authors should explicitly compare heterozygous K181E/+ organoids or another heterozygous TARDBP mutant line with homozygous K181E/K181E organoids to assess whether XPO1 inhibition has similar effects in a genotype that more closely resembles patient genetics. They should provide direct evidence that the K181E condensates in organoids are anisosomes through HSP70 core immunostaining, three-dimensional reconstruction, and FRAP measurements, and clarify whether KPT-276 is acting on anisosome-like structures or more generic cytoplasmic aggregates or puncta. Without this, the leap from a DLD1 2KQ cancer cell model to human ALS/FTD-relevant neurons is not convincingly supported.

      The reviewer is correct that the use of homozygous K181E organoids generates a background that is more sensitive for detecting phospho-TDP-43. The goal was to test whether XPO1 inhibition mitigates the phosphorylation of a TDP-43 disease mutant. For this purpose, we believe that our experimental setup is suitable. We agree that we should not extrapolate the result to over emphasize on its disease connection. We have revised the paper to tone down this section. We also remove the RNAseq data as it is not essential for our conclusions.

      It is also noteworthy that TDP-43 disease mutations are usually loss-of-function alleles. Although heterozygous background is sufficient to induce disease phenotype in aged humans, heterozygous background in experimental settings is usually unable to generate severe defects. Thus, it is quite common to study TDP-43 disease-related defects in homozygous knockout or RNAi-mediated depletion conditions (e.g. PMID: 35197626; 41120751; 38277467).

      Regarding the immunostaining signals in K181E organoids, we did not report them as anisosomes. As documented in the literature, p-TPD-43 is widely used as a marker to indicate pathological TDP-43 aggregation. P-TDP-43 is enriched in pathological aggregates in human ALS and FTD patients, colocalized with other aggregation signatures such as ubiquitin and other aggregation-prone proteins in the cytoplasm (PMID: 36008843), and is being used as a diagnostic marker for neurodegeneration (PMID: 31661037). The characterization of K181E organoid is reported in a pre-print by Zhang Q. et al., 2026 (PMID: 41292965), which is currently under revision for Science Advances. In Fig. 1I of this manuscript, we confirmed the cytosolic localization of p-TDP-43 in cells that were isolated from K181E organoids. In the current manuscript, Figure 7 is to show that nuclear export inhibition mitigates the accumulation of p-TDP-43 in a brain-like tissues. We revise the subheading and the corresponding text to avoid the confusion.

      (3) The title and framing assert that "nuclear export governs TDP-43 phase transitions." However, prior studies such as Pinarbasi et al. 2018 and Duan et al. 2022 indicate that TDP-43 is not a canonical XPO1 cargo and that its export is largely passive, with active nuclear import being the dominant determinant of nuclear localization. The authors cite these studies but still position XPO1 as a central, quasi-direct regulator. The data presented are largely correlative or based on pharmacologic manipulation and overexpression in an overexpression mutant background, with no direct evidence that XPO1 engages TDP-43 in a specific, regulated manner. Even if XPO1 does not engage WT TDP-43, it could still engage the 2KQ variant, which needs to be tested.

      We did not mean to conclude or imply that the regulation of TDP-43 by XPO1 is direct. In fact, we explicatively mentioned on page 8 of the original manuscript that the regulation is likely indirect and mediated by other factors. The sentence reads as “Since XPO1 does not bind TDP-43 directly (Pinarbasi et al., 2018), additional factors might link XPO1-mediated nuclear export to TDP-43 nuclear egression.”

      We now add new data in Figure 6, showing that in an in vitro reconstitution assay using semi-permeabilized cells, LMB treatment significantly stabilizes anisosomes in an RNA dependent manner. This new data suggests that XPO1 inhibition leads to increased nuclear RNA availability, which indirectly favors anisosome assembly and maturation (see discussion). We believe that this new finding has provided significant new insight into how nuclear transport modulates TDP-43 phase behavior. We have revised the title, the abstract and changed the framing according to the reviewer’s suggestion.

      (4) The XPO1 perturbations yield somewhat confusing phenotypes. XPO1 inhibition using Leptomycin B, KPT-276, and Verdinexor reduces anisosome number and enlarges remaining anisosomes, which remain liquid-like by FRAP recovery and fusion assays and stay nuclear. XPO1 overexpression causes fewer, enlarged puncta, but these are FRAP-impaired (gel-like) and redistribute to the cytoplasm. Thus, both decreased and increased XPO1 activity reduce anisosome number and enlarge puncta, but with opposite phase behaviors and subcellular localizations. The model presented in Figure 5L is relatively qualitative and does not resolve these issues. Moreover, XPO1 inhibition globally impairs nuclear export of many cargos and profoundly alters the nuclear environment, transcription, RNA processing, and chromatin. It is therefore difficult to conclude that the observed effects are specific to TDP-43 phase regulation as opposed to secondary consequences of broad nuclear export blockade.

      The reviewer correctly summarizes our data and interpretation: XPO1 loss-of-function and gain-of-function generate opposite phenotypes regarding TDP-43 phase regulation.

      Regarding the mechanism underlying XPO1-dependent TDP-43 phase regulation, as mentioned above, we developed a semi-permeabilized cell-based assay in which we used the pore-forming toxin streptolysin O to damage the plasma membrane after anisosome induction. We noticed that upon cell permeabilization and cytosol loss, anisosomes were mostly lost (Figure 6B, C). This is probably due to a reversible partition of TDP-43 into a less fluorescent soluble fraction. Supporting this idea, when permeabilized cells were incubated with cytosol plus an energy regenerating system, small puncta containing TDP-43 2KQ could be reformed in an energy dependent manner (Figure 6D, E). Interestingly, in LMB-treated cells, anisosomes remained stable despite cell permeabilization(Figure 3F). Since LMB treatment did not increase TDP-43 nuclear concentration (Supplemental Figure 1), this data suggest that nuclear export inhibition likely alter the nuclear environment to stabilize anisosomes. Indeed, when cells were permeabilized in the presence of a small RNAase, LMB-stabilized anisosomes also collapsed (Figure 6G).

      We now add more discussions on the potential effect of RNA on TDP-43 phase behavior in XPO-1 inhibited cells considering these new findings.

      (5) The authors show that anisosome induction depletes nucleoplasmic XPO1 signal and that mCherry-XPO1 can be seen in some TDP-43 puncta. However, antibody penetration into anisosomes is limited, so XPO1 depletion from nucleoplasm could reflect sequestration in the anisosome shell or core, but this is not demonstrated. There is no demonstration of physical interaction, even indirect interaction, between XPO1 and TDP-43 or a defined adaptor, nor identification of a specific mutant of XPO1 that selectively disrupts this putative interaction while preserving other functions. The known TDP-43 NES has been shown to be weak and not a functional XPO1-dependent NES in multiple studies. If XPO1 is acting through an adaptor that recognizes 2KQ or K181E specifically, that by itself would bring into question the generality of the mechanism for wild-type TDP-43.

      We agree that our data does not demonstrate an interaction between XPO1 and TDP-43. Considering our new data (mentioned above), it is possible that the effect of anisosome induction on endogenous XPO1 localization is also mediated by RNA. We now mention more explicitly that the regulation of TDP-43 by XPO1 is likely indirect (Page 8). We have revised our paper to separate any speculative statements from the data, and also discussed the possibility of alternative interpretations.

      (6) To support a mechanistic claim that nuclear export governs TDP-43 phase transitions, more targeted evidence is needed. The authors should test whether siRNA knockdown or CRISPR interference of XPO1 in the DLD1 2KQ model reproduces the effects seen with Leptomycin B and KPT-276, including FRAP and fusion phenotypes, and verify on-target effects by rescue with an siRNA-resistant XPO1 construct. They should demonstrate that canonical XPO1 cargos behave as expected under the inhibitor conditions used, as a positive control, and that the concentrations used are not grossly toxic. They should attempt to identify or at least constrain candidate adaptors that might enable XPO1-dependent export of TDP-43 through proteomic analysis of XPO1 co-purifying with 2KQ condensates or loss-of-function studies of candidate adaptors from the siRNA screen. Finally, they should test whether a TDP-43 mutant that cannot bind the proposed adaptor still responds to XPO1 manipulation.

      The anisosome enlargement phenotype upon XPO1 depletion was seen in our siRNA screens, which was identified by machine-based image analyses using 6 different siRNAs. This, together with the chemical inhibition experiments, demonstrate that the phenotype is specifically caused by XPO1 inactivation.

      When characterizing the effect of XPO1 inhibition on anisosome dynamics, we preferred chemical inhibitor because the effect is acute, and therefore less likely to be secondary.

      Regarding the inhibitor concentration, according to the literature, Leptomycin B was commonly used at 50-200 nM. We chose 200 nM to ensure a quick and complete inhibition of XPO1-mediated nuclear export (see Figure 3 in PMID: 9628873). This dose is also well tolerated by our cells.

      We did not suggest any specific adaptor that mediates XPO1 interaction with TDP-43. Whether there is an adaptor, and if so, the identity of such adaptor is out of the scope of this study. We revise our paper on page 8-9 to clarify these points.

      (7) Even with these data, what is currently shown is that global modulation of nuclear export capacity can alter the phase behavior and localization of a highly overexpressed RNA-binding-defective TDP-43 mutant and of K181E in organoids. This is important, but it is weaker than asserting that XPO1 directly governs TDP-43 phase transitions in physiological contexts. The title, abstract, and Discussion should be tempered to reflect that nuclear export is one of several pathways, alongside RNA splicing, translation, and proteostasis, that influence TDP-43 phase states in this model, and that the specific mechanism and cargo relationship between XPO1 and TDP-43 remain unresolved and may be indirect.

      We have revised the title, abstract, and main text to temper our conclusions.

      (8) The authors conclude that cytoplasmic TDP-43 aggregation plays only a modest role in TDP-43 proteinopathies because in homozygous K181E organoids, chronic KPT-276 treatment almost abolishes cytoplasmic pTDP-43 puncta, yet bulk RNA-seq shows only a relatively small fraction of dysregulated genes are rescued. There are several issues with this inference. Relying primarily on pTDP-43 antibody staining to define cytoplasmic TDP-43 aggregation is limiting. pTDP-43 antibodies label only phosphorylated species and may miss non-phosphorylated, oligomeric, or amorphous TDP-43 species that could still be toxic. Different pTDP-43 antibodies vary in epitope accessibility depending on aggregate conformation and subcellular location. More sensitive approaches, such as high-affinity TDP-43 RNA aptamer probes developed by Gregory and colleagues, biochemical fractionation for SDS-insoluble and urea-soluble TDP-43, and filter-trap assays, would provide a more quantitative assessment of cytoplasmic aggregation and its reduction by KPT-276. Without these, it is not safe to assume that cytoplasmic aggregation has been eliminated, as opposed to one antigenic subclass.

      We agree with the reviewer that p-TDP-43 may not represent all aggregate species. However, p-TDP-43 antibodies detect the pathologically validated species tightly associated with TDP-43 proteinopatheis. In human ALS and FTD-TDP tissues, cytoplasmic inclusions are strongly immunoreactive for phosphorylated TDP-43 (typically S409/410, as detected here). Additionally, p-TDP-43 immunohistochemistry is a routine diagnostic criterion in neuropathology. For these reasons, we believe that the observation that inhibition of XPO1 significantly reduces p-TDP-43 is a significant finding, as it suggests that inhibition of nuclear transport may rescue TDP-43 proteinopathy. We revised the text on page 9 to better explain the significance of p-TDP-43 staining.

      (9) The treatment window, spanning from day 87 to 122 with 20 nanomolar KPT-276, may be too late or too mild to reverse entrenched nuclear RNA-processing defects, even if cytoplasmic inclusions are cleared. Once widespread cryptic exon inclusion and alternative polyadenylation misregulation are established, many downstream changes may become self-sustaining or only partially reversible. Moreover, XPO1 inhibition will massively rewire nucleocytoplasmic transport of many transcription factors, splicing factors, and RNA-binding proteins. Thus, the lack of full transcriptomic rescue cannot be cleanly interpreted as evidence that cytoplasmic aggregates are only modest contributors. It may instead reflect that nuclear dysfunction is primary and XPO1 inhibition does not correct, and may even exacerbate, certain nuclear defects.

      We agree with the reviewer that the lack of rescue may be caused by some technical issues. We have removed the RNAseq data and the related texts since it is not essential.

      (10) To support a causal statement about the modest contribution of cytoplasmic aggregates, one would want more direct measures of neuronal health and function, such as cell death, neurite complexity, synaptic markers, and electrophysiology before and after KPT-276, not only transcriptomics. A way to selectively reduce cytoplasmic aggregation without globally inhibiting nuclear export would allow comparison of outcomes.

      We have removed the discussion regarding the role of cytoplasmic aggregates in disease.

      (11) Given these caveats, the concluding statements that cytoplasmic TDP-43 aggregation is only a modest contributor should be substantially softened. A more defensible interpretation is that in this homozygous K181E organoid model, chronic global XPO1 inhibition reduces pTDP-43-positive cytoplasmic puncta but only partially normalizes the steady-state transcriptome, suggesting that persistent nuclear RNA-processing defects and other pathways continue to drive pathology.

      We agree with the review and have removed the RNAseq part.

      (12) The screens are a major strength but need more rigorous validation for key hits, especially nuclear transport factors. For the siRNA screen, hits are filtered by anisosome number per nucleus, but there is no direct demonstration in the main text that XPO1 or CSE1L knockdown is efficient at the messenger RNA or protein level. For the highlighted genes, Western blot or quantitative polymerase chain reaction validation and phenotypic rescue would strengthen confidence. For small-molecule hits, it is not systematically shown that anisosome modulation is independent of changes in total TDP-43 2KQ expression or gross toxicity. Translation inhibitors are tested for this, but for many other hits, including proteasome, HSP90, and kinase inhibitors, expression and general nuclear structure should be monitored. Given the reliance on anisosome count as a readout, secondary screens that specifically distinguish changes in TDP-43 expression levels, changes in nuclear morphology or cell cycle, and specific changes in anisosome phase behavior, including FRAP and fusion for top hits, would greatly increase interpretability.

      For the siRNA screen, each positive hit was confirmed by two rounds of screen with 6 independent siRNAs in total. Although we did not validate the knockdown efficiency due to the large number of hits, we routinely include a positive siRNA control in our study (Cell death siRNA), which targets several essential gene. Transfection efficiency was controlled by measuring cell viability after knocking down of these genes. In addition, the identification of XPO1 as a positive regulator of TDP-43 phase behavior was independently validated by our chemical genetic screens with three XPO-1 inhibitors. We feel confident that XPO1 is a key modulator of TDP-43 phase behavior.

      For chemical treatment experiments, the anisosome fusion phenotypes could be detected as early as 5 h post treatment. Given the relatively short treatment, we do not expect a significant change in protein level or toxicity. To alleviate this reviewer’s concern, we performed an immunoblotting experiment to measure the total TDP-43 protein levels in drug-treated cells. Except for VLX, we did not detect any significant changes in the level of TDP-43 after drug treatment (Supplemental Figure 1).

      (13) The classification of condensates as liquid versus gel-like or solid is based almost entirely on FRAP recovery or lack thereof. While FRAP is appropriate, interpretations could be made more robust by including half-region-of-interest bleach controls and assessing mobile fractions and recovery kinetics more quantitatively across conditions. Complementing FRAP with other phase-behavior assays such as sensitivity to 1,6-hexanediol, shape relaxation after deformation, and coarsening behavior over longer timescales would strengthen the analysis. At present, some assignments, such as that XPO1 overexpression drives a gel-like transition, are reasonable but somewhat qualitative.

      In this study, we used two types of FRAP assays. We either bleached TDP-43 within anisosomes or bleached the surrounding TDP-43 molecules(Figure 2). The two complementary methods yield consistent results that allow unambiguously distinguish between TDP-43 LLPS state and gel-like condensation.

      In XPO1-related experiments, the two types of condensates formed by TDP-43 2KQ can be distinguished by several features including their subcellular localization, shape, and the fluorescence recovery kinetics. We feel that these combined data clearly segregate these puncta into two distinct types of assemblies. The proposed half-region-of-interest bleach is technically challenging for small anisosomes under normal conditions. However, whenever possible, (e.g. anisosomes enlarged by Leptomycin B), we did perform both whole anisosome bleach and partial bleach (Figure 5D, I). Both assays demonstrate that TDP-43 in these enlarged anisosomes is highly mobile.

      (14) For the Leptomycin B and KPT-276 experiments in cells and organoids, it would be important to confirm that canonical XPO1 cargo proteins accumulate in the nucleus and that the concentrations used are within a range that is not overtly toxic over the experimental timeframe. Assessing nuclear morphology, chromatin condensation, and general transcriptional activity through global RNA synthesis or key reporter genes would ensure that observed effects are not secondary to severe global nuclear export collapse.

      In Leptomycin B treatment experiments, we carefully chose a dose that was previously validated (see Figure 3 in PMID: 9628873). Based on our DAPI staining, the nuclear morphology appears normal with no abnormal chromosome condensation (Figure 5A). Additionally, in cell line-based experiments, the effect of Leptomycin B on anisosomes was detected 6-8 hours post treatment. The change in global protein synthesis because of RNA changes should be relatively minor at this stage. Indeed, our new immunoblotting experiment showed that LMB treatment did not affect TDP-43 protein level (Supplemental Figure 1). Most importantly, the in vitro semi-permeabilized assay demonstrates a direct role for RNA in stabilizing anisosomes.

      (15) In the organoid section, it is not clear how many independent iPSC clones and organoid batches were used per condition, nor whether batch effects were assessed in the bulk RNA-seq analysis. This should be fully specified and ideally controlled with isogenic wild-type and K181E clones. For transcriptional rescue, it is important to know whether the changes in wild-type organoids treated with KPT-276 are negligible. A direct wild-type comparison with or without KPT-276 is important to disentangle general drug effects from K181E-specific rescue. More detailed quantification of total TDP-43 and pTDP-43 in both nuclear and cytoplasmic fractions, including biochemical fractionation if possible, would strengthen the assertion that KPT-276 specifically reduces cytosolic pTDP-43 aggregates while sparing nuclear TDP-43.

      The organoid experiment was performed with two batches per condition to reduce the effect of batch variation. The wildtype cells and K181E mutant are derived from the same genetic background. This information is now included in the method section on page 14. Given the criticisms by review 1 and 2 on the RNAseq data, we have removed this non-essential data. 

      (16) Beyond the core issues above, several additions could greatly enhance the impact. The manuscript currently emphasizes XPO1, but the genetic and chemical data clearly implicate RNA splicing, translation, and proteostasis as equally strong or stronger regulators of TDP-43 phase states. A more integrated model that explains how these pathways intersect, for example, how splicing factor availability, ribosome loading, and proteasome capacity co-govern anisosome nucleation, growth, and hardening, would be valuable.

      We now discuss a new model in discussion based on our new Figure 6, which integrates the role of RNA splicing and nuclear transport in TDP-43 phase regulation on page 10. We agree with the reviewer that other questions are also important for future studies.

      (17) A key unresolved question is whether XPO1 is acting directly on TDP-43, or instead primarily regulates anisosomes by exporting other factors that more proximally control TDP-43 phase behavior. Given that TDP-43 is not a canonical XPO1 cargo and prior work indicates that its nuclear export is largely passive, it seems at least as plausible that XPO1 inhibition alters the nuclear concentration or localization of splicing factors, RNA-binding proteins, chaperones, or other modifiers identified in the screens, and that changes in these proteins secondarily reshape anisosome dynamics. In other words, XPO1 may be exporting a more direct regulator of anisome formation and hardening, rather than exporting TDP-43 itself in a specific, regulated way. The current data do not distinguish between these possibilities. Systematic identification of XPO1-dependent cargos that colocalize with or biochemically associate with anisosomes, combined with targeted perturbation of their nuclear export, would be needed to determine whether the relevant XPO1 substrate in this system is actually TDP-43 or an upstream modulator of its phase behavior.

      As discussed above, our new data regarding the role of RNA in TDP-43 phase regulation should alleviate this concern, although we cannot exclude the possible involvement of splicing factors in this process. We also clearly state that there is no evidence to support a direct interaction between TDP-43 and XPO1 on page 8.

      (18) Testing whether identified modifiers converge on nuclear TDP-43 concentration would be informative. Since phase separation is concentration-dependent, measuring nuclear versus cytoplasmic TDP-43 levels across key perturbations, including splicing inhibition, translation inhibition, proteasome inhibition, HSP90 inhibition, and XPO1 modulation, would help determine whether modifiers mainly work by changing nuclear TDP-43 concentration or by altering interaction networks and the material properties of condensates.

      In the newly performed immunoblotting experiment, we measured the TDP-43 levels in drug-treated cells but found no effect by most drugs (Supplemental Figure 1).

      (19) Examining other ALS-relevant RNA-binding proteins would be valuable. Given the role of XPO1 and other hits, it would be informative to briefly test whether similar principles apply to FUS, hnRNPA1, or other ALS-relevant RNA-binding proteins in the same cellular context, to argue for generality versus TDP-43-specific idiosyncrasies of the 2KQ system.

      We agree that this is an important issue but we feel the proposed experiments are beyond the scope of the study.

      (20) The Introduction sometimes implies that anisosomes are common and well-established intermediates en route to pathology. It would be helpful to more clearly state that, to date, anisosomes are primarily observed in overexpression and mutant systems and have not yet been unequivocally demonstrated in human patient tissue. The link between PDGFRβ, PAK4, GSK-3β, and YAP and TDP-43 phase dynamics is intriguing but only briefly mentioned. The authors should either expand on this or tone down the emphasis in the Results section.

      We have revised the introduction and added the following sentence on page 4. “The 2KQ-containing anisosomes, observed mostly in the nucleus under overexpression conditions, have not been validated in human patient samples.”

      (21) In the organoid methods, the authors should consider clarifying whether doxycycline is continuously used, which might alter TDP-43 expression and nuclear transport in a non-negligible way.

      The organoid model does not involve protein overexpression or doxycycline treatment. We measured endogenous p-TDP-43, which is why we feel this experiment is very significant. Unlike many other p-TDP-43 detection studies that rely on TDP-43 overexpression or exposing cells to excess stressors, we could detect substantial p-TDP-43 in 3D organoids grown under normal conditions, whereas the same cells grown and differentiated in 2D culture do not show p-TDP-43 (Zhang Q. et al., BioRxiv 2025).

      (22) For statistical methods, it would be beneficial to indicate whether multiple-comparison corrections were applied for the many FRAP, anisosome count, and size comparisons beyond DESeq2 internal corrections for RNA-seq.

      We have added more statistical information to the figure legends.

      (23) Some figure legends could more clearly indicate whether the images shown are single z-planes or maximum intensity projections and how the thresholding for anisosome detection was performed.

      We revised the figure legends to include this information. As for anisosome detection, because they are so obvious, standard thresholding combined with automated counting was sufficient to identify them.

      (24) In its current form, the manuscript contains an impressive set of screens and some nicely executed imaging of TDP-43 condensates, highlighting nuclear export among other pathways as a modulator of TDP-43 phase behavior. However, the physiological relevance is undercut by heavy reliance on an acetylation-mimetic, RNA-binding-defective TDP-43 mutant and a homozygous K181E organoid model. The mechanistic link between XPO1 and TDP-43 remains largely inferential and partly at odds with prior work. The conclusion that cytoplasmic TDP-43 aggregation is only a modest contributor to disease is not firmly supported by the available data.

      We agree with the reviewer that the strength of the study is our unbiased approach that identifies pathways capable of modulating TDP-43 phase behavior. In the revised paper, we included several experiments using an in vitro semi-permeabilized cell system to further dissect the role of nuclear export in TDP-43 phase separation. We believe that these new results should provide significant mechanistic insight that links nuclear export and RNA transcription and splicing to TDP-43 phase regulation. Additionally, we have revised our paper carefully to discuss the physiological relevance and the limitation of our study.

      (25) With substantial additional mechanistic work, particularly around XPO1, rigorous validation in more physiological TDP-43 contexts, more sensitive detection of cytoplasmic TDP-43 aggregates, and a tempering of the central claims, this study could make a meaningful contribution to understanding how nucleocytoplasmic transport and other cellular pathways influence TDP-43 phase transitions and aggregation. The work should be reframed as an important screening study that identifies nuclear export as one among several cellular processes that modulate TDP-43 phase behavior in a model system, rather than as a definitive demonstration that nuclear export governs pathological TDP-43 aggregation in disease.

      We now reframe the study as an important screening study that identifies nuclear export among several other pathways as modulators of TDP-43 phase behavior. We also propose a model that links RNA splicing to nuclear export in TDP-43 phase regulation.

      Reviewer #2 (Public review):

      Summary:

      This manuscript addresses an important and timely question in TDP-43 biology by systematically identifying regulators of TDP-43 anisosome formation, with a particular focus on nuclear export via XPO1. Using a combination of unbiased chemical screening, genetic perturbation, and advanced imaging approaches, the authors propose that inhibition of nuclear export modulates the abundance and biophysical properties of TDP-43 anisosomes. The study is conceptually innovative and has potential relevance for neurodegenerative diseases characterized by TDP-43 pathology. However, significant concerns regarding experimental controls, reporting transparency, and model translatability currently limit the strength of the conclusions and the interpretability of several key findings.

      We thank the reviewer for acknowledging the significance and innovation of our study.

      Strengths:

      (1) The study employs an unbiased, hypothesis-free compound screen to identify regulators of TDP-43 anisosome formation, which is a major strength and reduces confirmation bias.

      (2) The authors combine chemical and genetic screening approaches, providing orthogonal validation of key pathways and increasing confidence in the biological relevance of top hits.

      (3) The focus on biophysical properties of TDP-43 assemblies, assessed through imaging and FRAP, moves beyond simple presence/absence of aggregates and provides mechanistic insight into the biophysical states of TDP-43.

      (4) The use of multiple experimental modalities, including live-cell imaging, FRAP, pharmacological perturbation, and transcriptomic analysis, reflects a technically sophisticated and ambitious study design.

      (5) The authors attempt to extend findings beyond immortalized cancer cell lines by incorporating organoid models, demonstrating awareness of disease relevance and translational importance.

      Overall, the manuscript is clearly written and logically structured, making complex experimental workflows accessible and the central hypotheses easy to follow.

      Weaknesses:

      Despite its strengths, the manuscript has several major limitations that affect data interpretation and confidence in the conclusions.

      (1) Lack of appropriate controls for overexpression experiments:

      A central concern is the absence of proper controls for TDP-43 and XPO1 overexpression. Prior studies (including those cited by the authors, Archbold et al.2018) show that overexpression of WT TDP-43 alone is toxic to neurons. Thus, the experimental system itself may induce anisosome formation independently of the mechanisms under study. Similarly, XPO1 overexpression lacks a suitable control (e.g., mCherry alone or mCherry fused to a protein known to be independent of TDP-43). The near-complete colocalization of XPO1 with TDP-43 anisosomes upon overexpression raises the possibility that these structures reflect non-physiological protein accumulation rather than regulated assemblies.

      As mentioned in our response to reviewer 1, point 1, we have added more discussions to justify the use of acetylation mimetics in our study. We agree with the reviewer that these large puncta (both anisosomes and gel-like structures) likely resulted from TDP-43 overexpression. Nevertheless, in a titration experiment done by Yu et al. 2020 (PMID: 33335017), they showed that ectopic TDP-43 undergo demixing even at concentrations lower than endogenous TDP-43, although the demixed puncta were very small. Their result suggested that overexpression per se does not change TDP-43 phase behavior, only enlarge the demixed TDP-43 structures, which is necessary for our screen and imaging-based characterization.

      For XPO1 overexpression, we have done the mCherry alone control but due to space limit in Figure 5, we did not include it. We now include the data in Supplemental Figure 4. This figure shows that overexpression of mCherry did not change TDP-43 localization or anisosome structures.

      (2) Insufficient experimental and analytical transparency:

      The manuscript frequently lacks clear reporting of experimental details. In multiple figures, the stated number of independent experiments does not match the number of data points shown, making it difficult to assess statistical validity. Concentrations used in the compound screen are not clearly defined, nor is it stated whether multiple concentrations were tested. It is unclear how many wells, cells, or independent cultures were analyzed. The criteria used to reduce 1,533 screening hits to 211 candidates via STRING analysis are not explained. Knockdown and overexpression efficiencies are not reported.

      We apologize for these omissions. We have added more experimental details to the figure legends and the method. For the imaging experiments, data points reflect randomly selected individual cells imaged in 2-3 independent biological repeats. This is now stated in the figure legends. For chemical screens, we screened against NCATS libraries was first done at top concentration (10 mM) to ensure inhibitory efficacy for all potential hits. In the follow-up validation study, we validated the top hits using a series of concentrations, as shown in Figure 1B. Drug concentrations are provided in Figure 2A, 4A, C, E, F, 5A-D, F, Figure 6F, G, Figure 7A)

      We explain the STRING analysis in more detail now. Basically, STRING is a protein-protein interaction network that reports all potential interactions between any proteins in human proteome. Given the potential off-target effect of siRNA, we assume that if the screen identifies multiple components of a protein interaction network or pathway, the result is more likely to be real.

      We did not check XPO1 knockdown efficiency in high through-put screens (HTS) for several reasons. Firstly, the large number of positive hits makes it impossible to check knockdown efficiency for all of them. Secondly, the effect of XPO1 knockdown on anisosomes was seen with 6 different siRNAs in two rounds of screens. Thirdly, in the HTS protocol, we routinely included a transfection control (siRNAdeath) to control transfection efficiency. We would only process the data if siRNAdeath control killed > 90% of the cells. Lastly, the XPO1 knockdown result was independently validated by small molecule inhibitors. For TDP-43 overexpression, the study by Yu and colleagues suggested that the expression is more than 20-fold higher than endogenous TDP-43, but they showed that anisosome formation is not an artifact of protein overexpression. When the expression level was titrated down, they could still detect anisosomes.

      (3) RNA-seq concerns:

      The RNA-seq experiments are particularly problematic. The number of biological replicates per condition is not stated, and heatmaps suggest that only one sample per group may have been used, which would preclude statistical analysis. No baseline comparison between WT and mutant TDP-43 is shown. Given that TDP-43 is an RNA-binding protein, splicing analyses would be far more informative than gene expression alone, yet no splicing data are presented. Moreover, nuclear retention of TDP-43 does not preclude nuclear aggregation, which may still impair its splicing function.

      We apologize for the lack of clarity regarding the RNA-seq design. For each condition, organoids of two independently differentiated batches were treated in triplicate. What we showed before was averaged expression levels. We pooled the organoids of the same treatment from the two batches to reduce the impact of batch variation.

      Given the criticisms from both reviewers 1 and 2 on the limited interpretation power of the RNAseq study, we have removed this data from the revised manuscript.

      (4) Limited translatability to neuronal biology:

      All anisosome analyses are performed in a cancer cell line, raising concerns about relevance to post-mitotic neurons. While organoids are used as a secondary model, the assays performed do not overlap with those used in cancer cells, making it difficult to assess whether anisosome-related mechanisms are conserved. Neuronal toxicity, a critical outcome given known TDP-43 biology, is not assessed. Prior work has shown that WT TDP-43 overexpression alone is toxic to neurons, yet this is not addressed.

      We agree with the reviewer that the model used in this study is not directly relevant to neurodegeneration. However, as pointed out by the reviewer, neurons are much more sensitive to TDP-43-associated toxicity. By contrast, the cell line used in this study can tolerate TDP-43 overexpression with no detectable cytotoxicity. This feature makes it feasible to evaluate how different cellular processes modulate TDP-43 phase behavior without the confounding effect from cytotoxicity. Notably, the processes identified by our screens are all house-keeping pathways that are conserved in neurons. Thus, we believe that the reported findings are likely applicable to neurons. That being said, we have revised our paper to ensure that we don’t overstate the clinical relevance of our work.

      (5) Conceptual and interpretational gaps:

      The authors quantify anisosome number but also report conditions in which anisosome number decreases while size increases. The biological interpretation of larger anisosomes is not discussed, and whether this reflects improvement or worsening of pathology is unclear. Compounds targeting the same mechanism (e.g., nuclear export inhibition) are inconsistently used across experiments (KPT compounds, verdinexor, leptomycin B), raising concerns about reproducibility. In organoids, the experimental paradigm shifts to long-term treatment (35 days vs. 16 hours), further complicating interpretation.

      We thank the reviewer for these critical points. As pointed out by the reviewer 1 in point 4 above, we do not have evidence to establish a convincing correlation between the size of anisosomes and clinical phenotypes. Regarding the use of different drugs for different experiments, the initial screen identified KPT and Verdinexor because they are investigational drugs, but Leptomycin B was not in our library. In the follow-up studies, we switched to Leptomycin B because 1) it is highly potent and specific; 2) it was better characterized and more commonly used as inhibitors of XPO1 according to the literature. However, for the organoid study, we had to switch back to KPT because of the toxicity issue associated with long-term application of Leptomycin B.

      (6) Overinterpretation of rescue effects:

      Although the authors state that they aim to test whether nuclear export inhibition rescues neuronal defects, no functional neuronal readouts are provided (e.g., viability, morphology, axon outgrowth, or electrophysiological measures). RNA-seq alone is insufficient to support claims of rescue.

      Our interpretation of the RNA-seq data was that the rescue effect by nuclear export inhibition was limited and probably insignificant. Given that this negative data is not conclusive, we have removed it from the revised manuscript.

      (7) Finally, the model does not appear to exhibit cytosolic TDP-43 aggregation at baseline. It remains unclear whether longer induction would produce cytosolic gel-like assemblies and whether these would be prevented by nuclear export inhibition. Long-term data are shown only in organoids, yet anisosome formation is not assessed there.

      The expression system used in the study reaches a steady state after 24 h of induction. Prolonged expression up to 48 h did not alter the number of anisosome, nor does it change TDP-43 phase behavior. We now clarify this point on page 4.

      Reviewer #3 (Public review):

      Summary:

      TDP-43 proteinopathy is broadly found in neurodegenerative diseases. This manuscript investigates how nuclear export influences the biophysical properties of TDP-43. The authors use a combination of chemical screening and genome-wide siRNA screening to identify pathways that modulate TDP-43 liquid-to-solid transitions. Overall, the study employs a broad array of approaches and addresses an important question in TDP-43 pathobiology. The identification of nuclear export as a central regulator is compelling and conceptually aligns with the emerging view that TDP-43 nucleocytoplasmic trafficking is a major defect in neurodegeneration.

      Strengths:

      This work integrates chemical and genetic screening to identify novel modifiers. The candidates were validated in both reporter cell lines and iPS-differentiated organoids. The findings support the nucleocytoplasmic transport is important for the biophysical properties of TDP-43.

      We thank the reviewer for acknowledging the significance and strength of our study.

      Weaknesses:

      The mechanisms underlying the connection between nuclear export and phase transition need further clarification. Broader consequences of XPO1 inhibition are not addressed.

      We agree that our previous manuscript did not address how nuclear export inhibition affect TDP-43 phase behavior. As discussed in our paper, we proposed that the effect of nuclear export inhibition on TDP-43 phase separation is likely indirect. The most likely scenario is that inhibition of nuclear export changes the nuclear environment over time, which affects TDP-43 phase separation. We have tried to isolate nuclear extracts from control and LMB-treated cells and used mass spectrometry to identify proteins that are differentially present in the nucleus. However, knockdown of the identified top candidates did not abolish LMB-induced phase alteration (not shown). Considering our observation that RNA splicing is another modulator of TDP-43 phase behavior, we reasoned that it is possible that it is the combined change of RNA and protein composition in the nucleus that alters TDP-43 phase behavior. In new experiments presented in Figure 6, we now used a semi-permeabilized in vitro system to demonstrate that LMB treatment stabilized anisosomes in an RNA-dependent manner (see response to point 4 by reviewer 1). This new data allows us to propose a new model that link RNA splicing and nuclear export in TDP-43 phase regulation (Discussion).

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Include appropriate controls for all overexpression experiments. In particular, overexpression of WT TDP-43 alone and suitable tag-only controls (e.g., mCherry alone or mCherry fused to a protein unrelated to TDP-43/XPO1) should be included to control for aggregation driven by non-physiological protein levels.

      In Supplemental Figure S4, we included a tag-only control, which shows that mCherry alone does not affect the localization of XPO1, neither did we see mCherry co-localizes with TDP-43.

      Since WT TDP-43 itself does not form anisosome and because the goal of the study was to test how anisosome dynamics is affected by various conditions, we did not repeat our experiments with WT TDP-43.

      (2) Address whether TDP-43 anisosomes form under endogenous or near-physiological expression levels. If possible, include experiments using lower expression systems or endogenous tagging to demonstrate that anisosome formation is not solely an overexpression artifact.

      As mentioned above, in a titration experiment done by Yu et al. 2020 (PMID: 33335017), they showed that ectopic TDP-43 undergoes demixing even at concentrations lower than endogenous TDP-43, although the demixed puncta are small. Their result suggested that overexpression per se does not change TDP-43 phase behavior. Instead, it only enlarges the demixed TDP-43 structures, which is necessary for our screen and imaging-based characterization.

      (3) Clearly define biological versus technical replicates throughout the manuscript and report exact n-numbers for all experiments in figure legends and/or methods. Resolve discrepancies between stated and displayed n-numbers (e.g., figures showing more data points than the number of independent experiments reported). Further, include how data points were defined (e.g., cells, fields of view, wells).

      We now state clearly the biological repeats in figure legends. We did not use N number to specify technical replicate. The discrepancy between the stated N number (biological repeats) and the data points is because for imaging experiments, data points usually represent single cells collected from 2-3 biological replicates (N=2 or 3). Data points are now clearly defined in the figure legends (anisosome, cell, imaging field, or independent experiment).

      (4) The authors state that they identified a list of compounds that reduced anisosomes. Please clarify how the threshold was determined: Was this a statistical analysis or a specific threshold that has been used?

      For both siRNA screen and chemical genetic screen, we calculated the Z-score and used Z-score>2 as a cutoff. This is mentioned in the method.

      (5) Provide a complete list of compounds used in the chemical screen, including concentrations tested and whether multiple doses were evaluated.

      As mentioned above, the initial screen was done with just one concentration (10 mM). Identified positive hits were re-tested with multiple doses as shown in Figure 1. The compounds are from a commercial library (LOPAC R1280, Sigma #LO4200). The list of compounds can be found at vender’s website.

      (6) Clearly explain the criteria used to reduce the initial 1,533 screening hits to 211 candidates following STRING analysis, including cutoffs and prioritization logic.

      We now explain that the Z-score was used to further narrow down the hit (page 6). Additionally, we provide an explanation on how we use STRING to further narrow down the list. The sentence reads as “To further narrow down the list, we performed a STRING protein network analysis based on the assumption that a protein interaction network bearing multiple positive hits would be more likely to be a true effector.”

      (7) Report knockdown and overexpression efficiencies for all genetic perturbations used in the study.

      For TDP-43 overexpression, the study by Yu and colleagues suggested that the stable cell line expresses 20-fold more TDP-43 than endogenous one, but they showed that anisosome formation is not an artifact of protein overexpression. When the expression level was titrated down, they could still detect anisosomes (Yu, H. et al., Science 2021). For knockdown efficiency, since the screen used 6 different siRNAs for each identified target (a few hundred), it is technically challenging to validate the knockdown efficiency of each siRNA by conventional qRT-PCR. To control knockdown efficiency, we transfected cells in parallel with siRNA-death that contains a mixture of siRNAs targeting several essential genes (Qiangen, #1027299). We would only process the data if siRNAdeath control killed > 90% of the cells, indicating good knockdown efficiency.

      (8) Clarify the biological interpretation of changes in anisosome size versus number, particularly in conditions where fewer but larger anisosomes are observed. Discuss whether larger assemblies are hypothesized to be protective, neutral, or deleterious.

      Live cell imaging was used to dissect why cells treated with certain drugs such as XPO1 inhibitors have fewer but larger anisosome. Figure 5F shows that this is caused by the fusion of small anisosomes. Our data does not suggest that the size of anisosomes can differentiate between protective or deleterious state, but rather it is the LLPS state and subcellular localization of these assemblies that may play a more critical role in determining whether TDP-43 forms deleterious protein aggregates. The discussion is on page 10.

      (9) Specify whether all anisosomes induced by XPO1 overexpression were gel-like or whether this applied only to a subset. If only a subset was affected, please provide quantifications, otherwise state clearly that all anisosomes in XPO1 overexpression were gel-like.

      All TDP-43 puncta mislocalized to the cytoplasm in XPO1-overexpressing cells are gel-like because the FRAP experiment in Figure 5I was done with randomly selected TDP-43 puncta mislocalized to the cytoplasm.

      (10) Clarify which anisosomes (nuclear vs cytosolic; gel-like vs non-gel-like) were selected for FRAP analyses in Figure 5I.

      For Figure 5I, the control anisosomes in untreated cells are nuclear while under mCh-XPO1 expressing condition, only those in the cytoplasm were randomly selected for photobleaching.

      (11) The translatability of the conclusion based on cancer cell lines to brain organoids is not convincingly shown and could be strengthened by including additional assessment of anisosomes. While this might not be feasible in 3D cultures, the authors could alternatively use 2D cultured neurons to perform the same assays as performed in the cancer cell line. Additionally, the same treatment strategy should be applied. The reasoning for increasing treatment to 35 days in the organoids is unclear.

      In another manuscript that is currently under revision, we compared 2D iNeuron culture with 3D organoids. A pre-print is available at https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full. In this study, we found that endogenous TDP-43 K181E mutant do not undergo phosphorylation-dependent transition to aggregate in 2D cultures. Only when these cells were grown into 3-D organoids, TDP-43 phosphorylation could be detected. (see supplemental Fig. S1c, d in https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full). Thus, it is not possible to repeat the experiments in this study in 2D iNeuron cultures. We agree with the review that there is a gap between the study using the cancer cell line and the use of K181E iPSC-derived 3D organoids. We have toned down our conclusions throughout the text.

      (12) Address neuronal vulnerability explicitly by assessing toxicity, viability, or functional neuronal readouts, particularly given prior reports that WT TDP-43 overexpression alone is neurotoxic.

      We agree that this is an important point, but the main goal of this study was to dissect the cellular pathways/mechanisms that govern TDP-43 phase separation. We feel that the requested experiments are beyond the scope of the current study.

      (13) Clearly state the number of biological replicates used for each RNA-seq condition. Establish baseline transcriptional differences between WT and mutant TDP-43 prior to assessing the effects of nuclear export inhibition. Include PCA plots and heatmaps, including all samples.

      As mentioned above, we have decided to remove the RNAseq data from the manuscript to save room for new results.

      (14) Given the role of TDP-43 as an RNA-binding protein, consider including splicing analyses to assess whether nuclear export inhibition preserves or disrupts TDP-43-dependent RNA processing.

      We thank the reviewer for this suggestion. However, we feel that the proposed experiments are beyond the scope of the current study.

      (15) Improve clarity of transcriptomic visualizations (e.g., GO-term plots) and explicitly define all group labels used (e.g., Group A vs Group B).

      We have removed the RNAseq data.

      (16) Ensure consistent use of disease terminology (ALS vs FTD) throughout the manuscript, e.g., lines 222 and 244.

      We have checked the usage of these terms to make sure they are accurately used.

      (17) Correct figure and axis labeling errors (e.g., Figure 3A x-axis range).

      Figure 3A indicates the Z score distribution of the entire human genome. As stated on page 6, 21,404 genes were targeted.

      (18) Avoid overstatements in the Discussion that are not directly supported by the presented data, particularly regarding the interpretation of proteasome inhibition and gel-like anisosome states.

      We have revised our discussion substantially to tone down our conclusions.

      (19) Clarify the rationale for switching between different nuclear export inhibitors across experiments and discuss whether results were consistent across compounds.

      In the acute experiments down with the cancer cell line, we used LMB because it is potent and well characterized. In organoid experiment, we switched to KPT-276 because it is better tolerated by organoids, especially during longer treatment.

      Reviewer #3 (Recommendations for the authors):

      Major concerns that require clarification or further strengthening:

      (1) The connection between nuclear export and liquid-solid phase transition is not clear. The 2KQ mutant forms nuclear anisosomes. The manuscript does not provide data about its nuclear-cytoplasmic distribution normally, nor how the distribution is changed upon nuclear export inhibition or enhancement. In Figure 5I, it is unclear whether the anisosomes are in the nucleus or cytoplasm. The dynamics of nuclear vs cytoplasmic anisosomes should be measured separately. What is the mechanism that promotes nuclear export and changes the dynamics, especially nuclear anisosomes?

      As mentioned by the reviewer, the 2KQ mutant forms anisosomes only in the nucleus. This was documented in Yu, H. et al., Science 371 (2021), and also shown in our Figure 4A, F, Figure 5A. Figure 5A also shows that nuclear export inhibition does not change anisosome localization, only making them bigger while reducing the numbers. For Figure 5I, the control anisosomes in untreated cells are nuclear while under mCh-XPO1 expressing condition, only those present in the cytoplasm were randomly selected for bleaching.

      (2) Figure 5J, no obvious XPO1 is sequestered to anisosomes, as described in lines 208-209.

      Unlike Figure 5G, this experiment studied the localization of endogenous XPO-1 by immunostaining. As discussed in Yu et al., Science 371 (2021), proteins inside anisosomes could not be stained by antibodies due to an accessibility problem. This explains why we could only detect reduced XPO1 after anisosome induction.

      (3) Figure 6A, the localization of phosphor-TDP-43 is not clear. And it is not clear what cell types contain the aggregates. Higher-resolution images need to be included. The mechanism by which XPO1 inhibition reduces TDP-43 aggregation requires further validation. It remains unclear whether it is directly mediated through altered nucleocytoplasmic transport of TDP-43.

      We agree that it is technically challenging to visualize the precise subcellular localization of p-TDP-43 in 3D organoids. In the manuscript that reports the characterization of the 3D organoids, we dissociated cells from the 3D organoids by trypsin digestion and plated them out in 2D before immunostaining and imaging. We could clearly see p-TDP-43 co-localizes with the neuronal marker TUJ1 and is localized outside of nucleus (see figure 1 of https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full)

      In the newly added Figure 6, we used a semi-permeabilized cell system to dissect the phase separation dynamics of TDP-43 2KQ in cells treated with the nuclear export inhibitor LMB. Our data suggests that nuclear export inhibition alters the nuclear environment, making it more favorable for the liquid phase of TDP-43. This is dependent on nuclear RNA.

      (4) XPO1 controls the export of numerous essential proteins, and its inhibition can produce broad, potentially toxic effects unrelated to TDP-43. The manuscript should include a discussion of these off-target consequences.

      We thank the reviewer for this point. Given the new data in Figure 6, we now add some more discussion on the potential mechanism by which nuclear export inhibition modulates TDP-43 phase separation. This can be found on page 10.

      References:

      Zhang, Q. et al. A human forebrain organoid model phenocopies dysregulated RNA and protein homeostasis in ALS/FTD-associated TDP-43 proteinopathies. bioRxiv (2025). (https://www.biorxiv.org/content/10.1101/2025.11.09.687455v1.full

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer #1 (Public review):

      Summary:

      This preprint investigates the molecular mechanism by which warm temperature induces female-to-male sex reversal in the ricefield eel (Monopterus albus), a protogynous hermaphroditic fish of significant aquacultural value in China. The study identifies Trpv4 - a temperature-sensitive Ca²⁺ channel - as a putative thermosensor linking environmental temperature to sex determination. The authors propose that Trpv4 causes Ca²⁺influx, leading to activation of Stat3 (pStat3). pStat3 then transcriptionally upregulates the histone demethylase Kdm6b (aka Jmjd3), leading to increased dmrt1 gene expression and ovo-testes development. This work aims to bridge ecological cues with molecular and epigenetic regulators of sex change and has potential implications for sex control in aquaculture.

      Strengths:

      (1) This study proposes the first mechanistic pathway linking thermal cues to natural sex reversal in adult ricefield eel, extending the temperature-dependent sex determination paradigm beyond embryonic reptiles and saltwater fish

      (2) The findings could have applications for aquaculture, where skewed sex ratios apparently limit breeding efficiency

      Weaknesses:

      Although the revised manuscript represents an improvement over the original version, substantial weaknesses remain.

      We thank you for the critical comments. We have responded to your concerns by a point by point manner, and please see detail below.

      Scientific Concerns

      (1) Western blot normalization and exposure: The loading controls (GAPDH) in Fig. S3C appear overexposed, as do several Foxl2 blots. Because these signals are likely outside the linear range, I am not convinced that normalization is reliable. This raises concerns about the validity of the quantified results.

      We thank you for the concerns. We have repeated the experiments, and new blots were loaded in Fig.S3C.

      (2) Antibody validation and referencing (Line 776): The authors need to refer explicitly to figures demonstrating antibody validation. At present, these data are provided only as a supplementary file that is not cited in the manuscript. In addition, the Sox9a antibody appears to yield indistinguishable signals in control and RNAi conditions, suggesting that it may not recognize eel Sox9a. This issue is not addressed by the authors. Furthermore, antibody validation Western blots should be quantified.

      We thank you for the comments. We have repeated the siRNA experiments to show the specificity of the antibodies used. This file, named as the supplementary file 1, is now cited in “WB analysis” in the Materials and Method part. As required, the antibody validation of WB are uploaded in the supplementary file 1. Antibody validation for WB are now quantified, and please see the new figure 3 and supplementary Figure 3.

      (3) Unclear sample sizes (N values): Sample sizes remain unclear for several figures:

      (a) Fig. 3F - No N value is provided. Each graph shows three data points; does this indicate that only three samples were quantified? If ten samples were collected, why were all not quantified?

      We apologize for the confusion. Three data points were previously used to shown data of 3 replicates. In new figure 3F, 10 randomly selected sections were imaged, and the data are shown. In the revised manuscript, the sample numbers (the N values) are added, and all the information can be found in the figure legend.

      (b) Fig. 4 - No N values are reported.

      Now N values are added. Please see the figure legend.

      (c) Fig. 5A - Again, only three data points are shown per group, despite the apparent availability of twelve samples. The rationale for this discrepancy is not explained.

      We apologize for the wrong data representation. Now all the data points are shown in Figure 5.

      (4) qRT-PCR normalization: The manuscript does not specify the reference gene(s) used for qRT-PCR normalization. Although expression levels are reported as "relative," neither the identity of the reference gene(s) nor the justification for their selection is provided.

      We now have specify the reference gene in “Quantitative real-time PCR (qPCR) experiments” part in the Materials and Methods section.

      (5) Specificity of key antibodies: While the authors have made some effort to validate anti-Amh, anti-Sox9, and anti-Dmrt antibodies, the results remain incomplete. The Amh and Dmrt antibodies detect reduced protein levels following knockdown of their respective targets, which is encouraging. However, the Sox9a antibody shows no difference between control and RNAi conditions, suggesting it does not recognize eel Sox9. This is not acknowledged in the manuscript. In addition, no validation data are presented for Foxl2. Antibody validation data must be clearly referenced in the main text and presented in an interpretable and quantitative manner.

      The antibody specificity is very important. For that reason, we have generated at least two different antibodies for each target protein, using full-length or small peptide as antigen. We have repeated the experiments for key antibodies such as Dmrt1 and Sox9a. IF and WB results clearly showed the specificity of the antibodies.

      Author response image 1.

      Foxl2 antibody has also been reported in ricefield eel (Hu et al. SCIENTIFIC REPORTS | 4: 6884 | DOI: 10.1038/srep06884, Molecular cloning and analysis of gonadal expression of Foxl2 in the ricefield eel Monopterus albus).

      After short term warm temperature exposure, only a small portion of somatic cells in ovary may be induced to express the male markers. As different techniques have different capacity (sensitivity), some techniques were more easy to detect that change. For instance, qPCR and WB are ready to detect it, whereas IF is a little difficult in obtaining good quality data.

      (6) Immunofluorescence data quality: The immunofluorescence images remain difficult to interpret. I strongly encourage the authors to enlarge the image panels and to present monochrome images (white signal on black background). The current presentation severely limits interpretability.

      We thank you for the comments. We think that our IF images are of decent quality. Due to the limits of the Figure space (already busy for Figure 3), enlarging the image panels or presenting additional monochrome images will compromise the quality of other data. Alternatively, if you still concern its quality, we can put it in the supplementary.

      Author response image 2.

      (7) Unreferenced supplementary figure: Fig. S4 is included in the submission but is not referenced anywhere in the manuscript text.

      We now have renamed the supplementary Figures. And we have double checked the text to make sure all Figure information is correctly referenced. Figure S4 is removed, as it is not necessary.

      (8) Fig. 5B image resolution: The micrographs in Fig. 5B are too small to allow meaningful evaluation of the data.

      Now new Figure 5B images with higher resolution were shown.

      (9) Unexplained data inclusion (Fig. 5E): Fig. 5E includes a pERK blot that is not mentioned in the Results section. The rationale for including these data is unclear.

      Previous work have shown that FGF/ERK signaling may play a role in sex change of ricefield eel (in Chinese). We therefore examined the Erk activity to explore whether it is involved in sex reversal. The results showed that pErk was comparable between ovary and ovotestis. At your suggestion, we decided to remove the data.

      (10) Poor blot quality (Fig. S3C): The blots in Fig. S3C exhibit high background and overexposure. I am concerned about the reliability of the quantification shown in panel D.

      The experiments have been repeated at least three times, and similar results were obtained. We now have replaced some of the WB that were of high background or overexposure.

      (11) Poor blot quality (Fig. S5G): The Stat3 blots in Fig. S5G contain numerous white artifacts, raising concerns about their suitability for normalization in panel H.<br />

      We now have repeated the experiments, and uploaded a new representative blot with better quality.

      (12) Missing controls (Fig. 6E): Fig. 6E lacks controls for HO-3867 and Colivelin treatments alone. Without these controls, it is not possible to determine whether the reported effects are meaningful.

      We thank you for the comments. We now have added the data required (with HO-3867 and Colivelin treatments alone).

      (13) Graphical presentation: The use of a light blue-to-pink gradient in bar graphs throughout the manuscript does not aid interpretation. I recommend using more distinct colors (e.g., red, orange, green, blue, purple, gray, black) to improve clarity.

      We thank you for the comments. We now have changed the blue-to-pink gradient to more distinct color system to better present the data. Please see the detail in the revised Figures.

      In summary, the interpretation of the study remains limited by persistent issues related to data presentation, image quality, and reagent specificity.

      We thank you for the critical comments about our data, in particular for antibody specificity and image quality, and the detailed instruction for how to better present the data. Answering your questions have greatly improved the quality of the manuscript. We admit that due to the technique challenging (with different conditions and different doses of small molecules) and higher cost of animal experiments, some of the WB or IF experiments may not be of high standards.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      Editorial Concerns

      (1) Overstatement of conclusions: In lines 16-18, the authors state that Trpv4 "mediates" warm temperature-driven sex reversal. This claim is too strong given the data and should be toned down.

      We agree with our editorial comment about the overstatement. Now it reads “Trpv4 links environmental temperature to testicular differentiation in ricefield eel”.

      (2) Misuse of statistical language (Line 213): The term "significant" is used where statistical significance was not measured. The wording should be revised.

      We thank you for the point, and now have replaced “significant” to “marked”.

      (3) Terminology (Line 238): The term "co-expression" is inaccurate in this context. I suggest replacing it with "co-upregulation."

      We thank you for the point, and have changed it accordingly.

      (4) Drug description errors (Lines 241-242): The manuscript incorrectly identifies which drug functions as an agonist and which as an antagonist. This caused considerable confusion and must be corrected.

      We have carefully checked the sentence, and it was correct, as RN1734 and GSK1016790A are known Trpv4 specific antagonist and agonist, respectively.

      (5) Gene examples missing (Lines 247-250): The authors should explicitly name the testis-biased and ovary-biased genes referred to in this section.

      We thank you for the point, and now it reads “warm temperature exposure increased the expression of testicular differentiation genes such as dmrt1 and gsdf, accompanied by moderately decreased expression of ovarian differentiation genes such as cyp19a1a and foxl2”.

      (6) Lack of experimental context (Lines 322-324): Rather than simply listing the drugs used, the authors should briefly explain what each compound inhibits or activates and why it was employed.

      We have described this in the manuscript. The information of pStat3 activator and inhibitor has been described in Lines 305-309, as “HO-3867, a curcumin analogue, is a selective pStat3 inhibitor, which blocks pStat3 activity by directly binding to Stat3 DNA binding domain, and Colivelin is a potent synthetic peptide activator of pStat3, which increases pStat3 levels by acting through the GP130/IL6ST complex”, and the rationale has been stated in lines 32--322 as “To functionally demonstrate that pStat3 signaling is downstream of Trpv4, rescue experiments were performed by injecting into ovaries with individual and combined small molecules”.

      (7) Discussion of evolutionary differences: The Discussion misses an important opportunity to address why Stat3 activates kdm6b in ricefield eel but represses it in turtles. It is difficult to reconcile how the same transcription factor could exert opposite effects on the same gene during sex determination without additional context. A comparison of kdm6b regulation and sequence conservation between turtles and ricefield eel would strengthen this section.

      We have downloaded the promoter sequences of red eared turtle and ricefield eel. Based on the DNA sequences (Author response image 3), the similarity (conservation) was low between the two species.

      Author response image 3.

      It was appeared that DNA around the Stat3 binding sites in turtle are GC rich (CpG island), which may be subjected to DNA methylation modification, whereas the DNA in ricefield eel are not GC rich.The observations imply that the role of pStat3 is to promote the repression of kdm6b in turtle but the activation of kdm6b in ricefield eel.

      Moreover, our unpublished data showed that Trpv4-controlled calcium signaling is required to remove the repressive histone modification H3K27me3 at the kdm6b gene. If pStat3 is downstream of Trpv4 in this case, it supports again that Trpv4-pStat3 axis activate kdm6b in ricefield eel.

      Warm temperature promotes female sex in turtle but male sex in ricefield eel. If pStat3 is mediating Trpv4, it is not surprising that it represses kdm6b in turtle but activate it in ricefield eel.

      Based on above, we have added some sentences in the discussion part, and it reads “We reasoned that a yet-unidentified co-factor may determine whether Stat3 is a transcriptional repressor or activator. A comparison of promoter sequences of kdm6b between turtle and ricefield eel supported this”.

      (8) Supplementary figure formatting: Supplementary figures should be provided in accordance with eLife formatting guidelines.

      We have now formatted the supplementary figures that are in accordance with eLife formatting requirement. Please see the new uploaded supplementary figures.

      In sum, the interpretations are still limited by the above concerns regarding data presentation and reagent specificity.

      We thank our editor for the inspiring comments. We believe we have addressed all the major concerns by our editor.

    1. Author response:

      eLife Assessment

      This study provides a valuable advance in understanding how disordered proteins interact with cell membranes by identifying the sequence rules that enable aromatic residues to penetrate deeply into the membrane interior. The integration of complementary computational approaches, including molecular simulations, large-scale sequence analysis, and the development of an online prediction server, makes the work potentially impactful for the membrane protein and intrinsically disordered protein communities. The evidence supporting the main conclusions is generally convincing, although its transferability across diverse membrane compositions and its validity as a prediction tool for real protein-membrane systems remain to be further established.

      We thank the editors for recognizing our study as a valuable advance. This work lays a solid foundation for future developments to account for diverse membrane compositions and further refinements after additional experimental tests.

      Public review:

      Reviewer #1:

      A primary limitation is the heavy reliance on computational modeling. Training for AroMIP is generated using PPM rather than direct experimental measurements, and so the model may primarily reproduce PPM behavior rather than true membrane insertion thermodynamics. Moreover, all simulations use a single lipid composition (POPC:POPS:PIP<sub>2</sub> 70:25:5), but biological membranes vary substantially in cholesterol, cardiolipin, and acidic lipid content. Whether AroMIP's predictions transfer to diverse lipid environments remains untested. The 5% PIP<sub>2</sub> concentration used in the simulations is higher than that of a normal mammalian cell and may therefore overemphasize electrostatic contributions. Applicability beyond short 9-residue motifs is unclear, as longer-range interactions or secondary structure in full-length IDRs could modulate insertion in ways the current model does not capture. This could be considered for future development.

      The reviewer’s point on our reliance on PPM for training, a single lipid composition, and potential effects beyond a 9-residue motif is well taken. Regarding PPM, we chose it as the optimal compromise for high-throughput data. However, we complemented the high-throughput PPM data with experimental data on an initial set of 10 peptides. Moreover, we validate AroMIP on an additional 12 IDRs (intrinsically disordered regions; Table S2). On membrane composition, we now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph). On potential effects beyond a 9-residue motif, we now add justification and note neglected factors for future developments (paragraph running from p. 19-20), as suggested by the reviewer.

      Reviewer #2:

      (1) Aromatic residues have been shown to partition preferentially to the headgroup region of the lipid bilayer. Most of the papers on this problem were published in the mid 1990s to early 2000s. Some of the most important papers in this regard are the following: von Heijne, Annu. Rev. Biophys. Biomol. Struct. 1994, 23, 167-192; Doyle et al. Science 1998, 280, 69-77; Landolt-Marticorena, et al. J. Mol. Biol. 1993, 229, 602-608; Killian & von Heijne, TIBS 2000, 25, 429-434; Marx & Fleming J. Am. Chem. Soc. 2021, 143, 764-772. Strangely enough, none of these articles is cited.

      We have now citations to the Landolt-Marticorena paper and the von Heijne reviews (refs 25-27). The Doyle paper is not particularly relevant. As for the Fleming paper, we cited a 2016 JACS paper (original ref 27; now ref 30) that specifically dealt with aromatic residues.

      (2) This is the most important point and the most serious weakness. The authors find that the PPM method is able to reproduce the results from MD simulations, and the AroMIP model is able to perform well in comparison with PPM and MD, after training AroMIP on a large set of IDR sequences (intrinsically disordered protein regions) of the human proteome. The defining feature of the AroMIP calculation is the recognition of the importance of flanking residues in the membrane-insertion propensity of a sequence containing a central aromatic residue. All this sounds good. However, this is all theoretical. There is no connection to experiment or to any method that draws from experiment. The entire approach relies on the assumption that the MD simulations produce the correct results. There is no proof of the correctness of anything. As one of the greatest physicists of our times, Richard Feynman, wrote, "The test of all knowledge is experiment. Experiment is the sole judge of scientific "truth".”

      We emphasize that we have presented substantial experimental support for AroMIP. It correctly predicts the membrane insertion status of the initial set of 10 peptides, which were characterized experimentally. In addition, we validated AroMIP on an additional set of 12 IDRs (Table S2), most of which were characterized by experimental techniques including solution and solid-state NMR, fluorescence, H/D exchange, and cryo-EM. Lastly, we now show good correlation between our insertion scores and binding free energies calculated from the scale determined experimentally by White and co-workers (new Figure S10; p. 15, second paragraph).

      (3) The drawings in Figures 2 and 3 are incorrect and misleading. The size of the Tryptophan side chain is about 5.5 Å, whereas one-half of the bilayer ("a monolayer") thickness is about 15 Å. But in the figures, the lipid length and the Trp side chain seem about the same size. This is incorrect even in a qualitative sense.

      We have now revised these figures.

      Reviewer 3:

      (1) Membrane composition and lipid shape characteristics: The authors chose to use a model membrane bilayer of a distinct lipid composition, POPC: POPS: PI4,5P2 (70:25:5 molar ratio), for their all-atom simulations of the various model peptides. While this may be pertinent for some of these peptides, it is not for many, such as sequence 2 derived from Drp1, which preferentially binds target conical lipids such as cardiolipin (CL) and phosphatidic acid (PA). The rationale behind using PI4,5P2, which can induce positive membrane curvature when sequestered, versus CL and PA, which both induce negative membrane curvature, is not explained.

      We now acknowledge the limitation of our work based on a single composition and point to future developments of AroMIP involving membrane-specific parameterization (p. 19, 3rd paragraph). In this Discussion paragraph, we also speculate that conical lipids, by promoting membrane defects, may facilitate membrane insertion.

      (2) Parallel vs. perpendicular peptide orientation of sequence 2 in peripheral Drp1-lipid interactions: On page 11, the authors state that their simulation results of sequence 2 derived from Drp1 "contrasts with a transmembrane orientation proposed by Mahajan et al." However, upon review, a transmembrane orientation for this region has never been proposed anywhere. Drp1 is a peripheral membrane protein that reversibly binds CL- and PA-containing membranes via its intrinsically disordered variable domain containing an aromatic-centered WRG motif. Indeed, the model presented in Figure 9 of Mahajan et al. displays a peripheral and parallel orientation of the transiently helical WRG-containing motif rather than a transmembrane (i.e., across the bilayer) orientation. While the authors can distinguish between a parallel vs. perpendicular orientation of this sequence relative to the plane of the membrane bilayer surface from their simulations, suggesting that previous studies indicated a transmembrane orientation for Drp1 is disingenuous and misleading. The term "transmembrane" should be removed or replaced, as it presents a wrong image.

      We have now deleted the sentence mentioning “transmembrane orientation”.

      (3) Mutational analysis of W vs. F in membrane insertion of W-centered insertion motifs and vice versa: The PPM-based workflow suggests that F-centered sequences have the highest membrane insertion properties as opposed to W-centered ones. A W552F mutation in the WRGML sequence of Drp1 was, however, found to impair function. How do the authors rationalize this? A cross-mutational analysis of W vs. F in W-centered motifs and F-centered motifs is warranted.

      AroMIP predicts a membrane insertion propensity of 0.782 for the WRGML sequence and a moderately higher propensity, 0.837, with a W552F mutation. This increase contradicts the experimental observation of a 3.6-fold increase in membrane binding affinity by Mahajan et al. We now speculate that the specific lipid, cardiolipin, as the reason for the discrepancy (p. 19, 3rd paragraph). This discrepancy provides a concrete example for the need to account for membrane composition in future developments.

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This valuable study combined careful computational modeling, a large patient sample, and replication in an independent general population sample to provide a computational account of a difference in risk-taking between people who have attempted suicide and those who have not. It is proposed that this difference reflects a general change in the approach to risky (high-reward) options and a lower emotional response to certain rewards. Evidence for the specificity of the effect to suicide, however, is incomplete, which would require additional analyses.

      We thank the editors and reviewers for this important assessment. Based on clinical interviews, we included patients with and without suicidality (S<sup>+</sup> and S<sup>-</sup> groups). However, in line with suicidal-related literature (e.g., Tsypes et al., 2024), two groups also differed substantially in the severity of symptoms (see Table 1). To address the request for evidence on specificity to suicidality beyond general symptom severity, we performed separate linear regressions to explain in gambling behaviour, value-insensitive approach parameter (β<sub>gain</sub>), and mood sensitivity to certain rewards (β<sub>CR</sub>) with group as a predictor (1 for S<sup>+</sup> group and 0 for S<sup>-</sup> group) and scores for anxiety and depression as covariates. Results remained significant after controlling anxiety and depression (ps < 0.027; Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) on the clinical questionnaire to extract the orthogonal components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. We then performed linear regressions using these components as covariates to control for anxiety and depression. Our main results remained significant (ps < 0.027; Table S9). We believe that these analyses provide evidence that the main effects on gambling and on mood were specific to suicide.

      Moreover, as Reviewer 3 pointed out, these “absence of evidence” cannot provide insights of “evidence of absence”. Although we median-split patients by the scores of general symptoms (e.g., depression and anxiety-related questionnaires) and verified no significant differences in these severities (Figure S11), we additionally conducted Bayesian statistics in gambling behavior, value-insensitive approach parameter, and mood sensitivity to certain rewards. BF<sub>01</sub> is a Bayes factor comparing the null model (M<sub>0</sub>) to the alternative model (M<sub>1</sub>), where M<sub>0</sub> assumes no group difference. BF<sub>01</sub> > 1 indicates that evidence favors M<sub>0</sub>. As can be seen in Table S7, most results supported null hypothesis, suggesting that general symptoms of anxiety and depression overall did not influence our main results. Overall, we believe that these analyses provide compelling evidence for the specificity of the effect to suicide, above and beyond depression and anxiety.

      Beyond these specific findings, this work highlights the broader utility of computational modelling and mood to better understand behavioral effect, showing how to use both mood and choice data to better comprehend a psychiatric issue.

      Please see Tables S7, S8, S9 and our revisions below:.

      Page 17:

      “Within patients, this group effect on gambling rate remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.024; also see Figure S11, Table S7 and Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, (ps < 0.001), we performed Principal Components Analysis (PCA) to extract main components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. To further control for anxiety and depression, linear regression using these components as covariates revealed that the group effect on gambling rate remained significant (p = 0.024; Table S9).”

      Pages 18-19:

      “Within patients, this group effect on the approach parameter remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.027; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on approach parameter remained significant (p = 0.027; Table S9).”

      Page 21:

      “Within patients, this group effect on βCR remained significant after controlling for gambling rate, earnings, mood-related outcome effect, mood drift effect, sex, illness duration, family history, diagnosis, and various medications use (ps < 0.032), as well as general symptoms (e.g., depression and anxiety; p = 0.001; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on this mood parameter remained significant (p = 0.001; Table S9).”

      Page 27:

      “Beyond these specific findings, this work highlights the broader utility of computational modelling and mood to better understand behavioral effect, showing how to use both mood and choice data to better comprehend a psychiatric issue.”

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors use a gambling task with momentary mood ratings from Rutledge et al. and compare computational models of choice and mood to identify markers of decisional and affective impairments underlying risk-prone behavior in adolescents with suicidal thoughts and behaviors (STB). The results show that adolescents with STB show enhanced gambling behavior (choosing the gamble rather than the sure amount), and this is driven by a bias towards the largest possible win rather than insensitivity to possible losses. Moreover, this group shows a diminished effect of receiving a certain reward (in the non-gambling trials) on mood. The results were replicated in an undifferentiated online sample where participants were divided into groups with or without STB based on their self-report of suicidal ideation on one question in the Beck Depression Inventory self-report instrument. The authors suggest, therefore, that adolescents with decreased sensitivity to certain rewards may need to be monitored more closely for STB due to their increased propensity to take risky decisions aimed at (expected) gains (such as relief from an unbearable situation through suicide), regardless of the potential losses.

      Strengths:

      (1) The study uses a previously validated task design and replicates previously found results through well-explained model-free and model-based analyses.

      (2) Sampling choice is optimal, with adolescents at high risk; an ideal cohort to target early preventative diagnoses and treatments for suicide.

      (3) Replication of the results in an online cohort increases confidence in the findings.

      (4) The models considered for comparison are thorough and well-motivated. The chosen models allow for teasing apart which decision and mood sensitivity parameters relate to risky decision-making across groups based on their hypotheses.

      (5) Novel finding of mood (in)sensitivity to non-risky rewards and its relationship with risk behavior in STB.

      Weaknesses:

      (1) The sample size of 25 for the S- group was justified based on previous studies (lines 181-183); however, all three papers cited mention that their sample was low powered as a study limitation.

      We thank the Reviewer for rising this concern. We agree that the sample size for S<sup>-</sup> group (n=25) is modest, and the prior studies we cited also acknowledged limited power. We wanted to point out that we obtained a comparable sample size to a prior study. In the revision, we therefore updated the section to justify this sample size in which we acknowledge the limited power of our study in the limitation section. Please see our clarification below:

      Page 32:

      “Third, despite replicating our main results in an independent dataset (n=747), the modest S<sup>-</sup> subgroup size (n=25) has a limited statistical power.”

      (2) Modeling in the mediation analysis focused on predicting risk behavior in this task from the model-derived bias for gains and suicidal symptom scores. However, the prediction of clinical interest is of suicidal behaviors from task parameters/behavior - as a psychiatrist or psychologist, I would want to use this task to potentially determine who is at higher risk of attempting suicide and therefore needs to be more closely watched rather than the other way around (predicting behavior in the task from their symptom profile). Unfortunately, the analyses presented do not show that this prediction can be made using the current task. I was left wondering: is there a correlation between beta_gain and STB? It is also important to test for the same relationships between task parameters and behavior in the healthy control group, or to clarify that the recommendations for potential clinical relevance of these findings apply exclusively to people with a diagnosis of depression or anxiety disorder. Indeed, in line 672, the authors claim their results provide "computational markers for general suicidal tendency among adolescents", but this was not shown here, as there were no models predicting STB within patient groups or across patients and healthy controls.

      Thank you for these thoughtful comments. Our study focuses on why adolescent patients with suicidality have increased risk behavior, aiming to provide a mechanism-based target for suicide prevention. Therefore, our dependent variable in the mediation model was gambling behavior. We also agree that the clinically relevant question is whether suicidality can be predicted from task-derived behavior/parameters. We thus used risky behavior and the potential mental parameters to predict STB. Linear regressions showed that gambling behavior, as well as the value-insensitive approach parameter, can predict suicidal symptom scores among patients (former: β = 9.189, t = 2.004, p = 0.048; latter: β = 5.587, t = 2.890, p = 0.005). In healthy controls, these predictions failed (gambling behavior: β = 1.471, t = 0.825, p = 0.411; approach: β = 0.874, t = 1.178, p = 0.241). These results suggest that clinical relevance of these findings apply exclusively to people with a diagnosis of depression or anxiety disorder. We found same patterns for the mood parameter (mood sensitivity to certain rewards: patients: β = -28.706, t = -2.801, p = 0.006; healthy controls: β = -2.204, t = -0.528, p = 0.599). In sum, we believe that our statement of “computational markers for general suicidal tendency among adolescents” is reasonable now. Please see our revisions below:

      Page 17:

      “Furthermore, linear regression showed that gambling rate can predict the current suicidal ideation score (BSI-C, β = 9.189, t = 2.004, p = 0.048) among patients, but not among HC (β = 1.471, t = 0.825, p = 0.411), suggesting that gambling behavior has patient-specific predictive utility for suicidal symptoms.”

      Page 19:

      “Furthermore, linear regression showed that approach parameter can predict the current suicidal ideation score (β = 5.587, t = 2.890, p = 0.005) among patients, but not among HC (β = 0.874, t = 1.178, p = 0.241), suggesting that value-insensitive approach parameter has patient-specific predictive utility for suicidal symptoms.”

      Page 21:

      “Furthermore, linear regression showed that mood sensitivity to CR can predict the current suicidal ideation score (β = -28.706, t = -2.801, p = 0.006) among patients, but not among HC (β = -2.204, t = 0.528, p = 0.599), suggesting that mood sensitivity to CR has patient-specific predictive utility for suicidal symptoms.”

      (3) The FDR correction for multiple comparisons mentioned briefly in lines 536-538 was not clear. Which analyses were included in the FDR correction? In particular, did the correlations between gambling rate and BSI-C/BSI-W survive such correction? Were there other correlations tested here (e.g., with the TAI score or ERQ-R and ERQ-S) that should be corrected for? Did the mediation model survive FDR correction? Was there a correction for other mediation models (e.g., with BSI-W as a predictor), or was this specific model hypothesized and pre-registered, and therefore no other models were considered? Did the differences in beta_gain across groups survive FDR when including comparisons of all other parameters across groups? Because the results were replicated in the online dataset, it is ok if they did not survive FDR in the patient dataset, but it is important to be clear about this in presenting the findings in the patient dataset.

      Thank you for raising the important issue of multiple testing and for asking us to clarify exactly which tests were covered by the FDR procedure. In the clinical dataset we conducted a large number of inferential tests (χ<sup>2</sup>, t-tests, ANOVAs, regressions) spanning: (i) group differences in demographic/clinical characteristics; (ii) sanity checks (e.g., anxiety/depression questionnaires); (iii) primary hypotheses (e.g., group differences in risky behavior); (iv) model-based analyses (parameter checks and between-group contrasts); and (v) control/sensitivity analyses. Post-hoc t-tests were performed only when the three-group ANOVA was significant. This yielded >150 p-values. FDR was applied using all these p-values. Please see Supplementary Note 8.

      (4) There is a lack of explicit mention when replication analyses differ from the analyses in the patient sample. For instance, the mediation model is different in the two samples: in the patient sample, it is only tested in S+ and S- groups, but not in healthy controls, and the model relates a dimensional measure of suicidal symptoms to gambling in the task, whereas in the online sample, the model includes all participants (including those who are presumably equivalent to healthy controls) and the predictor is a binary measure of S+ versus S- rather than the response to item 9 in the BDI. Indeed, some results did not replicate at all and this needs to be emphasized more as the lack of replication can be interpreted not only as "the link between mood sensitivity to CR and gambling behavior may be specifically observable in suicidal patients" (lines 582-585) - it may also be that this link is not truly there, and without a replication it needs to be interpreted with caution.

      Thank you for these important comments. This study focused on cognitive and affective computational mechanisms underlying increased risky behavior in STB. Accordingly, we compared patients with STB (S<sup>+</sup>) with patients without STB (S<sup>-</sup>) and healthy controls (HC) to examine the effects of STB on risky behavior. Therefore, group comparison, instead of dimensional measure of suicidal symptoms by Beck Scale for Suicidal Ideation, can answer our research questions directly.

      To enhance consistency between the clinical and replication datasets, we included all participants in each dataset when performing the mediation analysis. Given that S<sup>-</sup> and HC did not differ in gambling behavior or the approach parameter in the clinical dataset, we merged these two groups. In the replication dataset, to mirror the S<sup>+</sup> vs. S<sup>-</sup> contrast used clinically, we categorized the general sample into S<sup>+</sup> and S<sup>-</sup> based on BDI item 9. The mediation results remained significant in both datasets (the clinical dataset: a×b = 0.321, 95% CI = [0.070, 0.549], p = 0.016; the replication dataset: a × b = 0.143, 95% CI = [0.016, 0.288], p = 0.031), suggesting that STB is associated with increased risk behavior via stronger approach motivation.

      We also acknowledge the non-replication of the correlation between gambling behavior and mood sensitivity to certain rewards in the online sample. While this pattern might indicate that the link is specific to suicidal patients, it may also reflect sample-specific or unstable effects; thus, we now state this explicitly and interpret the finding with caution. Please see our revisions below:

      Page 15:

      “We next verified our results in an independent dataset, including the same task and BDI questionnaire in 747 general participants (500 females; age: 20.90±2.41)[46]. One item in BDI involves the measurement of STB. In item 9 of BDI, participants chose one option that describes them best: Option 1, “I don't have any thoughts of killing myself.”; Option 2, “I have thoughts of killing myself, but I would not carry them out.”; Option 3, “I would like to kill myself.”; Option 4, “I would kill myself if I had the chance.”. In line with the current definition of S<sup>+</sup>/S<sup>-</sup> in the clinical dataset, we identified S<sup>+</sup> group as choosing Option 2, 3, or 4, while participants selecting Option 1 were categorized as S<sup>-</sup> group.”

      Page 19:

      “Given significant correlations between group, approach parameter, and gambling rate for gain trials (ps < 0.017), we further conducted a mediation analysis with the assumption of the mediating effect of approach motivation of suicidality on the risk behavior. Given that we aimed to test the effect of STB, with S<sup>-</sup> and HC as controls, and given that S<sup>-</sup> and HC did not differ in gambling behavior or in the approach parameter, we merged these two groups for the mediation analysis. Results supported our hypothesis (a×b = 0.321, 95% CI = [0.070, 0.549], p = 0.016; Figure 2C), confirming that suicidal thoughts and behavior increase risk behavior through stronger approach motivation.”

      Page 26:

      “However, we did not observe any significant correlation between mood sensitivity to CR and gambling behavior (ps > 0.389), which suggests that the link between mood sensitivity to CR and gambling behavior may be specifically observable in suicidal patients. Alternatively, this non-replicated result may also reflect sample-specific or unstable effects, which needs to be interpreted with caution.”

      (5) In interpreting their results, the authors use terms such as "motivation" (line 594) or "risk attitude" (line 606) that are not clear. In particular, how was risk attitude operationalized in this task? Is a bias for risky rewards not indicative of risk attitude? I ask because the claim is that "we did not observe a difference in risk attitude per se between STB and controls". However, it seems that participants with STB chose the risky option more often, so why is there no difference in risk attitude between the groups?

      Thank you for pointing out the ambiguity. In our manuscript, “motivation” and “risk attitude” are defined at the computational level. Following prior work with this task Rutledge et al., (2015, 2016), we decompose observed gambling into (i) value-dependent valuation parameters that capture risk attitude (e.g., risk aversion and loss aversion, which scale the subjective value of outcomes), and (ii) value-insensitive, valence-dependent biases that capture approach/avoidance motivation. Accordingly, a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups which is what we observe for S<sup>+</sup> vs. controls. We have clarified this point in the computational modeling section.

      Pages 12-13:

      “Please note that a higher gambling rate does not imply a change in risk attitude per se: it can arise from an increased value-insensitive approach bias even when risk-attitude parameters are comparable between groups. Risk attitude is indeed conceptualized in economics as the curvature of the utility function (i.e., the subjective value) of the objective outcomes, with concave curves associated with risk aversion, and convex curves associated with risk seeking [54,56]. By contrast, the approach or avoidance bias apply to all the value. A possible interpretation of the approach bias is that participant approach the option with the highest possible gain (the lottery) in the gain frame; the avoidance bias would then reflect a tendency to systematically avoid the highest potential losses (the lottery) in the loss frame.”

      Reviewer #2 (Public review):

      Summary:

      This article addresses a very pertinent question: what are the computational mechanisms underlying risky behaviour in patients who have attempted suicide? In particular, it is impressive how the authors find a broad behavioural effect whose mechanisms they can then explain and refine through computational modeling. This work is important because, currently, beyond previous suicide attempts, there has been a lack of predictive measures. This study is the first step towards that: understanding the cognition on a group level. This is before being able to include it in future predictive studies (based on the cross-sectional data, this study by itself cannot assess the predictive validity of the measure).

      Strengths:

      (1) Large sample size.

      (2) Replication of their own findings.

      (3) Well-controlled task with measures of behaviour and mood + precise and well-validated computational modeling.

      Weaknesses:

      I can't really see any major weakness, but I have a few questions:

      (1) I can see from the parameter recovery that the parameters are very well identified. Is it surprising that this is the case, given how many parameters there are for 90 trials? Could the authors show cross-correlations? I.e., make a correlation matrix with all real parameters and all fitted parameters to show that not only the diagonal (i.e., same data is the scatter plots in S3) are high, but that the off-diagonals are low.

      Thank you for raising these thoughtful concerns. The current task consisted of 90 choices and 36 mood ratings. There were 5 choice parameters and 4 mood parameters. The apparently strong identifiability is not unexpected, as 90 choice trials and 36 mood ratings are comparable to those in prior computational modeling literature (Blain & Rutledge, 2022).

      As suggested, we computed cross-scorrelations between all generating (“true”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery. Please see Supplementary Pages 2-3.

      “Parameter recovery: Figure S3 shows good parameter recovery for both choice and mood winning model (choice: rs > 0.91, ps < 0.001; intraclass coefficients > 0.78; mood: rs > 0.90, ps < 0.001; intraclass coefficients > 0.86). Moreover, we computed cross-correlations between all generating (“true”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery.”

      Page 10:

      “The numbers of choice trials and mood ratings were comparable to those in prior computational modeling studies [34,35].”

      (2) Could the authors clarify the result in Figure 2B of a correlation between gambling rate and suicidal ideation score, is that a different result than they had before with the group main effect? I.e., is your analysis like this: gambling rate ~ suicide ideation + group assignment? (or a partial correlation)? I'm asking because BSI-C is also different between the groups. [same comment for later analyses, e.g. on approach parameter].

      Thank you for pointing out the lack of clarity. We performed group difference analysis and correlation of suicidal ideation analysis, separately. We first performed group difference analysis to test our hypothesis of STB effects. We then conducted correlational analysis to further specify our findings.

      (3) The authors correlate the impact of certain rewards on mood with the % gambling variable. Could there not be a more direct analysis by including mood directly in the choice model?

      Thank you for this insightful suggestion. As suggested, we tried to integrate mood into choice models by adding mood bias component(s) in line with previous literature (Vinckier et al., 2018). The first model (mcM1) assumes that mood biases choice, building on cM3 (the winning choice model). cmM2 further separated the mood bias parameter into two components according to participants’ choices.

      However, model comparison using BIC supported cM3 (Table S6), that is, without consideration of mood in choice modeling. This can be due to the lack of block design in our experimental design unlike e.g., Vinckier et al., (2018) and Eldar & Niv, (2015). Please see Supplementary Note 6.

      (4) In the large online sample, you split all participants into S+ and S-. I would have imagined that instead, you would do analyses that control for other clinical traits. Or, for example, you have in the S- group only participants who also have high depression scores, but low suicide items.

      Thank you for this insightful suggestion. Following prior suicide-related literature (Tsypes et al., 2024), we controlled for depression by including them as covariates. Note that depression scores were derived from our established bifactor model (Wang et al., 2025), which decomposed depression from the anxiety. These results remained largely significant (ps ≤ 0.050), except a marginally significant effect of group on gambling behavior (p = 0.059). Despite a trend, this effect with covariates of depression-related questionnaires is strong in our clinical cohort (p = 0.024; Table S8). This suggests that the link between suicidality and risky behavior persists above and beyond general depressive symptoms.

      Please see our clarifications below:

      Page 26:

      “After controlling for depression severity using our established bifactor model (see ref 60 for details), these results remained significant (ps ≤ 0.050), except a marginally significant effect of group on gambling behavior (p = 0.059). Despite a trend, this effect with covariates of depression-related questionnaires is strong in our clinical cohort (p = 0.024; Table S8). This suggests that the link between suicidality and risky behavior persists above and beyond general depressive symptoms.”

      Reviewer #3 (Public review):

      This manuscript investigates computational mechanisms underlying increased risk-taking behavior in adolescent patients with suicidal thoughts and behaviors. Using a well-established gambling task that incorporates momentary mood ratings and previously established computational modeling approaches, the authors identify particular aspects of choice behavior (which they term approach bias) and mood responsivity (to certain rewards) that differ as a function of suicidality. The authors replicate their findings on both clinical and large-scale non-clinical samples.

      (1) The main problem, however, is that the results do not seem to support a specific conclusion with regard to suicidality. The S+ and S- groups differ substantially in the severity of symptoms, as can be seen by all symptom questionnaires and the baseline and mean mood, where S- is closer to HC than it is to S+. The main analyses control for illness duration and medication but not for symptom severity. The supplementary analysis in Figure S11 is insufficient as it mistakes the absence of evidence (i.e., p > 0.05) for evidence of absence. Therefore, the results do not adequately deconfound suicidality from general symptom severity.

      Thank you for this important comment. Based on clinical interviews, we included patients with and without suicidality (S<sup>+</sup> and S<sup>-</sup> groups). However, in line with suicidal-related literature (e.g., Tsypes et al., 2024), two groups also differed substantially in the severity of symptoms (see Table 1). To address the request for evidence on specificity to suicidality beyond general symptom severity, we performed separate linear regressions to explain in gambling behaviour, value-insensitive approach parameter (β<sub>gain</sub>), and mood sensitivity to certain rewards (β<sub>CR</sub>) with group as a predictor (1 for S<sup>+</sup> group and 0 for S<sup>-</sup> group) and scores for anxiety and depression as covariates. Results remained significant after controlling anxiety and depression (ps < 0.027; Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) on the clinical questionnaire to extract the orthogonal components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. We then performed linear regressions using these components as covariates to control for anxiety and depression. Our main results remained significant (ps < 0.027; Table S9). We believe that these analyses provide evidence that the main effects on gambling and on mood were specific to suicide.

      As pointed out, these “absence of evidence” cannot provide insights of “evidence of absence”. Although we median-split patients by the scores of general symptoms (e.g., depression and anxiety-related questionnaires) and verified no significant differences in these severities (Figure S11), we additionally conducted Bayesian statistics in gambling behavior, value-insensitive approach parameter, and mood sensitivity to certain rewards. BF<sub>01</sub> is a Bayes factor comparing the null model (M<sub>0</sub>) to the alternative model (M<sub>1</sub>), where M<sub>0</sub> assumes no group difference. BF<sub>01</sub> > 1 indicates that evidence favors M<sub>0</sub>. As can be seen in Table S7, most results supported null hypothesis, suggesting that general symptoms of anxiety and depression overall did not influence our main results. Overall, we believe that these analyses provide compelling evidence for the specificity of the effect to suicide, above and beyond depression and anxiety.

      Please see Table S7, S8 &S9 and our revisions below.

      Page 17:

      “Within patients, this group effect on gambling rate remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.024; also see Figure S11, Table S7 and Table S8). Given high correlations among anxiety and depression questionnaires (rs > 0.753, ps < 0.001), we performed Principal Components Analysis (PCA) to extract main components, where each component explained 86.95%, 7.09%, 3.27%, and 2.68% variance, respectively. To further control for anxiety and depression, linear regression using these components as covariates revealed that the group effect on gambling rate remained significant (p = 0.024; Table S9).”

      Pages 18-19:

      “Within patients, this group effect on the approach parameter remained significant after controlling for sex, illness duration, family history, diagnosis, and various medications use (ps < 0.05), as well as general symptoms (e.g., depression and anxiety; p = 0.027; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on approach parameter remained significant (p = 0.027; Table S9).”

      Page 21:

      “Within patients, this group effect on βCR remained significant after controlling for gambling rate, earnings, mood-related outcome effect, mood drift effect, sex, illness duration, family history, diagnosis, and various medications use (ps < 0.032), as well as general symptoms (e.g., depression and anxiety; p = 0.001; also see Figure S11, Table S7 and Table S8). Linear regression using PCA components as covariates revealed that the group effect on this mood parameter remained significant (p = 0.001; Table S9).”

      (2) The second main issue is that the relationship between an increased approach bias and decreased mood response to CR is conceptually unclear. In this respect, it would be natural to test whether mood responses influence subsequent gambling choices. This could be done either within the model by having mood moderate the approach bias or outside the model using model-agnostic analyses.

      Thank you for this important suggestion. As suggested, one interesting question was whether mood responses influence subsequent gambling choices and how to model them. First, we median-split mood responses (except the final rating) to compare gambling rate. Results showed a trend for less gambling rate in higher mood (t = -1.971, p = 0.050). However, there was no significant group difference (F = 0.680, p = 0.507). Second, with the assumption that mood biases choice, we constructed mcM1 based on cM3 (the winning choice model). Based on our finding of the negative correlation between mood sensitivity to certain rewards and gambling rate in S<sup>+</sup>, we separated β<sub>Mood</sub> parameter into β<sub>Mood-CR</sub> and β<sub>Mood-GR</sub> (cmM2). Model comparison using BIC supported cM3 (Table S6), that is, without consideration of mood in choice modeling. This can be due to the lack of block design in our experimental design unlike e.g., Vinckier et al., (2018) and Eldar & Niv, (2015). Please see Supplementary Note 6.

      (3) Additionally, there is a conceptual inconsistency between the choice and mood findings that partly results from the analytic strategy. The approach bias is implemented in choice as a categorical value-independent effect, whereas the mood responses always scale linearly with the magnitude of outcomes. One way to make the models more conceptually related would be to include a categorical value-independent mood response to choosing to gamble/not to gamble.

      We apology for the unclear statement. The approach bias is implemented in choice as a continuous value-independent effect, ranging from -1 to 1.

      It was true that the mood responses always scale with the magnitude of outcomes, since mood ratings were request after the outcomes. Therefore, mood parameters and the approach bias were both continuous.

      We also attempted to integrate mood into choice modelling. See Response 2 for Reviewer 3 for details.

      (4) The manuscript requires editing to improve clarity and precision. The use of terms such as "mood" and "approach motivation" is often inaccurate or not sufficiently specific. There are also many grammatical errors throughout the text.

      Thank you for this important suggestion. We have now explained motivation and mood in the Introduction section and the computational modeling section. Please see our clarifications below:

      Pages 3-4:

      “A growing literature indeed shows that risky behavior can be far better explained after adding value-insensitive approach and avoidance components to prospect theory [18,19], that is by including a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. This class of models highlights the important role of value-insensitive motivational components in decision making in addition to risk attitude-driven valuation (e.g., loss/risk aversion) [20].”

      Page 5:

      “Although mood is thought to persist for hours, days, or even weeks [30–33], momentary mood, measured over the timescale in the laboratory setting, represents the accumulation of the impact of multiple events at the scale of minutes [30,32,34–38]. Momentary mood external validity is demonstrated e.g., through its association with depression symptoms [37]. Mood is different from emotions, which reflect immediate affective reactivity and is more transient (e.g., from surprise to fear) [31–33,39].”

      We have corrected grammatical errors throughout the manuscript.

      (5) Claims of clinical relevance should be toned down, given that the findings are based on noisy parameter estimates whose clinical utility for the treatment of an individual patient is doubtful at best.

      Thank you for this comment. We agree that we did not evaluate the noise in our estimate e.g., by assessing the test-retest reliability on the task parameters, which is outside the scope of the study, and it is indeed possible that parameter estimate is somehow noisy. Therefore, we tone down the clinical relevance of our results. Please see our revision below:

      Page 32:

      “Next, we did not evaluate the noise in our estimate e.g., by assessing the test-retest reliability on the task parameters and it is indeed possible that parameter estimate is somehow noisy.”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) Title: I believe "aberrant mood dynamics" is both too general and overstating the results of this study, which did not measure mood dynamics longitudinally. "Aberrant" is also overly pathologizing. I would suggest sticking more directly to the results, for instance, "Insensitivity of momentary mood to non-risky rewards in adolescent suicidal patients".

      Thank you for this suggestion. We have now corrected it.

      (2) Abstract: in line 61, "Our study uncovers the cognitive and affective mechanisms" suggests that these are the only ones, and you uncovered them. Of course, there could be more mechanisms contributing to risk behavior in STB, so I would suggest removing the word "the" or adding "one of the".

      Thank you for this suggestion. We have now corrected it.

      (3) One major weakness of this study is that suicidal thoughts and behaviors were not assessed via a clinical instrument such as the Columbia Suicide Severity Rating Scale - this should be mentioned upfront.

      Thank you for this comment. According to medical records and information from family and friends by the researcher and psychiatrists, patients with suicidal thoughts and behaviors were categorized as suicidal group (S<sup>+</sup>), while patients without suicidal thoughts and behaviors were identified as control group (S<sup>-</sup>). Note that medical records and information were recorded from clinical interviews where the psychiatrists were vigilant for signs of suicidal ideation and inquired about suicidal-related thoughts and behaviors from both the patients and their families. Therefore, the current group operation was possibly comparable to Columbia Suicide Severity Rating Scale.

      (4) Table 1: female/male are sex, not gender (gender is man/woman/transgender/non-binary).

      Thank you for this suggestion. We have now corrected it.

      (5) Equation 1: It would be good to clarify what happens in gain-only or loss-only trials (the other value is then 0, but this can be clarified as it is not technically a loss or a gain).

      Thank you for this suggestion. We have now corrected it. Please see below for our revision:

      Page 12:

      “Please note that V<sub>gain</sub> is 0 in gain trials and V<sub>loss</sub> is 0 in loss trials.”

      (6) Figure 1E: The model prediction is not informative here. Given the linear regression model, there is no other option except that the mean prediction would overlap with the mean empirical measurement (unless the model was specified incorrectly). The same is true in Figure 2A.

      Thank you for this suggestion. We have now removed plots for model prediction.

      (7) Figure 1G: There was no analysis of the differences between groups in terms of earnings, given that the ANOVA was not significant. Still, if the claim is that risky behavior is sometimes suboptimal in this task, it would be good to show that there is a correlation between, say, symptoms of STB across groups and 1) risky behavior and 2) earnings.

      Thank you for this insightful comment. In the patient cohort, risky behavior (gambling rate)—but not earnings predicted the current suicidal ideation score (BSI-C, β = 9.189, t = 2.004, p = 0.048; earnings, β = 0.001, t = 0.582, p = 0.562). The lack of association for earnings is consistent with the task design, in which there is no stable optimal policy and payouts are only a coarse proxy for decision quality. Future work in learning paradigms, where optimality is well defined, may be better suited to test earning-based links to STB. We have clarified this point below:

      Page 32:

      “Second, although we assumed that increased risky behavior in STB was suboptimal, the current task was not suited to test this, given the task design of random feedback for gambling option. Future work in learning paradigms, where optimality is well defined, may be better suited to test earnings-based links to STB.”

      (8) Line 290: "beta_gain: -1-1" is unclear. I believe you meant beta_gain \in [-1,1].

      Thank you for this suggestion. We have now corrected it to make it clear.

      (9) The gain and loss biases are modeled as minimum and maximum probabilities for choosing the gamble. This is a legitimate choice for value-agnostic biases, but it is not the traditional choice (as far as I know). I wonder if the same results would hold with the more traditional formulation of the bias as an added constant to the utility of the gamble, i.e., p(gamble) = 1/(1+ exp(-mu(U_gamble + beta_gain - U_certain)). I believe in this case, you would also not have to specify different equations for positive or negative biases, or to limit the bias to the range of [-1,1] (indeed, the bias would be in reward-equivalent units).

      Thank you for this suggestion. The winning choice model we used here was consistent with previous literature (Rutledge et al., 2015 & 2016), which decomposed the decision process into risk-attitude-driven valuation (e.g., loss and risk aversion) and value-insensitive motivational components. These approach/avoidance parameters are a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference.

      As suggested, we also compared the traditional bias choice model. Model comparison did not support this. Please see Supplementary Page 4.

      (10) Also, for equations 5-8, it seems that 5-6 are identical to 7-8 except for the use of beta_gain versus beta_loss. You might want to consider simplifying by putting beta in the equations and specifying in the text that, depending on the trial type (loss or gain), the relevant beta is used.

      Thank you for this suggestion. We have now simplified it. Please see our revision below:

      (11) It is not clear what equations are applied to mixed trials in cM3.

      Sorry for the confusion. We have now clarified this point.

      Page 12:

      “Approach/avoidance parameters are not applied to in mixed trials.”

      (12) Model comparison: the mood models are nested within each other (e.g., mM3 can be derived from mM1 by setting beta_EV = beta_RPE). In this case, model comparison can use the likelihood ratio test instead of BIC, which can be too conservative (and therefore does not support the extra beta parameter for RPE, different from previous results in the literature). I wonder if a likelihood ratio test would lead to results more in line with previous findings with this task?

      Thanks for this suggestion. We agree that mM1 (CR+EV+RPE) and mM3 (CR+GR) are nested. However, our model space also included unnested models, such as mM5 (CR+GR<sub>better</sub>+GR<sub>worse</sub>). Therefore, it was not reasonable in our model space to use likelihood ratio tests.

      (13) Line 346: The replication sample is described as "healthy participants," however, their health (or mental health) status was not assessed, and they may as well have mental health concerns. I would suggest calling this a general sample or an undifferentiated sample - but not a healthy sample.

      Sorry for the confusion. We have now corrected this phrase.

      (14) Line 363: "in addition to the replication of previous findings in the validation dataset" is unclear. Are those tests not two-tailed?

      Sorry for the unclear statement. In the replication analyses, we used one-tailed t-tests because the direction of the effect was revealed on the clinical dataset. Please see our clarification below:

      Page 15:

      “For the replication of previous findings in the validation dataset, we used one-tailed tests in line with our clinically motivated directional hypothesis.”

      (15) Line 372: "validating our group manipulation" - the presented work does not have a manipulation. Maybe you meant "validating our grouping of participants"?

      Thank you for this suggestion. We have now corrected it to make it clear.

      (16) Figure 2B: It is not clear how the data were binned for illustration purposes only, and why this binning is necessary (I have not seen it in other papers) - presenting the data from each subject and the correlation line with error margins (as is done here) should be sufficient.

      Thank you for flagging this. For illustration only, we binned the data proportional to group sizes: in the patient sample (S<sup>-</sup> n = 25; S<sup>+</sup> n = 58; ≈1:2), we displayed 3 bins for S<sup>-</sup> and 6 bins for S<sup>+</sup>. We agree that binning is not necessary; all statistics were computed on raw, unbinned data. The binned panel was included solely for visualization, consistent with our prior work (Blain et al., 2023).

      (17) Table 2: delta BIC should be presented per subject (that is, divided by the number of subjects in each group), as the groups are of different sizes, so as presented now, the columns are not comparable across groups.

      Thank you for the helpful suggestion. Our goal in Table 2 is not to compare ΔBIC magnitudes across groups, but to identify the winning model within each group. The ΔBICs are aggregated at the group level solely to rank models for that group. Dividing by the number of participants would rescale each group’s column by a constant and would therefore not affect the within-group ranking or the conclusion that cM3 is the best model in all groups. For this reason, we retain the current presentation and interpret each column within group rather than across groups.

      (18) Line 640 - the effect of expectations and prediction errors on mood was not only shown in healthy people, but also in people with depression (Rutledge et al., 2007, https://pubmed.ncbi.nlm.nih.gov/28678984/)

      Thank you for this comment. Indeed, Rutledge et al., (2017) showed evidence for CR+EV+RPE mood model in adult people with depression. However, our study recruited adolescents with depression or anxiety, given that adolescent period might provide a developmental window for opportunities for early intervention of suicidality. Therefore, it is also possible that the current winning model was specific to adolescents. Please see our clarifications below:

      Page 28:

      “It is also possible that the current winning model was specific to adolescents. Given that Rutledge et al., (2017) supported the “CR-EV-RPE model” in adults with depression, our study with adolescent populations may suggest a developmental change for mood sensitivities.”

      (19) Supplemental material: Is the R2 section about R-squared? Perhaps you can use superscript on the 2 to make that clearer? For Figure S2, how was model recovery determined? Should I interpret the confusion matrix as suggesting that the winning model for each and every simulated subject was the generating model, or was the winning model determined for the whole simulated population in each of the 100 simulations? Traditionally, confusion matrices use the former measure, but the results of 100% recoverability make me suspect the latter was used here. In Figure S3, should we not be looking at simulated parameters and recovered parameters? What are "real parameters" here?

      Thank you for these important comments. We now consistently denote the coefficient of determination as R<sup>2</sup> (with a superscript 2) throughout the manuscript and Supplementary Materials.

      For the model recovery analysis in Figure S2, we have clarified that the confusion matrix is computed at the population level. Specifically, for each of the 100 simulations we generated a full dataset under each candidate model, fit all models to that dataset, and selected the winning model based on group-level model evidence (BIC). Each cell in the confusion matrix therefore reflects the proportion of simulations in which model j was selected as the best-fitting model when the data were generated by model i. This operation was reasonable because the decision of the winning model is made on the population-level dataset rather than on individual subjects.

      In Figure S3, the term “real parameters” referred to the parameters used to generate the simulated data. To avoid confusion, we now relabel these as “simulated (generating) parameters” and explicitly describe the figure as showing the relationship between simulated (generating) parameters and recovered parameters. Please see Supplementary Pages 2-3:

      “Model recovery: We generated 100 simulated datasets for each model (3 choice models and 8 mood models) using the fitted parameters of each model as the ground truth. Each dataset contained 201 trials and included 3 (or 8) sets of simulated data corresponding to the respective models. For each simulated dataset, we then fit all models and determined the winning model at the population level based on group-level BIC, yielding a confusion matrix in which each entry represents the proportion of simulations in which model j was selected as the best-fitting model when the data were generated by model i. As shown in Figure S2, all models are highly identifiable, indicating excellent recovery performance for both the choice and mood models.”

      “Parameter recovery: Figure S3 shows good parameter recovery for both choice and mood winning model (choice: rs > 0.91, ps < 0.001; intraclass coefficients > 0.78; mood: rs > 0.90, ps < 0.001; intraclass coefficients > 0.86). Moreover, we computed cross-correlations between all generating (“generating”) and recovered (“fitted”) parameters. The resulting matrix showed high diagonal (choice winning model: rs > 0.91; mood winning model: rs > 0.90) and low off-diagonal (choice winning model: abs(rs) < 0.63; mood winning model: abs(rs) > 0.40) correlations, further supporting parameter recovery.”

      Typos:

      (1) Line 90: original → originate

      (2) Line 596-598 - the same phrase is repeated twice.

      (3) Line 616: on the other word → hand.

      Sorry for the mistakes. We have now corrected them throughout the manuscript.

      Reviewer #2 (Recommendations for the authors):

      For people unfamiliar with interpersonal theory or motivational-volitional model, or three-step theory (lines 105-106), could you briefly explain the key idea of mood and suicide before going to the decision-making tasks? And from this, maybe motivate the predictions in your task? In particular, in the abstract and introduction, the phrasing could be a bit more concise and simpler. In the abstract, sentences were sometimes quite long. In the introduction, some paragraphs are somewhat repetitive. In the discussion, there were some typos.

      Thank you for these suggestions. We have now explained the key idea of mood and suicide before going to the decision-making tasks in the introduction, which can be seen below:

      Pages 4-5:

      “Contemporary theories of suicide converge on the idea that STB is initially caused by low mood experience. The interpersonal theory of suicide proposes that suicidal desire arises when people simultaneously feel socially disconnected (“thwarted belongingness”) and like a burden on others (“perceived burdensomeness”), experiences that are tightly linked to chronically low mood [25]. The motivational–volitional model [26] and the three-step theory [27,28] similarly emphasize that when negative mood and feelings of defeat or entrapment are experienced as inescapable, they can give rise to suicidal ideation, and that the progression from ideation to suicide attempts depends on additional factors such as reduced fear of death, increased pain tolerance, and a tendency to act impulsively under intense affect. Some official organizations, e.g., National Institute of Mental Health, have also listed mood problems as warning signals [8]. Interestingly, within the framework of decision making under uncertainty, gambling on lotteries with a revealed outcome has been found to induce high mood variance [29], providing an opportunity to assess the relationship between deficient mood and increased gambling decisions in STB.”

      We have also refined the wording and corrected typos throughout the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) Since many readers might only read the abstract, it is important that it is both informative and accurate. I have two suggestions in this respect. First, for the abstract to be more informative, it may be helpful to indicate already there that these are value-insensitive approach-avoidance parameters, in the sense that they favor/disfavor the gamble regardless of the potential outcomes' magnitude or probability. This issue is also present throughout the text, where the phrases "approach and avoidance motivation" are referred to as if they have established and precise computational definitions. In my view, these terms could just as easily be interpreted as parameters that multiply the value of potential gains or losses, which is not what the authors mean. It would be helpful to clarify this terminology.

      Thank you for these suggestions. In line with previous literature (Rutledge et al., 2015 & 2016), approach and avoidance motivation are indeed defined at the computational level, referring to a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. We have cited these papers in the manuscript. We also make it clear to further clarify approach and avoidance parameters in the abstract and introduction. Please see our revisions below:

      Page 2 (Abstract):

      “Using a prospect theory model enhanced with value-insensitive approach-avoidance parameters revealed that this rise in risky behavior resulted only from a heightened approach parameter in S<sup>+</sup>.”

      “Altogether, model-based choice data analysis indicated dysfunction in the approach system in S<sup>+</sup>, leading to greater propensity for gambling in the gain domain regardless of the lottery expected value.”

      Page 3 (Introduction):

      “A growing literature indeed shows that risky behavior can be far better explained after adding value-insensitive approach and avoidance components to prospect theory [18,19], that is by including a decision bias in favor of the highest gain (approach) and another decision bias against the lowest loss (avoidance), above and beyond options value difference. This class of models highlights the important role of value-insensitive motivational components in decision making in addition to risk attitude-driven valuation (e.g., loss/risk aversion) [20].”

      (2) The statement "our study uncovers the cognitive and affective mechanisms contributing to increased risk behavior in STB" is overstating the findings, as the study may have uncovered some contributing mechanisms, but likely not all of them. Removing the word "the" would fix this issue.

      Thank you for this suggestion. We have now corrected it.

      (3) Since mood is typically defined as lasting hours, it's inappropriate to refer to ratings that only reflect the last few trials as self-reports of mood. To be sure, I view the distinction between emotions and moods as quantitative, not qualitative, so I do not think there is a problem studying the former to understand the latter, but to avoid confusion, the terminology should follow common usage.

      Thank you for this suggestion. We follow previous work and operational definitions regarding mood (Rutledge et al., 2014, Eldar & Niv, 2015, Vinckier et al., 2018). Emotion is usually a very brief response to a specific stimulus (Emanuel & Eldar, 2023), e.g., leading to rapid changes like surprise then fear. In contrast, mood is defined as a diffuse state that is not specific to one stimulus. Here, we operationally and computationally define mood as an affective state reflecting the recent history of safe and gamble outcomes. We now clarify that point in the main text. Please see our revision below:

      Page 5:

      “Although mood is thought to persist for hours, days, or even weeks [30–33], momentary mood, measured over the timescale in the laboratory setting, represents the accumulation of the impact of multiple events at the scale of minutes [30,32,34–38]. Momentary mood external validity is demonstrated e.g., through its association with depression symptoms [37]. Mood is different from emotions, which reflect immediate affective reactivity and is more transient (e.g. from surprise to fear) [31–33,39].”

      (4) Line 78: The phrases "increase in risk attitude", "decrease in loss attitude", and "decrease in value-independent choice biases" are unclear to me in terms of their directionality. An attitude might be avoidant or embracing. If it is the former then increasing it would decrease risk-taking.

      Thank you for pointing out the ambiguity. We have now corrected them throughout the manuscript. Please see our revision below:

      Page 4:

      “We therefore hypothesized that heightened approach motivation, or weakened avoidance motivation, would account for increased risk behavior in STB.”

      (5) Line 125: I was not sure why one would expect the mood response to gamble-related quantities (EV and RPE) to be lower in STB and not higher.

      Sorry for the typo. We hypothesized that mood would respond more strongly to gambling-related quantities expected value (EV) and reward prediction error (RPE)—in adolescents with STB than in controls, given prior evidence that STB is associated with greater risk-taking.

      (6) The text could use proofreading, as there are many typos. These are from the first 100 lines alone:

      (a) Abstract: regardless the lotteries -> regardless of the lotteries'.

      (b) Line 78: it remains whether.

      (c) Line 80: can each -> each can.

      (d) Line 90: may original from.

      Sorry for the mistakes. We have now corrected them throughout the manuscript.

      (7) The rationale for focusing on the S+ group for mood model comparison is incorrect. The purpose is to identify parameters that vary as a function of suicidality, and for that, the S- group is just as important.

      Thank you for this comment. We agree that the S<sup>-</sup> group is as important as the S<sup>+</sup> group. A direct comparison was complicated because the winning mood models differed (S<sup>+</sup>: mM3; S<sup>-</sup>: mM5; Table 3). To ensure comparability, we checked results from both model specifications (mM3 and mM5). The conclusions were convergent: mood sensitivity to certain rewards (CR) was lower in S<sup>+</sup> than in S<sup>-</sup> (see Fig. 3 for mM3 and Fig. S8 for mM5).

      (8) There appears to be a contradiction between the inclusion criteria, which include having experienced suicidal thoughts and behaviors, and the definition of the S- group as not having suicidality.

      Thank you for pointing out this mistake. The corrected version of inclusion criteria can be seen on Page 7:

      “Patients were included if they met the following criteria: 1) both the researcher and psychiatrists agreed on their group classification; 2) they had a current diagnosis of major depressive disorder (MDD; unipolar depression), generalized anxiety disorder (GAD), or bipolar disorder with depressive episodes (BD), confirmed by two experienced psychiatrists using the Structured Clinical Interview for DSM-IV-TR-Patient Edition (SCID-P, 2/2001 revision; see Supplementary Note 1 for details);3) they were between 10 and 19 years of age; 4) they had no organic brain disorders, intellectual disability, or head trauma; 5) they had no history of substance abuse; 6) they had no experience of electroconvulsive therapy.”

      (9) It would be helpful to specify whether mood modeling was based on objective or subjective values, and why.

      Thank you for this helpful suggestion. We have now clarified whether mood modeling was based on objective or subjective values, and why. Specifically, we constructed two model families: one in which mood was driven by objective monetary outcomes (objective values) and one in which mood was driven by subjective values derived from each participant’s fitted choice model (subjective values). We then used the VBA_groupBMC function in the VBA toolbox to perform family-wise model comparison, with 8 candidate mood models within each family. Consistent with previous literature, the objective-value family provided a clearly superior fit to the data (exceedance probability, EP = 1.000). Based on this result and for parsimony, we report and interpret the mood modeling results from the objective-value family in the main text. We have clarified this point in Supplementary Note 9.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study presents an interesting approach for finding electrophysiological models that match experimental patch-clamp data. The authors develop a new method for deriving optimized current clamp protocols by training a neural network on synthetic data. This optimized current clamp is then used on both computational training data and on experimental data to predict current gating and conductance parameters that correctly reconstruct the electrical phenotype.

      Strengths:

      (1) The fitting of gating variables through an optimized patch clamp protocol is interesting.

      (2) The inclusion of experimental data is important, and the approach is shown to be effective in fitting them.

      Weaknesses:

      (1) Some clarity is necessary on the generation and selection of variable IPSC models. With such a large variation in so many parameters, I would expect some resulting parameters to generate non-realistic phenotypes, quiescent cells, etc. Are all 200,000 or 1,100,000 generated cells viable? Or are they selected somehow for realistic cell properties?

      Thank you for this important point. We agree that broad parameter variation can generate non-physiological model behavior. Indeed, with the +/-40% perturbation range, some simulated cells produced non-realistic outputs, including quiescent behavior, and failure to generate a complete action potential. These cases were excluded from the dataset. As a result, only cells exhibiting physiologically meaningful and numerically stable behavior were retained for further analysis. We have clarified this selection procedure in the Methods section. We applied a large variation to ensure that all possible combinations and morphologies were included in the training and testing data so the model would readily ingest new data and perform robustly.

      (2) The error shown in Figure 4 between different population sizes is not completely explained in the text - there seems to be a minimal difference between a population of 1,000 and 10,000, followed by a very good fit at 200,000. Is there a particular threshold that needs to be crossed where the error drops off? Related, how was the 200,000 number chosen?

      Thank you for this observation. We agree that the decrease in error shows a gradual performance improvement as the population size increases, rather than a strict cutoff. As shown in Figure 4, the difference between 1,000 and 10,000 samples is small, but as we continue to increase and get to around 200,000 samples, we see strong error minimization. This indicates how much training data is needed for optimal model performance. This improvement is due to better coverage of the high-dimensional parameter space, which helps the network learn the nonlinear relationships between the parameters and outputs.

      We tested a range of training data sets and found that above 200,000 training data sets, the model consistently produced low, stable errors and good test-training agreement. The test error decreased with the training error as the population size increased, indicating better generalization and suggesting that the model accurately predicts unseen data rather than overfitting to the training set.

      (3) Related to the point above, the 1,100,000 population for fitting experimental data also needs a more complete explanation: how was this number chosen, and how does the error compare with the other population sizes shown in Figure 4?

      Thank you for this question. We found that at a training data set size of 1,100,000 we were able to cover the large parameter space induced by +/-40% parameter perturbation. iPSC-CM measurements are known to exhibit high variability, and we wanted to capture the full range in the training data set so the model could ingest a wide range of experimental data. It is trivial to generate new training data, for example, to capture different experimental conditions like temperature differences, mutations, drugs, or ionic variability. We view this flexibility as a substantial strength of the approach. But the large perturbations we show in this study (+/-40%) allow the generation of a very broad range of cellular phenotypes while maintaining physiologically realistic ionic current properties and action potential behavior. Consistent with Figure 4, increasing population size reduces prediction error and improves generalization. The larger dataset provided more stable, accurate predictions when fitting experimental data, without evidence of overfitting.

      (4) Why are the optimized current clamp protocols different between panels A and B in Figure 5? Are they somehow informed by experimental data?

      Thank you for this question. The stimulation protocol used in panels A and B is identical. Panels A and B show whole-cell currents recorded under the same stimulation conditions as in Figure 3. The differences reflect variability in the underlying whole-cell ionic currents of the model cells rather than differences in the applied protocol. This is exactly the idea: the exact same protocol will generate different whole-cell currents in individual cells, but the model can find parameter sets for all of them.

      (5) Figure 6D: Is the EAD risk in panel D specific to cell 1, 2, or the pooled variants of both?

      Thank you for this question. We have clarified this point in the revised manuscript. The EAD risk shown in panel D is computed from the pooled variants of both Cell 1 and Cell 2, rather than being specific to either cell individually.

      (6) How sensitive is the fitting to minor parameter variation? Further, if one were to pick, let's say, the next-best-fitting value, would that fall close to the best one? Is the solution found unique, or are there multiple sets with good fits?

      Traditional optimization methods, such as Nelder–Mead, directly fit the model to the observed data by iteratively minimizing the error for each dataset. As a result, the solution can depend on the initial parameter guess and may converge to different local minima. In contrast, our approach trains a deep learning model on synthetic data generated from the baseline model, learning a mapping from whole-cell currents to the corresponding 52-parameter sets by minimizing prediction error. The mean squared error (MSE) decreases from approximately 10⁻² to below 10⁻³, with training and test errors overlapping closely, indicating stable training, good generalization, and accurate reproduction of the observed signals.

      The model achieves very low MSE and reproduces the electrophysiological outputs with high fidelity. However, accurate reproduction of the outputs does not imply a unique parameter solution. This is illustrated in Figure S1, where baseline and predicted parameter values show close agreement overall, yet small deviations persist across parameters. This indicates that different parameter combinations can yield similar whole-cell behaviors due to parameter correlations and compensatory effects. In such cases, the model learns to predict a representative parameter set that is most consistent with the training data and loss function, rather than converging to a single unique solution within a fixed numerical tolerance.

      Reviewer #2 (Public review):

      Summary:

      The authors present a computational framework for generating "cell-specific" digital twins of human iPSC-CMs from a single optimized voltage clamp recording. Using deep learning trained on > 1 million artificial cells, the authors demonstrate that the model can infer 52 biophysical parameters governing 6 major ionic currents, and the resulting digital twins can reproduce experimentally recorded action potentials.

      Strengths:

      The framework has clear potential for understanding cellular heterogeneity in iPSC-CMs, predicting individual drug responses, and reducing the experimental burden of multiple patch clamp protocols.

      Weaknesses:

      There are several concerns about the validation of the model and its clarity. First, the biological variability being modeled in this manuscript is not defined well. It is unclear whether the framework addresses cell-to-cell differences within a single differentiation batch, variability across iPSC lines, or donor-to-donor differences. This ambiguity makes it difficult to interpret what the "digital twin populations" actually represent biologically. Second, the main claim, "the digital twins enable drug testing and arrhythmia prediction that would be impractical experimentally", is not experimentally validated. For example, the E-4031 simulations predict EAD rates, but no direct experimental head-to-head comparison is provided to confirm that these predictions are accurate. Third, technical reproducibility and biological representativeness are not assessed. Single voltage clamp recordings are inherently noisy. Without knowing how much variability comes from the recording process (technical variation) vs true biological differences, it is difficult to judge whether observed "cell-specific" parameter differences are meaningful. In addition, the optimized protocol is claimed to be superior to conventional approaches, but again, no experimental comparison is shown.

      The authors should address these concerns, with particular emphasis on clarifying the biological context and providing direct experimental validation. Below are detailed specific points:

      (1) Ambiguous definition of iPSC-CM heterogeneity. The authors model "typical iPSC-CM heterogeneity" by varying 52 parameters +/- 40% around a baseline model (Figure 1), generating > 1 million synthetic cells. However, the manuscript does not clearly state what biological variability this model is intended to capture. Is this modeling within-line, cell-to-cell variability (e.g., cells from the same dish or differentiation batch that differ due to stochastic gene expression or maturation state)? Or is this modeling between-line or between-donor variability (e.g., genetic background differences, reprogramming efficiency)? This distinction is critical for interpretation. If the goal is to understand why different cells in the same dish behave differently, then training data should reflect that. If the goal is to compare patient lines or disease models, the framework needs validation across multiple donors or lines.

      For example, the experimental validation in Figure 5 uses a single iPSC line (iPS-6-9-9T.B), but how many differentiation batches or dishes were tested, or whether cells came from the same preparation are unclear. Another example is that the wide AP diversity in the training population (Figure 1A) is impressive, but there is no demonstration that real experimental cells actually fall within this assumption range of +/- 40%.

      From a biological perspective, iPSC-CMs are known to be highly heterogeneous within lines (maturation state, metabolic differences, epigenetic variation, spatial differences within the same dish, etc) and between lines (different donor/genetic background). Thus, please explicitly state whether the +/- 40% variation is intended to model within-line or between-line heterogeneity, and justify this choice with wet experiment data (or reference to experimental literature on iPSC-CM variability). Please clarify how many dishes, differentiation batches, and time points post-differentiation were used for experimental recordings (Figures 5-6). If the framework is intended to generalize across lines from different donors, please test the model on multiple independent iPSC lines (from different donors).

      Thank you for this important and insightful comment. The selected ±40% range was chosen to broadly explore all physiologically plausible electrophysiological behaviors, not to match a specific experimental distribution. Our goal was to cover enough behaviors for the model to learn a reliable mapping between responses and ionic parameters.

      We recognize that this approach does not explicitly account for variability between lines or donors. We have a current project focused on extending the framework to include multiple iPSC-CMs from patient donors, but given that the model framework successfully reproduces such a broad range of cell phenotypes, we feel confident that it will readily apply to different genetic backgrounds from patient-specific cells. This study is underway.

      We have updated the manuscript to clarify how the modeled variability is interpreted and added a discussion of these limitations. Furthermore, we clarified the experimental conditions, such as the number of differentiation batches and recording settings, in the revised Methods section.

      (2) Biological representativeness of single-cell measurements.

      The framework generates digital twins from single voltage clamp recordings. The patch clamp recordings in iPSC-CMs are subject to substantial technical variability. The manuscript does not address a fundamental question: "How representative are the measurements from a single cell on the dish (or line)?" In other words, if I measure one cell from a dish of a million cells, does that cell's digital twin tell me something about the dish as a whole, or just about that one cell? The manuscript presents Cell 1 and Cell 2 (Figures 5-6) as distinct individuals, but it's unclear whether these differences reflect true biological heterogeneity or simply sampling variability. I think the authors should perform replicate recordings on multiple cells (e.g., > 10 cells) from the same dish (same differentiation batch) and quantify how much the inferred parameters vary, and then compare between lines.

      Thank you for this important comment. We agree that the representativeness of single-cell measurements and the impact of technical variability are important considerations in interpreting the results. In this study, the framework is designed to generate digital twins that reflect the electrophysiological properties of individual recorded cells, rather than to directly represent the behavior of the entire cell population within a dish.

      As such, differences observed between Cell 1 and Cell 2 are intended to reflect variability at the single-cell level, which may arise from a combination of biological heterogeneity and experimental variability. We agree that systematic replicate recordings across multiple cells are valuable to quantify the relative contributions of biological and technical variability, and to assess the consistency of inferred parameters. However, this is beyond the scope of the current study. We have added clarification in the manuscript to explicitly state this limitation and to outline this as an important direction for future work.

      (3) No experimental validation of the main claim that in silico populations can replace wet experiments.

      The most exciting claim in the manuscript is that digital twins enable drug testing and arrhythmia prediction "at scale" without requiring hundreds of patch clamp experiments. Specifically, the authors show that in silico populations derived from two experimental cells (Figure 6C) predict dose-dependent EAD incidence for the IKr blocker E-4031 (Figure 6D), with ~3% of cells showing EADs at 50 nM.

      However, this prediction is not validated experimentally. If I actually patch 20-30 real iPSC-CMs and apply 50 nM E-4031, will ~3% of them show EADs, as the model predicts? Without this validation, I think the drug testing framework is purely hypothetical. The model may be internally consistent (e.g., Cell 1's twin behaves differently from Cell 2's twin), but there is no evidence that these in silico populations reflect real biological variability in drug response. Please provide experimental validation that justifies the prediction by digital twins.

      Thank you for this important comment. We agree that experimental validation of population-level drug response will be valuable for establishing the quantitative accuracy of the predicted EAD incidence. The E-4031 simulations are intended as a proof-of-concept illustrating how the framework can identify susceptible subpopulations and quantify relative proarrhythmic risk in silico. We agree that direct comparison with large-scale experimental datasets is a key next step, and we are working hard to get the study funded so that we can perform those experiments and bring this technology to scale.

      (4) Experimental validation and head-to-head comparison of optimized protocol.

      The authors claim that their deep learning-optimized voltage clamp protocol (Figure 3, Figure 4A) is superior to conventional approaches, but they have not validated this experimentally by doing a head-to-head comparison. The manuscript does not compare the optimized protocol to any published voltage clamp designs. If the optimized protocol is genuinely easier to implement and more informative than existing approaches, this would be a major practical advance. But without side-by-side comparison, it is impossible to judge whether the optimization made a real difference.

      Thank you for your comment. We agree that comparing directly with traditional voltage-clamp protocols through experiments would be useful. In this study, our main aim was to show that the optimized protocol enhances parameter inference within the modeling framework, not to prove experimental superiority. We have clarified this point in the revised version.

      Reviewer #3 (Public review):

      Summary:

      This work uses a convolutional neural network to optimize a voltage clamp protocol to identify features and parameters from human pluripotent stem cell-derived cardiomyocytes.

      Yang et al. introduce an innovative experimental framework that integrates computational modeling and deep learning to generate a digital twin of human pluripotent stem cell-derived cardiomyocytes (hPSC-CMs).

      Strengths:

      The major strength is the methodology used to bridge in silico prediction of cell behavior and mechanistic insights from the experimental dataset.

      The approach used in this study represents a significant step toward precision medicine by enabling in silico prediction of cellular behavior and mechanistic insight from experimental datasets. The study addresses an important and timely challenge in stem cell-based and personalized medicine, and the authors compellingly leverage state-of-the-art methods alongside strong expertise in computational modeling and cardiac electrophysiology

      Weaknesses:

      While the overall approach is highly compelling and the potential impact is substantial, there are two areas where clarification and refinement, particularly in the phrasing and framing used throughout the manuscript, would further strengthen the work.

      (1) While the overall goal of the study is compelling, the manuscript would benefit from clearer articulation of how the proposed framework is intended to be used in practice. In particular, it is not entirely clear whether the authors envision this approach as:

      (a) a method to extract population-level trends that, when paired with biological data, enhance statistical power and interpretability, or

      (b) a strategy capable of constructing a population-based model from limited single-cell recordings. If the latter is intended, additional guidance on the number of action potentials required per cell and the assumptions underlying this extrapolation would greatly clarify the scope and applicability of the method.

      Thank you for this thoughtful comment. We agree that the intended use of the framework should be more clearly articulated. In this study, we generate a large synthetic population of iPSC-CM models by varying 52 biophysical parameters governing key ionic currents. A neural network is trained on simulated whole-cell current responses to learn a mapping between current profiles and model parameters. Experimental recordings are then used as inputs to this trained model to infer ionic parameters, rather than directly fitting the model to data. This enables individual recordings to be interpreted within a large, physiologically plausible parameter space and supports population-level analysis of electrophysiological variability. The primary goal of the framework is therefore to facilitate mechanistic interpretation of variability and relate experimental observations to underlying ionic currents. But the longer-term intended goal is to develop digital twins from patient-derived cell lines and then use populations constructed from patient-specific digital twins to screen therapeutics and identify arrhythmia marker vulnerability in a very thorough and high-throughput way. We have clarified this in the revised manuscript.

      (2) The manuscript would also benefit from a clearer explanation of how electrophysiological heterogeneity observed in hPSC-CMs is linked to inter-patient variability. Although the authors state that this framework can be generalized to compare patient-specific hiPSC-CM lines, it remains unclear how this generalization is achieved, given the substantial sources of variability intrinsic to hiPSC-CMs (e.g., batch effects, reprogramming strategy, differentiation protocol, and maturation state). As acknowledged by the authors, addressing this level of variability likely requires large datasets; further clarification of how the proposed approach mitigates or accommodates these challenges would strengthen the translational claims.

      Below are my suggestions that could help strengthen the claims in the manuscript:

      (1) Adding a dedicated section describing the electrophysiological phenotype of the hPSC-CMs used in this study would help justify the choice of the underlying ionic model and the selection of the six ion currents analyzed. These currents are not only developmentally regulated but may also vary substantially across different hPSC-CM lines, which has implications for generalizability.

      Thank you for this important suggestion. We agree that providing additional context on the electrophysiological phenotype of the hPSC-CMs strengthens the rationale for both the underlying ionic model and the selection of currents analyzed.

      We have expanded the Methods section to clarify this point. Briefly, the ionic currents were selected based on the Kernik-Clancy iPSC-CM model developed in our prior work, which was specifically designed to capture the range of electrophysiological variability observed within an iPSC-CM cell line using a population-based framework. In this model, variation in key ionic conductances is sufficient to reproduce the diversity of action potential morphologies, spontaneous activity, and repolarization dynamics commonly reported experimentally, while avoiding non-physiological behaviors.

      Accordingly, we focused on six primary ionic currents that are known to play dominant roles in shaping action potential characteristics and variability in iPSC-CMs. This selection reflects a balance between model parsimony and physiological relevance, enabling the framework to capture the expected spectrum of variability within a given cell line. We also note that the framework is extensible, and additional currents or alternative parameterizations can be incorporated to account for differences across cell lines, donors, or experimental conditions in future studies. See updated discussion.

      (2) If feasible, inclusion of patch-clamp data from an additional hPSC-CM line would significantly strengthen the claim that this framework can harmonize and generalize across datasets and cell sources.

      Thank you for this helpful suggestion. We agree that adding data from more hPSC-CM lines would improve the framework's generalizability. In this work, our goal was to show that the digital twin framework is data-driven and can easily be expanded to include more hPSC-CM lines, allowing for cross-line comparisons in future studies. We have clarified this and included a discussion of this limitation in the revised manuscript. We are currently seeking funding for patient-specific lines as well to allow scalability.

      (3) The authors note that the experimental cells exhibited high variability in action potential morphology. This is an important observation that directly supports the motivation for the study and should be explicitly presented, even if only in the supplementary materials.

      Thank you for this suggestion. We agree that explicitly showing the variability in experimental action potential morphology strengthens the motivation for this study. We have now added a section in the discussion discussing this and referencing the many prior studies that focused on iPSC-CM variability, including the studies upon which our initial model (Kernik-Clancy) was based.

      (4) In the hERG-blocker experiments, further clarification is needed regarding the biological relevance of the reported 3% incidence of early after depolarizations (EADs). Additionally, an interrupted sentence in this section makes it unclear whether the goal is to demonstrate that the digital twin can capture rare arrhythmic risk events or whether the digital twin is necessary to determine whether this level of risk is clinically meaningful.

      Thank you for this important comment. We agree that more clarification is needed on the ~3% EAD incidence and the digital-twin role. This analysis aims to show that electrophysiological variability can create a small, susceptible subpopulation under drug effects, not to set a clinical risk threshold. The observed ~3% EAD incidence reflects the emergence of such a susceptible subpopulation under hERG block. While relatively small, this fraction is important because it arises from modest, physiologically plausible variation in ionic properties and would be difficult to capture using single-cell or small-sample approaches. As described in the Discussion, this variability-driven emergence of EADs provides a quantitative measure of proarrhythmic risk at the population level. The digital-twin framework enables systematic identification and quantification of these rare events, linking cell-level variability to population-level responses. We have revised the manuscript to clarify this point.

      (5) The manuscript states that some action potentials were excluded from the experimental dataset. A brief explanation of the exclusion criteria, along with guidance on how to distinguish high-quality from low-quality recordings, would improve transparency and reproducibility.

      Thank you for this comment. We agree that the definition of failed recordings should be clarified. We have now specified the exclusion criteria in the Methods section.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) It would be helpful if the network cartoon in Figures 2 and 3 were replaced with a simplified sketch of the actual neural network used.

      Thank you. We now have new figures 2 and 3.

      (2) Subsection title for the Introduction has a typo.

      Thank you. We have fixed it.

      Reviewer #2 (Recommendations for the authors):

      (1) Technical quality control criteria are not specified.

      The Methods section states that "any incomplete or failed recordings were excluded," but does not define what constitutes a failed recording. The criteria could be subjective.

      Thank you for pointing this out. We agree that the definition of failed recordings should be clarified. We have now specified the exclusion criteria in the Methods section.

      “Recordings were excluded if they exhibited no spontaneous firing, abnormally slow firing rates, or failed to capture a complete action potential waveform. These criteria were applied consistently across all recordings.”

      (2) "Cell-specific" may overstate the claim.

      The term "cell-specific digital twins" (title, throughout) implies that the inferred parameters reflect the true biological state of each cell. However, parameters are derived only from curve-fitting to electrophysiological data and do not reflect other biological components (e.g., gene expression, contractility, calcium handling, metabolism, etc). Please consider rephrasing to "electrophysiology-based digital twins", "voltage clamp-matched digital twins", etc.

      Thank you for this important comment. We agree that the term “cell-specific” could be interpreted as implying a complete representation of the biological state of each cell. We have also adjusted the wording in relevant sections to avoid over-interpretation.

      Reviewer #3 (Recommendations for the authors):

      (1) I would add the list of the 52 parameters in the method section/SI and not just in the reference. Additional justification of why the perturbation was set as +/- 40% for the 52 parameter or +/- 20% for the EAD population would also help.

      Thank you for this helpful comment. We have included model equations and highlighted the 52 parameters in the Supplementary Information and provided additional justification in the Methods.

      (2) In Figure 1B, might be helpful to add the axis of the Vm instead of the dotted line indicating 0 mV to show differences in the diastolic potential.

      Thank you! We have now updated Figure 1B.

      (3) Figure 1C-I might be more impactful to show traces from the AP shown in Figure B to reinforce the impact of a single current in the AP shape.

      We have now updated Figure 1C-I to include traces from the AP shown in Figure 1B.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In the manuscript by Winke et al, the authors present evidence that fear-induced analgesia is mediated by somatostatin projection cells from the vlPAG to the RVM. This study uses a mouse model of fear-induced analgesia, and incorporates optogenetic circuit manipulation with behaviour and electrophysiology to gain a meaningful insight into a novel circuit involved in fear-induced analgesia.

      Strengths:

      (1) This is a well-constructed study with appropriate controls and analyses.

      (2) Alternative interpretations of the data are systematically considered and eliminated via rational experiments. The authors are commended for a nice piece of experimental work.

      (3) The vlPAG is a known region of pain modulation, and this study adds valuable insight to the circuit involved in fear-associated analgesia.

      We are very thankful to the referee for these positive comments.

      Weaknesses:

      (1) Only male mice are included in this study.

      We thank the reviewer for this point. We used only males in this first study for practical reasons to work with a population as homogeneous as possible. However, taking sex differences in biological mechanisms into account, we included this restriction in the summary and discussion

      (2) Animals are excluded from analyses based on clearly defined criteria, but it is not clear how many mice were excluded from each group.

      We thank the reviewers for raising this point. As stated in the Methods, we applied strict inclusion criteria for mice undergoing the hot-plate test, specifically a discrimination index ≥ 0.4 and a conditioning index ≥ 0.3. Using these criteria, 23% of wild-type mice were excluded for failing to meet the discrimination criterion. In the transgenic groups, an average of 20% of mice failed to meet the learning criteria, and an additional 12% were excluded due to incorrect opsin injection or misplaced optic fiber placement.

      (3) The authors implement a pain sensitivity assay that involves a hot plate with progressively increasing temperature. The time to nociceptive responses is reported. Without reporting the actual temperature at which the mice respond, it makes it difficult to compare nociceptive responses to previously published work (which typically use a defined and static hotplate temperature).

      We thank the reviewer for this comment. We provided this information related to the actual temperature of the nociceptive response in the original manuscript in supplementary figures 1, 2 and 5.

      (4) The authors present evidence that inhibition of SST vlPAG cells enhances spinal nociceptive electrophysiological responses, but the corresponding pain sensitivity is not altered (Figure 2, CS- condition). The reason for the discrepancy between electrophysiological and behavioural responses is not clear.

      We believe this comment arises from a misunderstanding of our results. In our study, inhibiting SST+ vlPAG cells did not increase nociceptive electrophysiological responses. Instead, it decreased spinal nociceptive transmission, as evidenced by reduced nociceptive field potentials and WDR responses in Figure 4c,e. Consistent with this electrophysiological effect, photoinhibition of SST+ vlPAG cells also produced behavioral analgesia, as evidenced by increased nociceptive response latency in the hotplate test under both CS− and CS+ conditions (Figure 2f). Therefore, our electrophysiological and behavioral findings are not contradictory but instead support the conclusion that inhibiting SST+ vlPAG cells reduces pain sensitivity regardless of defensive state. We will revise the text to clarify this point.

      Reviewer #2 (Public review):

      Summary:

      Wenke et al. investigated the role of vlPAG somatostatin-expressing neurons in the mediation of analgesia during defensive states. A newly developed paradigm of cued fear-conditioned analgesia, which consists of a combination of an auditory fear retrieval session and a pain test, was used to evaluate this cell population's contribution to fear-mediated analgesia. Optogenetic manipulation of vlPAG SST+ neurons modulated the responses to a nociceptive cue (Hot Plate) presented concomitantly with an aversively conditioned tone. At the same time, alterations in the freezing levels could be observed during optogenetic activation of vlPAG SST+ neurons. In order to disentangle the impact of these cells on analgesia from their impact on the expression of defensive behaviors, the authors performed electrophysiological recordings from the dorsal horn in the spinal cord of anesthetized mice. A vlPAG-RVM-DH pathway was identified to trigger nociceptive C-fibers upon optic activation of the RVM. Finally, pathway-specific activation of SST+ vlPAG-RVM neurons could abolish CS-induced analgesia.

      Strengths:

      The study addresses a relevant topic, that is, brainstem circuits for pain-modulatory mechanisms as part of defensive states evoked by threat. This is important because the circuit mechanisms underlying pain are still not fully understood, and defining molecular markers of cellular circuit substrates may support the identification of potential pharmaceutical targets in treating pain. The authors confirm a previous study in that a somatostatin-positive cellular population presents a crucial vlPAG circuit element mediating anti-nociceptive effects. Key novelty aspects of the present study are the demonstration that these neurons seem to play a role specifically in threat-induced analgesia. This was possible by the elegant design and application of a novel fear analgesia paradigm, combined with cell- and pathway specific optogenetics.

      We thank the referee for such positive feedback.

      Weaknesses:

      Despite the convincing and rigorous experimental approach, the study leaves some interpretational room when it comes to the proposed circuit mechanism. This could either be addressed by additional experiments or by more discussion of alternative circuit layouts.

      Major Comments:

      (1) The paper by Zhang et al. (https://pubmed.ncbi.nlm.nih.gov/36641028/), which identified a role for vlPAG SOM+ neurons in mediating anti-nociception in neuropathic pain, needs to be referenced and its results discussed, if not reconciled. While functionally, both studies find an analgetic role of vlPAG SOM+ neurons projecting to the RVM, Zhang et al., using slice physiology, characterize those neurons as glutamatergic. In Figure 4E of Zhang et al. they find general (fear-independent) analgetic effects with PAG-RVM specificity by performing chemogenetic experiments.

      We thank the reviewer for highlighting this important point. We agree that the study by Zhang et al. is highly relevant and should be discussed in the revised manuscript. Their work shows that inhibiting vlPAG SST/SOM neurons with chemogenetic methods produces analgesia in a neuropathic pain model, and in our study, we similarly found that inhibiting SST+ vlPAG neurons increases hotplate response latency (Figure 2f), which aligns with an analgesic effect. Additionally, we observed that activating SST+ vlPAG neurons suppresses fear-conditioned analgesia.

      At the same time, there are important differences between the two studies that may explain the differences in interpretation. First, the behavioral paradigms are not identical. Zhang et al. used a hotplate protocol where animals were directly exposed to a nociceptive temperature, whereas in our study, we used a progressive temperature ramp and explicitly compared responses during a conditioned stimulus (CS+) and a non-conditioned control stimulus (CS−). These controls were important for us to distinguish fear-specific effects from more general effects related to stress, arousal, sensitization, or other non-associative processes.

      Second, the two studies differ in experimental context. Zhang et al. examined this circuit in a neuropathic pain model, whereas our study focused on acute nociceptive processing and fear-conditioned modulation of pain. We therefore believe that the apparent discrepancy might reflect differences in pain state and behavioral context, rather than a direct contradiction.

      Finally, Zhang et al. showed in slice recordings that SST+ vlPAG neurons provide excitatory input to RVM neurons. This is an important finding that we now address in the revised manuscript. At the same time, because the RVM contains heterogeneous neuronal populations with different projection targets and functions, these recordings alone do not prove that all recorded RVM neurons are part of the descending pathway controlling spinal nociception. Therefore, we have revised the Discussion to explicitly acknowledge Zhang et al. and to emphasize both the similarities and differences between the two studies.

      It can be argued that in addition to the two functionally distinct inhibitory SOM subtypes hypothesized by Winke et al., there is another, excitatory subpopulation. Also, the different experimental conditions (chronic vs. acute pain, non-threat vs. fearful cues/contexts may recruit different vlPAG SOM+ populations. All of this is conceivable, yet I wonder whether the contrasting findings could more parsimoniously be reconciled. The author's own results presented here in Supplementary Figure 3 suggests that SOM+ vlPAG cells are colocalizing with glutamate and thus could also be excitatory. In addition to this rather complementary piece of evidence, a more extensive characterization of vlPAG neurons using IHC and slice physiology would be needed to justify the unambiguous identification of their inhibitory nature.

      We thank the reviewer for this thoughtful comment. We agree that our current data do not support a definitive conclusion that all SST+ vlPAG neurons are inhibitory. As the reviewer notes, our Supplementary Figure 3 shows that SST+ vlPAG cells can also co-localize with glutamatergic markers, which is consistent with the possibility of cellular heterogeneity within this population. We also agree that different experimental conditions, such as chronic versus acute pain and non-threatening versus fear-related contexts, may activate different SST+ vlPAG subpopulations.

      Our intention was not to claim that SST+ vlPAG neurons constitute a uniform inhibitory population, but rather that SST+ cells are strongly represented among inhibitory neurons in the vlPAG. We agree, however, that more detailed characterization, including additional immunohistochemical analyses and slice physiology, is necessary to more definitively determine the neurotransmitter phenotype and functional connectivity of these neurons. We have therefore revised the text to temper our interpretation and to explicitly acknowledge the likely heterogeneity of SST+ vlPAG neurons, including the possibility of an excitatory subpopulation. We therefore modified the discussion accordingly:

      “Our results align with the parallel inhibition- excitation model, where inhibitory and excitatory cells form two distinct, parallel descending pathways for pain modulation.

      Indeed, previous research demonstrated the presence of an inhibitory pathway projecting throughout the PAG–RVM-spinal cord dorsal horn neuraxis. Our results complement this study by suggesting that one of these previously proposed parallel pathways is mediated by SST+ vlPAG cells and has a functional role in mediating analgesia. At the same time, our data indicate that vlPAG SST neurons are heterogeneous, with approximately one-third of these cells co-localizing with excitatory markers. Together with the recent observation that excitatory SST+ vlPAG neurons project to the RVM (Zhang et al., 2023), this raises the possibility that a subset of long-range SST+ vlPAG neurons contributes to an excitatory descending pathway within the PAG–RVM–spinal dorsal horn neuraxis. By contrast, local GABAergic SST+ vlPAG neurons may participate in local circuit mechanisms related to defensive-state expression, including freezing. Further anatomical and functional studies will be required to resolve these possibilities.”

      In the absence of a direct identification of these cells exclusively releasing GABA, an alternative explanation should be considered. What about looking at vlPAG SOM+ neurons as a putatively mixed bag of local, inhibitory interneurons and long-range, RVM-projecting excitatory cells? This model would then open up interesting questions as to the actual function of somatostatin as a modulator of vlPAG circuit activity and associated function, and from my perspective, would nicely fit into the view of PAG circuits as integrators of complex survival responses.

      We thank the reviewer for this insightful suggestion and agree that, in the absence of direct evidence that vlPAG SOM+/SST+ neurons are exclusively GABAergic, an alternative interpretation should be considered. In particular, we agree that this population may be heterogeneous and could include both local inhibitory interneurons and long-range excitatory neurons projecting to the RVM. We believe this is an important and constructive framework for interpreting our data, and we have revised the Discussion accordingly. In the revised text, we now explicitly acknowledge the likely heterogeneity of vlPAG SST+ neurons and discuss the possibility that distinct local and long-range SST+ subpopulations may contribute differently to defensive-state regulation and descending pain modulation. We agree with the reviewer on this point and have modified the discussion accordingly (see point above).

      (2) "Our data indicate that the optogenetic inhibition of SST+ vlPAG cells promotes analgesia irrespective of the animal's defensive state. In contrast, the optogenetic activation of long-range SST+ vlPAG cells that project to the rostral ventromedial medulla (RVM) abolishes the analgesia mediated by fear behavior." (lines 32-35). Consider toning down these conclusions, as contrasting activation with inhibition of two different (though overlapping) populations cannot be fully conclusive. Alternatively, a pathway-specific (vlPAG-RVM) inhibitory experiment could help to fully understand the circuit mechanism and verify the necessity of these neurons.

      We thank the reviewer for raising this point. We agree that inhibition of the entire SST+ vlPAG population and activation of the long-range SST+ vlPAG neurons projecting to the RVM population are not directly equivalent manipulations. Our conclusion was intended at the level of observed functional effects: inhibition of SST+ vlPAG neurons promotes analgesia regardless of the defensive state, while activating long-range SST+ vlPAG neurons projecting to the RVM suppresses fear-conditioned analgesia. This occurs regardless of whether the SST vlPAG neurons are excitatory or inhibitory. To address the excitatory or inhibitory nature of SST vlPAG neurons, we have revised the discussion to include a reference to the Zhang et al study.

      (3) Despite an overall very thorough reporting style, some information is missing from the manuscript:

      (a) In Figures 2d and f, what are the freezing levels during optogenetic manipulation? From Figure 3d, one can expect that freezing is inhibited during the hot plate test, which could bias the NC response towards shorter latencies.

      We thank the reviewer for this important comment. As shown in Figure 1e, we previously quantified freezing both at CS onset and at the time of the nociceptive response in the hot plate test. These analyses indicate that freezing levels at the time of the nociceptive response do not differ between the CS+ and CS− conditions. Therefore, the variation in hot plate response latency is unlikely to be due to differences in freezing at the time of response.

      We acknowledge, however, that freezing was not directly measured during optogenetic manipulation in this experiment. Based on the temporal profile of freezing shown in Figure 1e, we still consider it unlikely that the effect of optogenetic manipulation on nociceptive latency is mainly caused by a change in freezing behavior.

      (b) In Figure 5, the histological experiment showing the vlPAG-to-RVM pathway is presented by a qualitative image only. Here, some quantification would strengthen the finding.

      We thank the reviewer for this comment. The aim of the histological experiment in Figure 5 was to provide qualitative anatomical evidence that vlPAG projections reach the RVM and are positioned in close apposition to spinally projecting RVM neurons. We did not intend this experiment to serve as a quantitative characterization of connectivity. We agree that a more systematic quantification would be informative, but this would require additional dedicated experiments beyond the scope of the present manuscript.

      (c) In Figures 6 c and d "Consistently, activation of the SST+ vlPAG-RVM pathway during CFCA had no impact on CS-presentation, whereas the same manipulation performed during CS+ blocked the increase in NC response latency compared to GFP controls." (line 194-196). Is it possible that the NC response cannot be any lower than the one during CS-, thus constituting a floor effect?

      We are thankful to the reviewer for this important point. We agree with the reviewer that this is indeed a possibility. We have added a sentence in the discussion to acknowledge this limitation.“Another possibility is that our nociceptive test with a slow ramp of temperature induces a floor effect on nociceptive response latency, which may limit the detection of further decreases in latency under certain conditions.”

      (c) Connected to major point 1- this experiment is important for defining the circuit mode and therefore should be as convincing as possible. However, for the colocalization experiment in Supplementary Figure 3, the methodological description is missing and thus makes it hard to comprehend how this data set was generated (how many data points, etc.). The visual depiction of the results is non-standard and not easily graspable. Consider e.g., a Venn diagram.

      We apologize for this omission in the original manuscript. We have now provided this methodological information in the method section. We have now expanded the description of these data in the figure legend to ease the comprehension of the figure.

      Reviewer #3 (Public review):

      Summary:

      Conditioned analgesia refers to the ability of a learned fear cue to suppress pain-related behavior and neural activity. Understudied, the authors developed a novel conditioned analgesia procedure in which a cue that had been paired or unpaired with shock was played while a hot plate increased temperature. Compared to several control conditions, the authors found increased latency to a nociceptive response (paw licking). The authors identified somatostatin neurons in the periaqueductal gray as a likely mediator of the behavior. They then showed that: (1) stimulating vlPAG-SST neurons blocked nociceptive response latency increases to the CS+, (2) stimulating vlPAG-SST neurons suppressed fear retrieval freezing, (3) stimulating vs. inhibiting vlPAG-SST neurons drove opposing modulation of c-fibers and Aδfibers, (4) direct-projecting vlPAG SST neurons modulate freezing while RVM-projecting vlPAG SST neurons modulate conditioned analgesia.

      Strengths:

      These experiments have many strengths. The behavioral assay is chief among them. The assay is robust and controls for confounding factors to reveal a repeatable effect of a shock-paired cue to delay nociceptive responding. The optogenetic experiments provide the correct level of temporal precision, given the authors' time-specific interest in cued responding. Combining neuronal manipulations with spinal recordings is particularly innovative, especially in the context of more behavioral neuroscience-based assays. All-in-all, I found this to be an exceptionally strong set of experiments.

      Weaknesses:

      No obvious weaknesses were identified by this Reviewer.

      Recommendations for the authors:

      Comments from Reviewing Editor:

      Summary

      Three reviewers have assessed your manuscript on vlPAG somatostatin pathways contributing to conditioned analgesia. Conditioned analgesia refers to the ability of a learned fear cue to suppress pain-related behavior and neural activity. Understudied, the authors developed a novel conditioned analgesia procedure in which a cue that had been paired or unpaired with shock was played while a hot plate increased temperature. Compared to several control conditions, the authors found increased latency to a nociceptive response (paw licking). The authors identified somatostatin neurons in the periaqueductal gray as a likely mediator of the behavior. They then showed that: (1) stimulating vlPAG-SST neurons blocked nociceptive response latency increases to the CS+, (2) stimulating vlPAG-SST neurons suppressed fear retrieval freezing, (3) stimulating vs. inhibiting vlPAG-SST neurons drove opposing modulation of c-fibers and Aδ-fibers, (4) direct-projecting vlPAG SST neurons modulate freezing while RVM-projecting vlPAG SST neurons modulate conditioned analgesia.

      Strengths

      All three reviewers converged on multiple strengths. The assay developed was seen to be novel, rigorous, and included a variety of controls that convincingly demonstrated conditioned analgesia. Focusing on the ventrolateral periaqueductal gray, and more specifically on somatostatin-expressing cells, made prior sense, and the results more than justified this selection. Approaching the vlPAG and circuits with many converging methods provided further, compelling evidence for a role in conditioned analgesia.

      Weaknesses

      Specific weaknesses are described in the individual reviews. Generally, the following weaknesses were identified. The study only used male mice, a choice that should be better justified. Animals were reasonably excluded from analysis, but the final group ns for analyses were not always clear. Some statistical results lacked clarity. The relevance of these findings to prior work (particularly Zhang et al. 2023, Journal of Pain) was not always described. Relatedly, the results would be better contextualized by appreciating and describing the likely diversity of somatostatin functional types and projection types.

      Recommendations

      (1) Provide rationale for only using male mice, discuss the limitation of the exclusion of females, and note that male mice were the subjects in the abstract.

      Thank you for this recommendation, we have mentioned this information in the abstract and in the discussion. We have also mentioned the limitations of not including female mice in the abstract and the discussion of the revised manuscript.

      (2) Complete final report ns for each statistical analysis. If you have not already done so, please include full statistical reporting including exact p-values wherever possible alongside the summary statistics (test statistic and df) and, where appropriate, 95% confidence intervals. These should be reported for all key questions and not only when the p-value is less than 0.05 in the main manuscript.

      An extended table with all statistical tests and analysis for all figures has been provided in sup Table 1.

      (3) Include example videos of CFCA sessions, demonstrating optogenetic effects.

      We understand the editor’s request to include video material illustrating the behavioral responses. However, we would prefer not to include such videos in the manuscript, in accordance with our institution's guidelines and recommendations on the dissemination of animal experimentation footage. Importantly, all behavioral sessions were systematically video-recorded from both sides of the apparatus, allowing detailed offline analysis of the animals’ responses. These recordings were carefully examined by an experienced experimenter to assess nociceptive behaviors, including jumping responses and licking of the stimulated hindpaw. This procedure ensured a reliable and accurate evaluation of pain-related behavioral reactivity. While the videos themselves cannot be included in the manuscript for the reasons mentioned above, we believe that the behavioral scoring procedures described in the Methods section provide a clear and rigorous description of how these responses were assessed. In addition, Figure 1 includes an example image illustrating hindpaw licking behaviour, which is typically more subtle and more difficult to identify than jumping responses. We therefore believe that this visual example, together with the detailed description of the scoring procedure and the quantitative data provided, adequately supports the interpretation of the behavioural results.

      (4) Provide summary expression and ferrule placement figures.

      We thank the editor for this comment. We have now included schematic summaries of fiber placements for both SST and VIP mice used in this study, based on histological verification (Supplementary Figures 10 and 11). Representative images of viral expression are also provided (Figure 2a, Supplementary Figure 7b and f).

      (5) Detail how behavior judgments were made.

      We thank the editor for emphasizing this important methodological point. During all behavioral sessions, mice were video-recorded simultaneously from both sides of the apparatus, allowing a comprehensive and unobstructed view of the animals’ posture and movements throughout the experiment. These recordings were subsequently analyzed offline by an experienced experimenter trained to evaluate nociceptive behaviors. Pain-related behavioral responses were assessed based on well-established indicators of nociceptive reactivity. In particular, we quantified overt escape-like reactions such as jumping, which reflects a strong aversive response to the stimulus. In addition, we evaluated more localized nociceptive behaviors directed toward the stimulated limb, including licking of the hindpaw. These measures are commonly used in rodent pain assays and provide reliable behavioral readouts of nociceptive sensitivity. The combination of bilateral video recordings and expert behavioral scoring ensured that both subtle and robust nociceptive responses could be accurately detected and categorized during the analysis.

      (6) Provide the temperature at which nociceptive responses were initiated. Check grammar and references.

      The temperature at which nociceptive responses were initiated were originally reported in Supplementary Figure 1, 2 and 5.

      Reviewer #1 (Recommendations for the authors):

      (1) The authors use optogenetic manipulation of SST activity in the vlPAG to show that this cell type is involved in fear-induced analgesia. They include a valuable control to show that manipulation of another inhibitory cell type (VIP) also does not impact analgesia. It would be helpful to know the expression level of VIP cells in the vlPAG. Is this a predominant inhibitory projection cell in the vlPAG (besides SST)?

      We thank the reviewer for pointing this. While we did not quantify the expression level of VIP+ cells in the vlPAG in the present study, available data suggest that this population is relatively sparse compared to other inhibitory cell types. In particular, reference to the Allen brain atlas indicates that VIP gene expression in the vlPAG is limited and primarily localized around the fourth ventricle, within the lateral and ventrolateral PAG, rather than broadly distributed across the region. Consistent with this, we provide an example of viral expression in VIP-Cre mice in Supplementary Figure 7f, illustrating the restricted distribution of VIP+ neurons in the vlPAG. We have also provided a summary of ferrules placement for SST and VIP mice used in our study in Supplementary Figures 11 and 10, respectively.

      (2) The numbers of animals dropped from each experiment should be indicated - perhaps on the statistics table?

      We thank the reviewer for pointing this.

      As stated in the Methods, we applied strict inclusion criteria for mice undergoing the hot-plate test, specifically a discrimination index ≥ 0.4 and a conditioning index ≥ 0.3. Using these criteria, 23% of wild-type mice were excluded for failing to meet the discrimination criterion. In the transgenic groups, an average of 20% of mice failed to meet the learning criteria, and an additional 12% were excluded due to incorrect opsin injection or misplaced optic fiber placement.

      (3) Line 105: "...,which activity..." change to "..., whose activity..."

      Done

      Reviewer #2 (Recommendations for the authors):

      (1) Please also provide absolute temperature values of the nociceptive response threshold.

      The temperature at which nociceptive responses were initiated was originally reported in Supplementary Figure 1, 2 and 5.

      (2) It would be nice to see an example video of a CFCA session (with and without optogenetic manipulation).

      We understand the editor’s and reviewer’s request to include video material illustrating the behavioral responses. However, we would prefer not to include such videos in the manuscript, in accordance with our institution's guidelines and recommendations on the dissemination of animal experimentation footage. Importantly, all behavioral sessions were systematically video-recorded from both sides of the apparatus, allowing detailed offline analysis of the animals’ responses. These recordings were carefully examined by an experienced experimenter to assess nociceptive behaviors, including jumping responses and licking of the stimulated hindpaw. This procedure ensured a reliable and accurate evaluation of pain-related behavioral reactivity. While the videos themselves cannot be included in the manuscript for the reasons mentioned above, we believe that the behavioral scoring procedures described in the Methods section provide a clear and rigorous description of how these responses were assessed. In addition, Figure 1 includes an example image illustrating hindpaw licking behaviour, which is typically more subtle and more difficult to identify than jumping responses. We therefore believe that this visual example, together with the detailed description of the scoring procedure and the quantitative data provided, adequately supports the interpretation of the behavioural results.

      (3) Please provide a schematic summary of fiber placements and opsin expressions confirmed by histological examinations.

      We thank the reviewer for this comment. We have now included schematic summaries of fiber placements for both SST and VIP mice used in this study, based on histological verification (Supplementary Figures 10 and 11). Representative images of viral expression are also provided (Figure 2a, Supplementary Figure 7b and f).

      (4) "Valid nociception readout responses included jumping or licking the hindpaw." (Line 453). How was this evaluated- manually or automated, blinded etc.?

      We thank the reviewer for emphasizing this important methodological point. During all behavioral sessions, mice were video-recorded simultaneously from both sides of the apparatus, allowing a comprehensive and unobstructed view of the animals’ posture and movements throughout the experiment. These recordings were subsequently analyzed offline by an experienced experimenter trained to evaluate nociceptive behaviors. Pain-related behavioral responses were assessed based on well-established indicators of nociceptive reactivity. In particular, we quantified overt escape-like reactions such as jumping, which reflects a strong aversive response to the stimulus. In addition, we evaluated more localized nocifensive behaviors directed toward the stimulated limb, including licking of the hindpaw. These measures are commonly used in rodent pain assays and provide reliable behavioral readouts of nociceptive sensitivity.The combination of bilateral video recordings and expert behavioral scoring ensured that both subtle and robust nociceptive responses could be accurately detected and categorized during the analysis.

      (5) Line 226 REF33 doesn't seem to fit.

      The reference list has been updated. Related to this section in which we discuss the disinhibition mechanisms inducing nociception in chronic stress mice. We have cited the work of Samineni et al., 2015 (reference 15) and Tovote el al., (reference 23) both related to these disinhibition mechanisms.

      Full sentence for reference 33 (now 35): “Two independent previous studies found that long-range inhibitory inputs from the central medial amygdala contact inhibitory cells within the vlPAG, implicated in different roles: the modulation of fear behavior (23) and nociceptive transmission (35)”.

      Ref 35 - Yin, W. et al. A Central Amygdala–Ventrolateral Periaqueductal Gray Matter Pathway for Pain in a Mouse Model of Depression-like Behavior. Anesthesiology 132,1175–119 (2020)

      (6) Some minor language, semantic, and grammatical flaws.

      The manuscript has been evaluated for language, semantic and grammatical flaws

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #2 (Public review):

      (1) There are certainly some areas of the manuscript that would benefit from deeper exploration, such as electron microscopy/other imaging approaches to explore whether deletion of PfMSP2 has a visible impact on merozoite surface structure.

      We in principle agree with the reviewer that applying enhanced resolution microscopy approaches to understand structural and functional changes with loss of PfMSP2 could be of interest. However, based on our ongoing work, this represents a significant body of work in terms of experimental optimisation in an effort to gain the detail required to make meaningful insights. Therefore, this will remain outside the scope of this manuscript and we hope to provide these insights in future studies.

      (2) Further replicates of the video microscopy assays to see whether trends in the data could reach significance (although these are very time-consuming and technically difficult assays).

      Conclusions we have drawn from live-cell imaging data for MSP2 knock-out parasites encompass some 43 invading merozoites from 21 schizont ruptures for PfDd2 WT and 35 invading merozoites from 18 schizont ruptures for PfDd2 DMSP2 parasites. One of the leading studies to apply live-cell microscopy to film invading merozoites based conclusions of invasion kinetics on: 3D7 (number of merozoite invasion =63, number of schizont ruptures =23), D10 (invasions =33, ruptures =20) and W2mef (invasions =39, ruptures = 15; this line is of the same lineage as Dd2) (Weiss et al. PLoS Pathogens, 2015). Although there are variations within and between lines from this gold-standard study, our dataset is mostly comparable in terms of the number of schizont ruptures and merozoite invasions filmed and analysed to look at changes in kinetics. What we can say definitively is that there is no strong phenotype in the absence of inhibitory antibodies against other antigens for either live-cell or growth inhibition assays. Therefore, we have focussed the data interpretation in the manuscript to highlight the lack of statistical significance and limited phenotype seen, which given the previously believed importance of MSP2 to P. falciparum invasion of red blood cells is somewhat surprising.

      In order to address this suggestion, we have modified the discussion to better represent any non-significant changes in invasion and growth seen.

      “Despite the abundance of PfMSP2 on the merozoite surface and previous work suggesting a role in RBC invasion, we found merozoites invade and grow with similar kinetics to wildtype parasites in the absence of PfMSP2. This does not exclude a role for PfMSP2 in vivo where there are additional pressures, such as immune-effector mechanisms and flow dynamics, on merozoite invasion. However, given we have knocked-out PfMSP2 from two different P. falciparum isolates, our findings do not currently support a major role for PfMSP2 in the mechanics of merozoite invasion. Thus, it appears that the function of the two most abundant proteins on the merozoite surface, PfMSP1 (Das et al., 2015; Kals et al., 2024) and PfMSP2, are not obviously linked to merozoite binding to the RBC and subsequent invasion.”

      (3) Follow up of some of the genes where expression is changed by PfMSP2 knockout (as the authors point out, there are no candidates that have a very obvious link to invasion suggesting that they may be compensating for PfMSP2 function, although several are expressed in schizont stages).

      A thorough investigation of the genes where expression changes with PfMSP2 knock-out would require a substantial body of additional work, not least because they would all have to be investigated as there is no single likely candidate based on stage of expression, membrane binding properties or previous links to merozoite surface architecture. Given this, potential follow up of these proteins will be left for future studies.

      We also thank the reviewer for the recognition of the work provided in the manuscript and the modifications made that have improved the manuscript from version 1. The reviewer also recognises the value in our detailed characterisation, including data where phenotyping changes with MSP2 knock-out could not be seen, in defining the function of PfMSP2 as commented below:

      However, there is already a substantial amount of data in the manuscript, and more detailed follow-up is reasonable to leave to future work. Overall, with the modifications made through the review process, including the addition of new controls for key experiments, the claims and conclusions are justified by the data, and the manuscript generates important new information about a highly studied Plasmodium falciparum merozoite surface protein.

      Reviewer #3 (Public review):

      Major points:

      (1) Much of the manuscript describes negative results and this reviewer found it arduous to get through many negative or nonsignificant results before finally getting to the significant effect on AMA1 inhibitory antibodies, not presented until Figure 6! Computational studies in Fig. 1 could be a supplementary figure. Figs. 2 and 3. demonstrate knockout in 3D7 and Dd2, respectively and could be assembled into a single figure. (Notably Fig. 2A and 3A are almost identical with use of some different primers.) Fig. 2E, 2F, 3D-H, all of Fig. 4, most of Fig. 5 are all negative or insignificant results that could also be moved to supplementary data. As MSP4, MSP5, and SUB1 are presumably included in the whole genome RNA-seq experiments shown in Fig. 4C, it makes sense to remove Fig. 4A data from the paper fully. These consolidating changes would help highlight the key finding of improved binding and block of AMA1's role in invasion.

      We have chosen to not take the approach proposed by Reviewer 3 as it would leave the manuscript with only around 2.5 Figure panels and undersells the very significant amount of work that has been done to characterise PfMSP2 knock-out lines. Although, as noted by the reviewer, piggyBac mutagenesis studies predict PfMSP2 is dispensable, much of the field likely expect PfMSP2 to be essential to P. falciparum blood stage parasite growth due to the results of earlier reverse genetics approaches and many years of publications that have speculated on the importance of the protein. Therefore, we are also conscious of providing very clear and comprehensive evidence to support our findings. While this may delay highlighting the findings in Figure 6, we also note that the lengths we have gone to in characterising an important antigen with a difficult phenotype is still valued as evidenced by Reviewer 2 (Public Review Comments on the original manuscript):

      “PfMSP2 knockouts are made in two different strains, which is important as it is known that invasion pathways can vary between strains, but is a level of comprehensiveness that is not always delivered in P. falciparum genetic studies. The knockout strains are characterised very thoroughly using multiple different assays, and the authors should be commended for publishing a good deal of negative data, where no phenotype was detected.”

      (2) The potentiating effects on anti-AMA1 antibodies are shown with rabbit sera and purified antibodies, mouse monoclonal antibodies, and smaller i-bodies inspired by shark antibody-like receptors but not with human monoclonal antibodies (hmAbs). As naturally acquired hmAbs targeting AMA1 have been identified and characterized (PMIDs: 39632799, 40020675), would it not be important to test these antibodies in the ∆MSP2, especially as the authors emphasize the importance of their model in designing better human malaria vaccines?

      As the reviewer noted, we demonstrated enhanced inhibitory activities of antibodies to AMA1 using rabbit polyclonal antibodies, mouse mAbs, and i-bodies. We note that the WD34 i-Body we used was humanised to be IgG-like with a human Fc-region (IgG1 backbone). Rabbit IgG is very similar to human IgG1. Therefore, we have provided evidence of the enhancing effect using different types and sources of antibodies relevant to human immunity to support our conclusions. Our findings open new avenues for future research and we agree with the reviewer that future studies using panels of human mAbs to defined epitopes would be interesting and may further inform vaccine design; however this is beyond the scope of the current paper. We do not have the mAb mentioned by the reviewer to test in our system. To perform studies with human mAbs would take a substantial amount of time (many months), requiring the generation of different human mAbs and quantification of their activity and testing them for potentiation effects. While this would be an interesting future endeavour, we do not feel that such studies are needed at this stage to support our conclusions, and instead would be a future extension from our current paper. To acknowledge the reviewer's comment, we have extended our comment in the discussion about future studies with different panels of invasion inhibitory antibodies to include huMabs targeting AMA1 as follows:

      “Further investigation using the parasite lines developed in this study and a wider panel of antibodies that target different stages of the merozoite invasion process, including human monoclonal antibodies against AMA1 (Patel et al., 2025), could shed more light on this potentially novel mechanism of vaccine derived antibody efficacy.”

      (3) Fig. 7 presents quantitative fluorescence microscopy to measure anti-AMA1 binding and support a model where MSP2 serves to sterically hinder antibody access to AMA1 on individual merozoites. I understand that the negative WD33 control is useful to contrast to the positive WD34 antibody (both bind AMA1 but only WD34 exhibits parasite growth inhibitory effects), but it seems that use of smaller i-bodies rather than conventional larger mouse or ideally human monoclonal antibodies may compromise demonstration of steric hindrance by MSP2 because smaller i-bodies may be less hinder.

      The antibodies used in this experiment have fluorescent tags attached. So while the untagged WD33 and WD34 i-bodies are approximately 14 kDa, when fused to GFP or mCherry their expected size increases to approximately 42 kDa, approaching that of the Fc-tagged WD34 i-body (78 kDa) that shows increased growth inhibitory activity in the absence of MSP2. Therefore, we expect steric hindrance to be a significant factor with these fluorescently tagged antibodies.

      (4) Some explanation for why WD33 fails to inhibit growth despite targeting the same antigen as WD34 is needed. Are the epitopes known? Does one bind further from the RON2 binding pocket?

      As reported in Angage et al., Nature Communications 15, 7206 (2024). WD34 has been identified to bind to, and block, a site within the hydrophobic AMA1 and RON2 binding pocket found on Domain II of AMA1. In contrast, WD33 recognises a distinct conserved epitope in Domain II of AMA1 near to, but not overlapping with, the hydrophobic AMA1 and RON2 binding pocket. We have clarified this by including additional description when first describing the i-bodies as follows:

      “When we tested the i-body WD34 (Angage et al., 2024) which binds a highly conserved epitope that includes the PfRON2-binding pocket on PfAMA1 domain II, we observed a small potentiation of PfAMA1 specific activity with knock-out of PfMSP2 in Pf3D7 (1.3-fold; IC<sub>50</sub> PfD7 WT 0.012 mg/mL; IC<sub>50</sub> Pf3D7 DMSP2 0.009 mg/mL; p=0.08 Figure 6F).”

      Then

      “A second i-body, WD33 (Angage et al., 2024), which binds AMA1 between domain II and domain III but does not appear to overlap with the PfRON2-binding pocket on PfAMA1, had very limited invasion inhibitory activity against Pf3D7 parasites and did not show improved potency with knock-out of Pf3D7 MSP2 (0.9-fold; IC<sub>50</sub> Pf3D7 WT 1.02 mg/mL; IC<sub>50</sub> Pf3D7 DMSP2 1.1 mg/mL; p=0.8; Figure 6I).”

      Recommendations for the authors:

      Reviewing Editor Recommendations:

      Although providing microscopic images might require a lengthy process, including results based on human mAbs (if available) might enhance the strength of evidence. The reorganization of the figures and the presentation of results usually falls into the realm of personal preferences, however, if the comments/suggestions are useful, it might highlight your message.

      As covered in the Response to Public Reviewer Comments for Reviewer 2 and indicated by the editor, investigations of phenotypes found in this study using high-resolution imaging techniques (e.g. electron microscopy) will require very significant additional work and will be attempted in future studies. We also provide a response to Reviewer 3 in regards to the potential to test human monoclonal antibodies and believe this is best done more thoroughly in future studies. We have elected to not make substantial changes to the data presented as suggested by Reviewer 3. We have addressed additional comments as covered below.

      Reviewer #3 (Recommendations for the authors):

      Minor Comments

      (1) Scale bar in Fig. 7A is not resolved well. The image is too pixelated to resolve merozoites or the actual dimensions of the scale bar.

      We have updated this figure to provide improved clarity of the scale bar.

      (2) Lines 69, 216, 221, 253, 628-629, 648 all suggest that MSP2 was heretofore assumed to be essential. However, piggyBac insertional mutagenesis revealed that MSP2 is highly dispensable (MIS of 0.988, per PlasmoDb.org; PMID: 29724925). I would suggest to tone down this claim as it does not detract from the authors' production of useful ∆MSP2 clones.

      We agree with the reviewer that the piggyBac insertional mutagenesis study results should also be acknowledged and apologise for this oversight. To address this, we have reviewed the sentences highlighted by the reviewer and, where appropriate for the historical interpretation of PfMSP2 function, have added the following modified information through the text:

      P. falciparum merozoite surface protein 2 (PfMSP2), an antigen reported to be refractory to gene knock-out in P. falciparum (Sanders et al., 2006) but that has also been reported to be dispensable in a piggyBac mutagenesis study (Zhang et al., 2018), has been of long-term interest as a vaccine candidate.”

      “Given previous unsuccessful attempts to disrupt pfmsp2 (Sanders et al., 2006), and its high abundance on the merozoite surface (Gilson et al., 2006), PfMSP2 has been traditionally viewed as an essential P. falciparum protein with an essential function in merozoite invasion, although more recent piggyBac mutagenesis studies have called this understanding into question (Zhang et al., 2018).”

      We have chosen not to modify this text and it remains the same as below. The reason for not changing this text is the result that we could knock-out MSP2 from 3D7 was still unexpected given the published reverse genetics studies and results from piggyBac mutagenesis studies are also sometimes not reliable indicators of what happens when reverse genetics is performed. Therefore, the following text we believe is a reasonable description.

      “Unexpectedly, we confirmed successful disruption of pfmsp2 by replacing the coding sequence between 132 bp and 819 bp of the gene with a hDHFR drug selection cassette in the 3D7 P. falciparum laboratory-adapted line (Figure 2A and B), resulting in Pf3D7 DMSP2 parasites.”

      “As a previous reverse genetics study in 3D7 reported that PfMSP2 was essential for P. falciparum growth in vitro (Sanders et al., 2006), we investigated whether PfMSP2 could also be removed from PfDd2, an isolate of P. falciparum that differs from 3D7 in geographical origin, RBC receptor usage and allelic type of pfmsp2.”

      “However, CRISPR-Cas9 gene editing used in this work has shown that, in contrast to previous attempts to knock-out PfMSP2 (Sanders et al., 2006), PfMSP2 is not essential for P. falciparum blood stage parasite growth in vitro.”

      “Advancements in gene-editing techniques in P. falciparum have allowed us to directly demonstrate using reverse genetics in two different parasite lines that PfMSP2 is not essential for P. falciparum growth in vitro.”

      (3) Figs. 2B, 2C, 2D show PCR, immunoblots, and IFA with a ∆MSP2 clone but two clones (termed clone 1 and clone 2) are show in panels 2E and 2F. Which clone is used in each panel? Without clarification, readers may wonder if one clone was used for PCR but another clone gave a desired result in immunoblots? By convention, validation studies (PCR and immunoblots) should be performed and shown (in Supplementary figures) for all clones used for phenotype studies; alternatively, a single clone can be used throughout if all clones are presumed identical. Which of these clones was used for the RNA-seq experiments in Fig. 4C? Similar questions arise for the two knockout clones made in the Dd2 line (Fig. 3D).

      We agree with the reviewer that it would be helpful to have this information provided more clearly through the Results. To this end, we have updated the Figure legends across Figures 2, 3, 4, 5, 6, 7 and Supplementary Figure 5 as appropriate to specifically indicate the clones used for the downstream experiments. All clones were validated by PCR and, after growth characteristics were found to be the same, a single clone was used for all downstream experiments for PfMSP2 knock-outs in both 3D7 and Dd2.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public Review):

      Summary:

      This study aims to understand how cell fusion contributes to wound healing using a laser-induced injury in the notum epithelium of a developing fruit fly. The authors meticulously characterize the epithelial fusion events using a live imaging approach and report that syncytia arise by 'border breakdown' and 'cell shrinking'. The syncytial epithelial cells also appear to outcompete mononucleated cells and preferentially dissolve their tangential borders, which correlates with the accumulation of actin at the leading edge.

      Strengths:

      The strength of this study is the authors' live imaging approach to capture these dynamic fusion events that are a fundamental, yet poorly understood biological process.

      Weaknesses:

      A major weakness is that all the authors' conclusions are based on descriptive studies, in which the role of cell fusion is not directly tested. This is particularly important because other models of wound induced polyploidization have demonstrated that another cytoskeletal protein, myosin, was upregulated and dependent on endoreplication, and not cell fusion. Therefore it remains unclear to what extent cell fusion, endoreplication, or both are required to outcompete mononucleated cells as well as pool actin as described in this study.

      We thank the reviewer for appreciating our live imaging and meticulous approach. In this revision we have identified that the gene Atg1 is required for wound-induced fusion in the pupal notum: when Atg1 is knocked down, there is a reduction in wound-induced cell fusions, both border breakdown and cell shrinking. Analysis of Atg1 knockdown shows that the wounds close more slowly. This is a direct test of the role of cell fusion in speeding wound closure, presented in new Fig. 4.

      Reviewer #2 (Public Review):

      Summary:

      Overall, this study provides a thorough description of the formation of syncytia following wounding of the proliferation-competent diploid epithelium of the pupal notum. While this phenomenon has already been described briefly for this particular tissue by the Galko lab in Wang et al 2015, the authors provide a much more detailed description and characterisation of the process providing some novel insights (radial versus tangential border breakdown, cell shrinkage, timings, syncytia outcompeting mononucleated cells, etc.).

      Strengths:

      This paper provides an elegant, thorough, descriptive characterisation of syncytia-driven wound closure using state-of-the-art confocal live imaging of the pupal notum. The authors show that laserinduced wounding of this diploid, proliferation-competent epithelium results in the formation of syncytia of various sizes in the first few cell rows around the wound edge, which progressively become bigger as healing proceeds. This results in ~50% of cells becoming part of these syncytia. The cell fusion events were convincingly demonstrated by showing the disappearance of p120ctnRFP and E-Cadherin-GFP from cell-cell borders as well as cytoplasmic GFP mixing of GFPpositive cells with a GFP-negative cell.

      Apart from cell-cell fusion by border breakdown that mostly happens in the first 2h following wounding, the authors also found that at later stages of wound healing cell shrinkage following cytoplasmic mixing contributed to sycytia formation.

      Next, the authors provided some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound.

      The authors then show that radial border breakdown occurs much less frequently than tangential border breakdown. They suggest that radial border breakdown reduces the requirement for cell-cell intercalations. They also hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. To test this, the authors generate single-cell clones that overexpress Actin-GFP. They then show convincingly how a single Actin-GFP-positive cell in the second cell row fuses with one GFP-negative cell in the first cell row. The Actin-GFP signal then spreads in the fused cell and labels some previously unlabelled actin-rich structure near the wound edge which most likely is the actomyosin cable. This provides some evidence for resource sharing by cytoplasmic mixing following fusion.

      Weaknesses:

      The authors provide some convincing evidence that syncytia outcompete mononuclear cells for being positioned in the first cell row around the wound. The authors suggest that the syncytial cells might be better able to close the wound. However, some genetic studies would need to be done to establish this more convincingly. E.g. Could the authors genetically block syncytia formation and then show that these wounds now heal slower?

      We now present such data in new Fig. 4, which describes knocking down Atg1, previously shown by the Leptin lab to promote wound-induced fusions in larval epidermis. We quantify the resulting reduction in fusion in the pupal notum and show that the leading edge advances more slowly to heal the wound.

      The authors suggest that radial border breakdown reduces the requirement for cell intercalation. While this might be true it also raises the question of how the various syncytia facing the wound border change shape to allow the shrinkage of the first cell row over time to allow wound closure. None of the four movies included in the study shows the whole wound healing process until the later stages, making it hard to assess this. It would be good to include one such movie showing the syncytia in the whole wound and comment on this point.

      In response to the reviewer's request, we now extend Supplemental Video S1 out through 8 hours after wounding (same video as included previously but extended longer). In this video, as in many of the wounds, it is hard to determine the exact moment of closure because a syncytium extends across the wound whereas the nuclei do not. However, during the process of closure, one can clearly observe the large syncytia becoming more wedge-shaped – drastically reducing the section of their perimeter remaining in contact with the wound’s leading edge.

      In addition, we now explore how syncytia reduce the need for intercalation in a computational model, presented in new Fig. 7 and Supplemental Videos S5 and S6. One can observe the modeled syncytia becoming similarly wedge-shaped. The modeling shows that the presence of syncytia and their ability to reshape can speed closure by about 1/3 even if the syncytia have no special properties aside from their relative size.

      In both the experiments and models, some syncytia are also removed from the leading edge by intercalation, but the presence of syncytia reduces the total number of intercalations needed.

      The authors hypothesise that tangential border breakdown might allow fused cells to share resources and provide more resources to be used near the wound edge, e.g. for actomyosin cable formation. They show convincingly through the fusion of a single Actin-GFP-positive cell in the second cell row with a GFP-negative cell in the first cell row that Actin-GFP spreads in the fused cell and labels the previously unlabelled actomyosin cable. While the hypothesis of resource sharing to improve healing is intriguing and makes sense, this experiment doesn't necessarily prove the benefit of resource sharing. It does show cytoplasmic mixing following fusion, now allowing the GFPlabelled actin to diffuse and be incorporated into the actomyosin cable. In a wild-type condition, fusion would not increase the total concentration of resources, although it would increase the total amount of resources within this bigger fused cell. The question is whether resource sharing without increasing the protein concentration is beneficial and increases the efficiency of certain wound healing mechanisms. There might be a benefit of cell fusion, if for example certain resources were only present in limited amounts or if protein transport could increase the concentration locally. To provide better evidence for the hypothesis that resource sharing improves wound healing, maybe the authors could look at the actomyosin cable in a wounded epithelium (such as in Figure 4E, F), in which all cells express MyoII-GFP. The authors could compare the average intensity of the actomyosin cable at the wound edge in mononucleated cells versus in syncytia. If resource sharing is indeed beneficial, it might be that the actomyosin cable is stronger/brighter in syncytia or it forms quicker.

      We agree with the reviewer that we have not "proved the benefit of resource sharing". Because we cannot inhibit resource sharing while still allowing cell fusion, we can think of no rigorous way to test this hypothesis. We appreciate the reviewer's suggestion of quantifying the myosin at the leading edge cable, but we can imagine too many caveats to the interpretation to make it worthwhile. Rather, we accept the limitation that this is an untested, perhaps untestable, hypothesis -- but nevertheless intriguing.

      We do want to clarify ideas about the concentration of resources after fusion. We agree that the overall concentration of a given resource (mass/volume) throughout a syncytium would be the same as the overall concentration in the unfused progenitor cells; however, a syncytium would have a larger total resource mass to direct subcellularly, allowing for local subcellular concentration to be greater in a syncytium vs. an unfused cell. We demonstrate this subcellular localization of actin in a syncytium twice, in Fig. 7C and E (previously Fig. 6C,E), which we think is evidence for increased local concentration.

      The biggest limitation of this study is that the authors don't address how the formation of these syncytia is regulated. While the manuscript in its current form provides some valuable new insights into syncytial-driven wound closure, it would be much more informative if it also provided some mechanistic details. The authors could test if some of the mechanisms shown to regulate syncytial formation in other types of syncytia-driven wound healing are also involved here. E.g. Yorkie was shown to negatively regulate cell fusion in adult syncytial-driven wound closure (Losick et al 2013). The authors could test for the effect of Yorkie-RNAi in the epithelium on wound closure and syncytia formation. Expression of the dominant negative RacN17 also blocked cell fusion in adult syncytial-driven wound closure (Losick et al 2013).

      Moreover, JNK activation was shown to be needed in larval syncytial-driven wound closure (Galko and Krasnow 2004). The authors could test JNK pathway reporters to assess pathway activation or test if the JNK pathway is needed for syncytial-driven wound closure by expressing a dominantnegative form of Basket JNK in the epithelium.

      Or could syncytia formation be regulated by changes in Integrin-mediated adhesion as shown by the Galko lab in Wang et al 2015? They show that wounding provoked a striking relocalization of PINCH and ILK, indicating the disassembly of functional FA complexes concomitant with syncytium formation. Maybe the authors could investigate some of these.

      We investigated the role of JNK in fusion by expressing bsk<sup>DN</sup> on one side of the wound. Comparing the numbers of border-loss fusion on each side, we did not find a significant difference in our seven-sample cohort (see Author response image 1). If we had increased the sample size, we may have found a significant difference with a small effect size, but because of the small difference in fusions on each side we did not think this was worth pursuing. Instead, we include data that the autophagy gene Atg1 is required for cell fusion in new Fig. 4, which begins to address mechanism, and relates the wound-induced fusion described here in pupae to wound-induced fusion shown in larvae. A complete mechanism for wound-induced fusion is outside the scope of this paper, as we focus on the function of syncytia in healing wounds.

      Author response image 1.

      Another general question that the authors raise but don't address enough is whether syncytia-driven wound closure in proliferation-competent epithelia is any different from the one in post-mitotic, polyploid epithelia. Since the mechanism regulating the former is not known, this remains unclear.

      We now include a paragraph on this question in the discussion.

      Finally, it is not clear, whether syncytia in these proliferation-competent epithelia get resolved after wound healing. Do they get removed and replaced by mononucleated proliferation-competent cells or do the syncytia stay in the epithelium like a scar? The authors should provide some images of wound areas a few hours after wound closure is complete and comment on this.

      To answer the reviewer’s question: some but not all syncytia do get removed during wound closure by remarkable apoptotic/extrusion events. This will be the subject of a future manuscript, as it is outside the scope of this paper focusing on the function of syncytia in promoting wound healing.

      Minor points:

      Figure 3: It would be better to have the microcopy images alongside the quantifications.

      The images in Figs. 1 and 2 show the border breakdown and shrinking cells, and we do not see benefit in adding them in Fig. 3.

      Figure 4A: The syncytium at the wound edge here doesn't look straight but wavy. Does it not form an actomyosin cable that straightens the front? Or are there lamellipodia/filopodia?

      We assume the reviewer is asking about the wavy edge outlined at 400 min after wounding (now Fig. 5A). As shown by Jacinto and colleagues in the first pupal wounding paper (JCB 2013), the actin cable forms quickly, within 15 minutes; much later actin protrusions extend from the leading edge to close the wound. This result is consistent with the wavy edge 400 min after wounding.

      248: The authors suggest an interesting hypothesis that mitochondria or ER could be pooled in fused cells. It would be nice to see some evidence: e.g. by labeling mitochondria and assessing where they are in syncytia versus mononucleated cells and whether they are concentrated around the wound edge.

      Although we don't think that exploring mitochondria or ER is central to this manuscript, we agree it would be an interesting question for the future.

      141-145 (Figure 4B and C) This example is not completely convincing. First, it is hard to see where the wound edge is. Second, it would be good to include an even later time point when the cell is clearly no longer at the wound edge.

      We have revised this figure, now Fig. 5B,C, to include a later image at 360 min after wounding healing, and this additional panel clarifies that the smaller cell leaves the wound edge. As noted in the text, the wound edge is indicated by the cell borders lacking p120ctn.

      Reviewer #3 (Public Review):

      Summary:

      White et al. described laser-induced wound healing of the Drosophila pupal notum. They found that the epithelial monolayer is dynamically induced to form syncytia by cell-cell fusion as an important part of repair. They reveal two processes: cell shrinking and border breakage that occur as part of syncytia formation. Expression of GFP in the cytoplasms of some epithelial cells reveals that cytoplasmic contents mix following injury and the GFP rapidly diffuses between cells. Using live imaging they observe that syncytia expand towards the wound, maintain their positions close to the leading edge, and apparently displace smaller cells. They propose that syncytia redistribute cellular components towards the wound facilitating repair and show that labelled actin becomes concentrated at the leading edge.

      Strengths:

      The manuscript is interesting and on an important and emerging topic of wound healing in a genetically tractable organism. The manuscript is very well written.

      Weaknesses:

      There are three major issues that the authors must address: 1. Is cell-cell fusion sufficient to enhance/facilitate wound healing? 2. Characterization of "border breakdown"; Is this phenomenon disassembly of apical junctions following membrane fusion? 3. Are cells really shrinking or is it only the apical domains that "shrink" as the cells join the syncytium.

      We thank the reviewer for recognizing the importance of this topic. Our responses to the specific weaknesses are below.

      Recommendations for the authors:

      Reviewer #1 (Recommendations For The Authors):

      Major Components:

      (1) For syncytia measurements the nuclei are labeled with histone-GFP which is expressed in all cell types. How do you know the nuclei within the cell junctions are epithelial and not another cell type, such as immune cells recruited to the injury site? It would be helpful to verify the number of nuclei per cell using an epithelial-specific nuclear marker as well. This could be via epithelial Gal4-specific expression of a UAS-nls-GFP.

      This is an interesting point. In response to the reviewer's question, we investigated by doing the converse experiment, labeling immune cells with hml-Gal4, UAS-GFP, and observing what they do after wounding (analyzing six wounded pupae). They do get recruited to the wound, but they remain either in the wound center or at the basal side of the leading edge. Because they are labeled with cytoplasmic GFP, we would be able to ascertain whether they fused with epithelial cells because they would share their GFP with epithelial cells in the epithelial plane, and they did not. Thus we are confident that the many syncytial nuclei are not derived from immune cells. Our live tracking throughout the manuscript, and specifically of GFP-labeled clones, also supports our interpretation that syncytial nuclei derive from epithelial cells.

      (2) The manuscript focuses on cell fusion, but other mechanisms of cell enlargement have been observed to occur during wound healing via endoreplication. To what extent do epithelial cells in pupae notum endocycle or endomitosis post injury? It is unclear if the increase in syncytia size during a 1-2hr period could also be due to endomitosis, which would also increase nuclear number.

      Since the first submission of this manuscript, we published our results demonstrating limited wound-induced endoreplication after this type of explosive laser injury to the pupal notum (White et al, 2024, PMID: 38495588). We chose to publish this work separately because we could not offer the same degree of depth for endoreplication as we could for fusion: our pupal notum injury model is extremely well-suited to analyzing cell fusion and wound closure by live imaging; however, it is not particularly well-suited for analyzing endoreplication in fixed tissue. With respect to reviewer's question about endomitosis -- i.e. nuclear divisions that are not accompanied by cell divisions -- even after many years we have not observed an endomitosis event, which would be visible by live imaging, whereas we frequently and easily observe mitosis of diploid cells.

      (3) One of the major conclusions of this study is that cell fusion is necessary to pool resources at the leading edge. Therefore it is critical that authors identify a mechanism to inhibit cell fusion to test this assumption.

      We now include new Fig. 4, an analysis of the role of Atg1 in promoting wound-induced fusion and wound closure. These results build on the finding of the Leptin lab (Kakanj et al, 2022) that autophagy genes are required for fusion. Our results are consistent with the model that syncytia speed wound closure.

      (4) There is evidence that myosin increases in endoreplicating cells during wound healing hence it is, maybe equally - if not more - probable that the increase in resources (here actin-GFP) at the leading edge is dependent on endoreplication instead of cell fusion.

      Some of the new data we provide for this manuscript is a correlation between cell size and distance traveled, showing that larger cells travel more within the wound (Fig. 4F,G). Endoreplication would certainly be expected to contribute to increasing cell size, and our published 2024 data indicates that there can be one extra S-phase induced by these types of wounds. Doubling the genome is not a significant contribution to cell size compared to the 10s of nuclei we observe in syncytia from fusion. Nevertheless, we do not claim that actin is the only important resource that can be pooled subcelluarly for the benefit of the cell; we use it only as a proof-of-principle. Finally, we discuss the work on myosin in wound-induced endoreplicating cells (Losick and Duhaime, 2021).

      Reviewer #3 (Recommendations For The Authors):

      Major comments

      (1) Can induction of epithelial fusion enhance wound healing?

      Different epithelial cell-cell fusion processes have been well-characterized: i) Trophoblast fusion in the placenta mediated by Syncytins. ii) Viral induced cell-cell fusion mediated by diverse viral glycoproteins (e.g. gp41 from HIV, Hemaglutinin from Influenza, GP from Ebola, and G glycoprotein from VSV). iii) Epidermal, myoepithelial, and other epithelial cell-cell fusion in C. elegans mediated by EFF-1 and AFF-1. iv) Cell-cell fusion in the eye lens (unknown fusogens). The authors may want to compare and discuss the temporal dynamics and intermediates observed in the diverse processes of epithelial cell-cell fusion with the characterization of syncytia formation during wound healing of the Drosophila pupal notum. Since some of these characterized cell-cell fusogens can fuse heterologous cells, including Drosophila S2 cells (Shilagardi et al., 2013; https://pubmed.ncbi.nlm.nih.gov/23470732/), the authors may consider expressing these fusogens in Drosophila pupal notum before, during and after injury. This could determine whether syncytia formation is sufficient to stimulate efficient wound healing.

      We thank the reviewer for the suggestion of comparing and discussing temporal dynamics and intermediates observed in the many types of epithelial fusion that are well understood. Regretfully, we do not think this article is the right venue for such a complex discussion, especially since we have little by way of comparison in our own wound-induced fusion data. As for overexpression of fusogens, it is an intriguing idea to force cell fusion with a heterologous fusogen such as EFF-1 and then investigate any resulting changes in wound healing. However, since half the cells within 70 µm of the wound already fuse even without a heterologous fusogen, it seems unlikely we could meaningfully increase the level of cell fusion unless we expressed the fusogen universally, forcing the fusion of nearly all the epithelial cells as well as other cells throughout the body that express pnr-Gal4. Because the overexpression of EFF-1 in C .elegans results in lethality (PMID: 26854231), a widespread induction of fusion would be expected to cause other types of physiological problems that would interfere with the interpretation of wound closure rates. Further, the conditional expression tools in Drosophila allow excellent spatial control, but temporal control is still somewhat low-resolution, so that we would have difficulty expressing EFF-1 before, during, and after wounding at times that would be relevant to understanding wound healing.

      (2) The phenomenon of "border breakdowns" described here is not clear. The authors are probably studying the disassembly of the apical junctions following the initiation of membrane fusion and pore expansion. This should be clarified by using membrane labels to directly observe membrane fusion. Researchers have used electron microscopy and membrane fluorescent probes to follow cell-cell fusion. For example, GPI-mCherry, FM4-64, lipid-modified-GFPs (e.g. PH-domain fluorescently labeled proteins) DiO, DiI, and many others. See for example: Markosyan et al., 2016; https://pubmed.ncbi.nlm.nih.gov/26730950/; Mohler et al., 1998; https://pubmed.ncbi.nlm.nih.gov/9768364/; Meng et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32668210/.

      We agree completely with the reviewer, that border breakdowns represent the disassembly of apical junctions following initiation of membrane fusion and pore expansion. Direct evidence for this order of events is found in the video stills of Figure 1 panel I and video S2, which show that cytoplasmic GFP is transferred to the fusion partner 14 minutes before there is a visible decrease in the apical adherens junction marker p120ctn. The reproducibility of this order of events is documented in Fig. 3: among 107 GFP-labeled cells, 30 of them first visibly shared GFP with a fusion partner, and then 11/30 displayed border breakdown, 16/30 displayed cell shrinking, and 3/30 did not fuse. This last category is consistent with a fusion pore that closed rather than expanded productively. Although we have obtained TEM images of wound-induced fusion pores, these are included in another manuscript currently in revision and so cannot be included here, and further these EM images do not shed light on border breakdown per se, as only live imaging can establish the relationship between border breakdown and pore formation (GFP-sharing).

      (3) The observation of cell shrinking may be misleading. The process the authors describe as "cell shrinking" may involve shrinking of the apical domain, maintaining the cell volume. To clarify this process, the authors may simultaneously label the apical and basolateral domains. It is possible that fusion pore formation occurs in the basolateral, apical, or both domains. The apical shrinking could reflect the migration of the apical junctions following fusion. A similar process has been described in epidermal and vulval cells of C. elegans and other nematodes (Mohler et al., 1998; https://pubmed.ncbi.nlm.nih.gov/9768364/; Sharma-Kishore et al., 1999; https://pubmed.ncbi.nlm.nih.gov/9895317/; Kolotuev and Podbilewicz 2008; https://pubmed.ncbi.nlm.nih.gov/18031720/).

      We thank the reviewer for pointing out these examples of cell fusion in nematodes, and we now compare our findings to Mohler et al, 1998. In Fig. 2D, we specifically investigated what happened to the cell volume of these shrinking cells, and we hope we have now clarified both the text and the annotations on the figure to make our findings more clear. In the X-Z plane, the entire cell volume of two shrinking cells is visible from cytoplasmic GFP labeling. For both cells, the cytoplasmic volume moves laterally into the neighboring syncytia, appearing to initiate the movement from the basal-most area of the cell so that 150 minutes after wounding, both cells have a reduced apical footprint and only a whisp of apically-oriented cytoplasm, with the remainder of the cytoplasm having moved into the syncytia. These images make it clear that fusion is occuring, and that when the apical area disappears the corresponding cytoplasm has also moved into the territory of the neighboring syncytium. In response to the reviewer's suggestion, we did try labeling basolateral domains, but the fluorescent proteins we examined are not restricted to the basolateral domain and are difficult to interpret.

      Minor comments

      (1) Lines 40-43. Repair of injuries has also been observed in non-proliferative syncytial epidermal cells and involves cell-cell fusogens. The authors may want to include this reference: Meng et al., 2020; https://pubmed.ncbi.nlm.nih.gov/32668210/.

      We thank the reviewer for the suggestion, and we have included this reference in the Discussion paragraph about fusogens.

      (2) Lines 128-130. Is "Shrinking fusion" an "artefact"?

      The apical junction shrinks not the cell. I suggest following basolateral membranes to see whether the cell is indeed shrinking as it fuses. The authors may want to share whether the cell volume is maintained but spills into an existing syncytium; the apical junction shrinks because it disappears/disassembles (see also Major comment 3).

      As discussed in Major comment 3, we do provide evidence that the cell cytoplasm spills into an existing syncytium. Perhaps the reviewer finds the term "shrinking cell" to be misleading, as we all agree that the cell contents do not disappear. We have updated the manuscript to use the term "apical shrinking" throughout.

      (3) Lines 157-159. Are these small cells or instead they are small apical junctions? The interpretation should include basolateral domains of the small cells to determine their size! It is also possible that some small cells have fused with the syncytia but on the basolateral domain without apical junction disassembly.

      We appreciate the reviewer's rigor. As noted above, we were not able to analyze the basolateral domains of these cells. Because our all analyses are live-imaging videos, we are able to identify the cells are undergoing apical shrinking and clearly delineate those from stable diploid cells. We now realize that the term "small cells" is confusing and can be mixed up with apical shrinking. These cells are not "small" but normal sized, small only in comparison with the gigantic syncytia around them. We have removed the term "small" from this description.

      (4) Lines 204-206. Many genes required for myoblast fusion in Drosophila have been shown to play a role in different stages of cell-cell fusion. Do they play roles in epithelia fusion during wound closure in the pupal notum?. For example, actin polymerization? Dynamin? Ig-domain and integrin cell adhesion machineries?

      We now provide a new Fig. 4 that shows that the autophagy gene Atg1 reduces wound-induced cell fusion, as it does in larvae (Kakanj et al, 2022), and importantly these wounds close more slowly. We have not analyzed mutants in actin polymerization because we are confident they would interrupt many aspects of wound healing. The Galko lab has identified that integrins suppress wound-induced cell fusion in larval epidermis, but we have not tested these. We have a manuscript in revision demonstrating a requirement for Dynamin and other endocytosis genes in wound-induced fusion, and without dynamin-mediated fusion, these wounds close more slowly.

    1. Author response:

      We sincerely thank the editors and reviewers for their time and thoughtful feedback on our manuscript. The reviewers' constructive comments have been very helpful in guiding our revision plan. Below, we outline our plan.

      In response to Reviewer #1's comments on clarifying the factors that affect image difficulty and categorization rules, we will implement several revisions. First, to clarify what drives image difficulty, we will test whether image typicality within categories, quantified using methods such as Kramer et al. (2023; Sci Adv 9.17: eadd2981), can explain monkey categorization performance. Second, we will also examine whether performance on generalization images depended on their similarity to specific repeated images and on their category typicality. Third, to address whether monkeys and humans apply similar category rules, we will focus on images for which monkeys consistently made errors and examine whether these same images also yielded lower performance (i.e., longer reaction times) in humans.

      Reviewer #1 also raised an important question about how well macaque IT representations and behavior align. The IT categorization performance estimated in our manuscript is currently lower than monkey behavior, but this may reflect the limited number of recorded neurons. We will estimate ceiling IT performance as a function of neuron count and compare it with monkey and human behavior.

      In response to Reviewer #2's suggestion to enhance narrative flow, we will reorganize the text and adjust the ordering of certain figures and sections to ensure smoother transitions between findings and analyses. Specifically, we will more clearly state which parts of the manuscript establish monkeys' categorization ability and which parts compare their behavior with models or humans before performing a triangular comparison across all three.

      Regarding Reviewer #2's suggestion to test DNN performance on control experiments (non-natural stimuli, arbitrary categorization), we agree this is an excellent addition. We will perform these analyses and plan to report the results in the revised manuscript.

      We believe these revisions will substantially strengthen the manuscript and fully address the reviewers' feedback.

    1. Author response:

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study constructed engineered NK-92 cell extracellular vesicles displaying CD19 single-chain variable fragment and evaluated their therapeutic efficacy in MRL/lpr mouse models of systemic lupus erythematosus, demonstrating that these vesicles could deplete B cells, alleviate lupus nephritis, and improve mouse survival. However, this strategy lacks significant innovation compared to existing research. The current results are not sufficient to provide strong support for the experimental hypotheses.

      Weaknesses:

      (1) This study proposes using engineered EVs displaying CD19 scFv to target B cells for SLE treatment. However, similar core therapeutic strategies have been reported in previous studies. For instance, recently, studies have reported engineered EVs for SLE therapy (J Control Release. 2025, 384:113886; Ann Rheum Dis. 2025, 84(11):1811-1821; J Nanobiotechnology. 2026, 24(1):203). Another research team from China also constructed engineered EVs displaying anti-CD19 scFv for SLE treatment, which is highly consistent with the present work in targeting strategy, delivery vehicle, and disease model (Mol Ther. 2026:S1525-0016(26)00080-8). Moreover, the human trial of allogeneic CD19-targeted CAR-NK therapy for SLE has been published (Lancet. 2026, 406(10522):2968-2979). This study has not made original improvements in therapeutic vectors, targeting modules, therapeutic mechanisms, and indications, and thus finds it difficult to meet the requirements of high-level journals for originality and novelty.

      J Control Release. 2025, 384:113886; Ann Rheum Dis. 2025, 84(11):1811-1821; J Nanobiotechnology. 2026, 24(1):203). Another research team from China also constructed engineered EVs displaying anti-CD19 scFv for SLE treatment, which is highly consistent with the present work in targeting strategy, delivery vehicle, and disease model (Mol Ther. 2026:S1525-0016(26)00080-8). Moreover, the human trial of allogeneic CD19-targeted CAR-NK therapy for SLE has been published (Lancet. 2026, 406(10522):2968-2979).

      Reviewer 1 mentioned 4 publications

      (1) J Control Release. 2025, 384:113886; Genetically engineered extracellular vesicles expressing decoy protein TACI provide a therapeutic effect in systemic lupus erythematosus mouse model

      (2) Ann Rheum Dis. 2025, 84(11):1811-1821; J Nanobiotechnology. 2026, 24(1):203)Genetically modified CD19-targeting IL-15 secreting NK cells for the treatment of systemic lupus erythematosus. –but not Evs

      (3) Lancet. 2026, 406(10522):2968-2979) Efficacy and safety of allogeneic CD19 CAR NK-cell therapy in systemic lupus erythematosus: a case series in China。

      (4) Anti-CD19 engineered exosomes enable B-cell targeted anti-BAFF mRNA delivery to alleviate lupus progression”, 

      We sincerely thank the reviewers for their valuable and constructive feedback. We fully acknowledge the important contributions made by the publications cited, and we respectfully submit that they do not invalidate our findings. A critical point to emphasize is that our study employed engineered NK-92 cell extracellular vesicles (EVs) not the cells themselves and we would like to respectfully reiterate the fundamental differences between whole cells and non-cellular EVs, particularly in terms of safety and efficiency profiles. Our safety hypothesis is further supported by the clinical use of inactivated NK-92 cells (as demonstrated in this study: [URL]), which we believe provides a strong and relevant precedent. We are also very grateful that the originality and novelty of our approach have been favorably recognized by Reviewers 2 and 3, which we take as an encouraging validation of our work.

      (2) Numerous core experiments are missing, including the validation of CD19 scFv fusion protein expression on EVs, systematic characterization of engineered EVs, verification of EVs functions and therapeutic mechanisms, and in vitro and in vivo safety assessments. The available data are insufficient to support complete conclusions.

      (3) The stable expression of CD19 scFv on EVs should be further verified by Western blot or flow cytometry. The anchoring of CD19 scFv on the outer membrane surface of EVs must be confirmed. In addition, the loading capacity of CD19 scFv on exosomes should be quantified for the dosage selection in SLE treatment.

      We sincerely thank the reviewers for raising these important points. We note that points (2) and (3) address essentially the same concern, and we fully agree that further validation of CD19 scFv fusion protein expression on EVs is necessary. We are pleased to confirm that we will present additional data on this in due course. Furthermore, we respectfully acknowledge that several other aspects—including the EVs' functions, therapeutic mechanisms, in vitro and in vivo safety profiles, and CD19 scFv loading capacity—remain to be thoroughly investigated. We are committed to addressing these important questions in our follow-up studies, and we hope to provide more comprehensive insights in future work.

      (4) In vitro experiments are required to confirm the specific targeting ability of CD19 scFv-EVs to B cells and clarify the precise mechanism of B cell depletion, particularly whether it is mediated by effector molecules carried by exosomes such as perforin and granzyme B.

      We are most grateful to the reviewer for raising this important point. We are happy to report that we have successfully obtained data demonstrating the specific targeting of CD19 scFv-EVs to B cells, and we will be pleased to include these findings in our revision. With regard to the mechanism of action, we respectfully acknowledge that perforin and granzyme B are recognized as key mediators of NK cell targeting. Nevertheless, we are not aware of any published evidence to date that supports the presence of this same machinery in NK exosomes. We consider this a valuable question for future exploration, and while it lies beyond the scope of the current work, we are diligently investigating it in related ongoing studies.

      (5) The key quality control parameters, such as the stability, purity, buoyant density, and particle/protein ratio of engineered exosomes, should be characterized and identified.

      Agreed, We will provide additional characterization data for the engineered EVs in our revision.

      (6) For the in vivo treatment experiments, the author needs to explain how the treatment dose of CD19scFv-EVs was determined in order to clarify the dose-effect relationship.

      We sincerely thank the reviewer for this valuable suggestion. We fully agree and will be happy to revise the dose calculation accordingly in the updated manuscript.

      (7) It is necessary to supplement with in vivo imaging and tissue distribution data to prove that the CD19 scFv-EVs can specifically accumulate in B-cell organs such as the spleen or lymph nodes. 

      We sincerely thank the reviewer for this valuable suggestion. We fully acknowledge that this is a challenging experiment for several reasons: (1) EV internalization is a rapid process and is therefore difficult to capture; and (2) currently, there is no reliable method available for labeling EVs. Nevertheless, we respectfully assure the reviewer that we will make every effort to attempt this experiment and will report our findings in due course.

      (8) The author needs to clarify the mechanism by which CD19 scFv-EVs reduce B cells in vivo and verify the caspase apoptosis pathway.

      We sincerely thank the reviewer for these valuable comments. We are pleased to confirm that we have successfully demonstrated the specific targeting ability of CD19 scFv-EVs to B cells, and we will gladly incorporate these results in our revised manuscript.

      Regarding the mechanism of action, we fully acknowledge that perforin and granzyme B are well-established mediators of NK cell targeting according to textbook knowledge. However, to the best of our knowledge, there is currently no evidence indicating that NK-derived exosomes are equipped with the same machinery. We respectfully recognize that this is an interesting and important question; while it lies beyond the scope of the present study, we are actively pursuing it in our ongoing parallel work.

      We also appreciate the reviewer's comment regarding the apoptosis pathway. We respectfully note that this aspect was not assessed in any of the publications mentioned by Reviewer 1, which suggests that such analysis may be considered optional rather than mandatory. Nevertheless, we fully agree that this is a worthwhile avenue for further investigation, and we are committed to exploring it in our future studies."

      (9) For the in vivo therapeutic experiments, the clinical first-line drugs and the free CD19scFv should be used to supplement the control group to highlight the advantages of the engineered EVs.

      We sincerely thank the reviewer for this thoughtful and constructive advice. We fully agree that if we were developing this approach for clinical trials, regulatory agencies such as the FDA would require it to demonstrate superiority over current first-line clinical drugs. However, we respectfully wish to clarify that the primary objective of the present study is to provide a proof-of-concept that this strategy is feasible. We fully acknowledge that efficacy and safety will need to be investigated more intensively in future studies before any clinical translation can be considered. We are grateful for this valuable perspective and will be sure to discuss these considerations more explicitly in the revised manuscript.

      (10) Safety assessment in this manuscript is completely absent. Routine toxicity examinations, including hepatic and renal function tests, routine blood tests, and histopathological analysis of major organs in mice, must be supplemented. In addition, the systemic inflammatory cytokine profile and anti-drug antibody levels should be determined to rule out critical safety risks such as cytokine release syndrome and immunogenicity. The authors only focused on alterations in B cells; the impacts of the treatment on T cell subsets, NK cells, and monocytes/macrophages should be further investigated.

      We sincerely thank the reviewer for this valuable advice. We fully agree and will be happy to provide additional data to address this point in our revised manuscript.

      Reviewer #2 (Public review):

      Summary:

      Sun and colleagues report the development of an engineered extracellular vesicle platform derived from NK-92 cells that display an anti-CD19 single-chain variable fragment (scFv) on their surface via fusion with LAMP-2B (V-CD19-Exo). In an MRL/lpr mouse model of SLE, the authors demonstrate that intraperitoneal administration of V-CD19-Exo reduces splenic CD19+CD20+ B cells, attenuates proteinuria and lupus nephritis pathology, downregulates pro-inflammatory cytokines (IL-17A, IFN-γ) and autoantibodies (anti-dsDNA, ANA), and improves survival from approximately 25% to 80%. The authors propose that this "cell-free" targeted extracellular vesicle strategy offers advantages over conventional cell therapies, including lower immunogenicity, scalable production, and no requirement for lymphodepletion.

      The study addresses an important question in autoimmune disease therapeutics: how to achieve targeted B cell depletion while avoiding the complexities and safety risks associated with CAR-T/CAR-NK cell therapies. The concept is novel, and the initial in vivo efficacy data are encouraging. However, several significant limitations in experimental design, mechanistic depth, and evidence rigor temper the strength of the conclusions.

      Strengths:

      (1) Novel conceptual approach.

      The adaptation of CAR targeting principles to extracellular vesicles represents a creative and potentially impactful strategy. By displaying CD19 scFv on NK-92-derived vesicles, the authors successfully confer B cell-targeting capability while retaining the cytotoxic effector functions of the parental NK cells. This "cell-free" concept addresses genuine limitations of live cell therapies, including the need for lymphodepletion, risks of cytokine release syndrome, and manufacturing complexity.

      (2) Comprehensive in vivo efficacy readouts.

      The study evaluates therapeutic effects across multiple clinically relevant endpoints: B cell depletion (flow cytometry), renal function (proteinuria, UPCR), renal histopathology (HE staining with semi-quantitative scoring), systemic inflammation (IgE, IL-17A, IFN-γ), autoantibody production (anti-dsDNA, ANA), and survival. This multi-dimensional characterization strengthens the phenotypic evidence for efficacy.

      (3) Appropriate control groups.

      The inclusion of non-targeted NK92-Exo as a control allows attribution of the observed effects to CD19-mediated targeting rather than non-specific vesicle-associated activities.

      (4) Significant survival benefit.

      The improvement in survival from 25% to approximately 80% in V-CD19-Exo-treated mice is substantial and represents arguably the most compelling evidence for therapeutic potential in this model.

      Weaknesses:

      (1) Mechanism of B-cell reduction remains unclear.

      The manuscript reports a dramatic reduction in splenic CD19+CD20+ B cells (from 10.53% to 1.51%) following V-CD19-Exo treatment. However, the authors do not establish whether this results from direct cytotoxicity (e.g., perforin/granzyme-mediated killing, apoptosis induction) or from functional suppression/downregulation of CD19 expression. The authors speculate that the effect is likely mediated by cytotoxic proteins carried by NK-92-derived vesicles, but no data are provided to support this mechanism. Essential experiments would include the detection of apoptosis markers (Annexin V, activated caspase-3/7) in B cells, assessment of perforin/granzyme B content within V-CD19-Exo, or in vitro co-culture assays demonstrating direct B cell killing.

      We sincerely thank the reviewer for raising this excellent question. We fully agree that it is an important point that truly needs to be addressed. We are pleased to confirm that we have already begun investigating this and hope to obtain meaningful results in due course.

      (2) Small sample sizes.

      Most experimental endpoints were assessed with n=5 per group, which is marginal for detecting modest effect sizes and may amplify the influence of individual biological variation. While the survival study had n=10 per group, the main mechanistic and endpoint analyses would benefit from larger cohorts (n=8-10) to increase statistical power and robustness.

      We are most grateful to the reviewer for this thoughtful and constructive comment. We completely agree that the sample size in our current analysis is somewhat limited for robust statistical evaluation. We are pleased to report that we have since collected additional data, which we will incorporate into our revised manuscript to strengthen the statistical power. If further data become available, we will gladly update them in subsequent revisions.

      (3) No dose-response or dosing optimization studies.

      All experiments used a single dose (10<sup>9</sup> particles per injection) and a fixed schedule (twice weekly for three weeks). The absence of dose-response data leaves unclear whether the observed effects represent maximal efficacy or could be achieved with lower doses, and whether alternative dosing regimens could improve outcomes or reduce potential off-target effects.

      We appreciate the reviewer's thoughtful and important question. We completely agree that this needs to be addressed, and we have already started working on it. We will be pleased to update our data in later comments once further results are obtained.

      (4) Lack of safety assessment.

      The authors emphasize the theoretical safety advantages of extracellular vesicles over cell therapies, but no systematic safety evaluation is presented. Key missing data include: histopathological examination of non-target organs (liver, lung, heart, gastrointestinal tract), assessment of off-target immune activation (T cell responses, cytokine profiles beyond those measured), and evaluation of potential accumulation or toxicity with repeated dosing.

      We appreciate the reviewer's careful and important observations. We fully agree that a systematic safety assessment is necessary.We are actively conducting these experiments and will update our manuscript with the findings as soon as possible.

      (5) Incomplete characterization of the engineered vesicles beyond targeting.

      While the manuscript successfully demonstrates CD19scFv display and vesicle enrichment of exosomal markers, it does not characterize whether V-CD19-Exo retains the full spectrum of NK-92 effector molecules (perforin, granzymes, FasL, TRAIL, cytokines such as IFN-γ) at functional levels. Quantitative or semi-quantitative comparison of cargo between V-CD19-Exo and parental NK-92 cells or non-engineered NK92-Exo would help contextualize the observed in vivo effects.

      We thank the reviewer for this valuable comment. We fully agree that further characterization of the engineered vesicles including NK-92 effector molecules and cargo comparison is needed. We are actively working on this and will update the manuscript as soon as the data become available.

      (6) Sex as a biological variable is not systematically addressed.

      The authors note in the Discussion that the same treatment showed more significant efficacy in male mice compared to females (data not shown), yet all main experiments were conducted exclusively in female mice. Given the strong sex bias in SLE epidemiology (approximately 9:1 female-to-male ratio) and potential differences in immune responses between sexes, this observation warrants systematic investigation rather than a footnote. Presenting the sex-differential data or alternatively, conducting adequately powered sex-stratified analyses would substantially strengthen the manuscript.

      We appreciate the reviewer's important comment. We agree that sex is a relevant biological variable, but a systematic analysis is beyond the current scope. We will consider this for future studies and will acknowledge this limitation in the Discussion.

      (7) Translational claims are premature.

      The manuscript repeatedly emphasizes advantages over cell therapy (low immunogenicity, scalable production, no requirement for lymphodepletion) as if these are established properties of V-CD19-Exo. However, no experiments directly compare V-CD19-Exo to CAR-NK or CAR-T cells in terms of efficacy, immunogenicity, or safety. Similarly, claims of "scalable production" and "high batch-to-batch consistency" are not supported by any manufacturing or quality control data. These statements should be toned down or supported with empirical evidence.

      We thank the reviewer for this important observation. We fully agree that our therapeutic claims are premature without direct comparative and manufacturing data. We will revise the manuscript to temper these statements and present them as potential advantages that warrant future investigation.

      Reviewer #3 (Public review):

      Summary:

      This manuscript describes the development of engineered NK-92-derived extracellular vesicles (EVs) displaying CD19scFv for targeted treatment of systemic lupus erythematosus (SLE). Using a CD19scFv-LAMP2B fusion strategy, the authors generated EVs intended to selectively target pathogenic B cells in the MRL/lpr lupus mouse model. The study reports reductions in CD19⁺CD20⁺ B-cell populations, improvements in proteinuria and renal histopathology, decreased inflammatory cytokines and autoantibody levels, reduced splenomegaly, and improved survival outcomes following treatment. The work aims to position engineered EVs as a cell-free alternative to CAR-T/CAR-NK therapies for autoimmune disease treatment. While the concept is interesting and potentially translational, the study currently lacks sufficient methodological rigor, EV purification standards, mechanistic validation, and comprehensive characterization to fully support many of the claims presented.

      Strengths:

      (1) The study addresses an important unmet clinical need in systemic lupus erythematosus and explores an innovative cell-free therapeutic strategy.

      (2) The concept of combining CAR-like targeting approaches with engineered EVs is interesting and potentially translational.

      (3) The manuscript includes both in vitro and in vivo experiments, including functional renal assessments, immune profiling, histopathology, and survival studies.

      (4) The authors attempt to evaluate multiple disease-associated readouts, including proteinuria, cytokines, autoantibodies, splenomegaly, and survival outcomes, which strengthens the overall biological relevance of the work.

      (5) The use of engineered NK92-derived vesicles as a scalable alternative to CAR-NK therapy represents a potentially attractive therapeutic platform.

      (6) The in vivo therapeutic observations in the MRL/lpr lupus model are encouraging and warrant further mechanistic investigation.

      Weaknesses:

      (1) The EV isolation strategy is not sufficiently rigorous for defining the isolated particles as "exosomes" according to current International Society for Extracellular Vesicles/MISEV guidelines. The precipitation-based workflow without density gradient purification or SEC raises major concerns regarding EV purity and identity.

      We thank the reviewer for this valuable and timely comment. We fully agree that our precipitation-based isolation does not meet MISEV guidelines for defining particles specifically as 'exosomes.' Since our characterization is based on shape, protein markers, and size, we will replace 'exosome' with 'extracellular vesicles' throughout the manuscript to more accurately reflect our methodology.

      (2) No direct validation was provided demonstrating successful surface localization or functional accessibility of CD19scFv on EV membranes.

      We thank the reviewer for this valuable point. We agree, and we are happy to confirm that we have obtained data on surface localization and functional accessibility of CD19 scFv, which we will include in the revision.

      (3) The characterization of EVs is incomplete and insufficient. Additional positive/negative EV markers, purity metrics, and orthogonal characterization methods are required.

      We thank the reviewer for this important point. We fully agree that more comprehensive EV characterization is needed. We are pleased to confirm that we have obtained data on CD19 scFv surface localization and accessibility, which we will include in the revision. We also acknowledge the need for additional markers and purity metrics, and will address this as a limitation in the Discussion.

      (4) The absence of density gradient ultracentrifugation is particularly concerning, given the systemic injection of EV preparations into mice, as contaminating soluble factors and non-vesicular particles may contribute to the observed therapeutic effects.

      We sincerely thank the reviewer for raising this important technical concern. We fully agree that density gradient ultracentrifugation is a more rigorous method for EV purification and that contaminating soluble factors or non-vesicular particles cannot be completely ruled out in our current preparation. We also acknowledge that even with gradient ultracentrifugation, absolute purity is not guaranteed. Nevertheless, we respectfully note that the therapeutic effect of CD19 scFv from EVs was evident when compared to appropriate controls, suggesting that the observed efficacy is attributable at least in part to the EVs themselves. We will add a clear statement of this limitation in the Discussion and will consider more stringent purification methods in our future studies.

      (5) The manuscript lacks adequate mechanistic studies explaining how engineered EVs mediate B-cell depletion or immune modulation.

      We thank the reviewer for this important point. We agree that mechanistic studies would be valuable, but we respectfully note that our current paper focuses on establishing a proof-of-concept. We plan to investigate the mechanisms of B-cell reduction and immune modulation in our future work.

      (6) The in vitro functional assays are weakly designed, particularly the use of A549 cells for evaluating CD19-targeted vesicle function.

      We thank the reviewer for this comment. We wish to clarify that the A549 experiment was intended to confirm that the engineered EVs retain their native function, not to validate CD19 targeting (which will be addressed in point (2). We will revise the manuscript to make this distinction clearer.

      (7) Important methodological details are missing, including EV normalization strategies, flow cytometry gating controls, blinding procedures, and randomization approaches.

      We thank the reviewer for this important observation. We agree that several methodological details were missing. We will reorganize and expand the Methods section to include EV normalization, flow cytometry gating controls, blinding, and randomization procedures.

      (8) Several figures, particularly TEM and western blot images, are of low quality and difficult to interpret.

      We thank the reviewer for this comment. We agree that the TEM and Western blot images are of low quality. We will provide improved, higher-resolution images in the revision

      (9) The study does not sufficiently exclude the possibility that observed therapeutic effects result from contaminating soluble immune mediators rather than EV-specific activity.

      We appreciate this concern. Based on our data, we believe the effects are EV-specific. We will acknowledge this limitation and plan additional controls in future work.

      (10) Broader immune profiling is lacking despite the systemic immune complexity of SLE.

      We thank the reviewer for this important point. We agree that broader immune profiling would be valuable, especially for clinical translation. However, our current study is designed as a proof-of-concept to establish feasibility. We will acknowledge this limitation in the Discussion and plan to address immune profiling in our future work.

      (11) The statistical analysis section includes tests that are not reflected in the Results section, creating concerns regarding data presentation and consistency.

      We thank the reviewer for pointing this out. We agree that the statistical tests in the Methods do not match those in the Results. We will revise both sections to ensure consistency throughout.

      (12) Overall, while the concept is interesting, the manuscript currently falls short of the experimental rigor expected for high-impact translational EV studies.

      We sincerely thank the reviewer for this thoughtful comment. We fully agree that this is a very early-stage translational study, and we acknowledge that considerable work remains before any clinical application can be envisioned. Nevertheless, we respectfully believe that our findings provide a valuable conceptual framework and an initial proof-of-concept that may inform and guide future translational development."

    1. Author response:

      We appreciate the reviewers’ positive assessment of the overall concept and the strength of the wild-type mouse data. We also agree with the main concern raised by the reviewers and editors: the Alzheimer’s disease model findings are more preliminary and should be distinguished more clearly from the stronger conclusions supported by the wild-type data. In the revised manuscript, we will soften the abstract, and discussion to avoid overstating disease-model efficacy, and will frame the AD-model results as suggestive and hypothesis-generating rather than definitive.

      We also plan to address the major methodological and interpretive issues raised in the reviews. We will add sex breakdowns to the figure legends and, where feasible, include sex in the analyses. We will further examine the existing EEG/EMG data to determine which additional sleep bout or spectral analyses can be included, while also clarifying the interpretation of increased dark-phase sleep as a redistribution of sleep and activity rather than a generalized improvement in sleep. We will also clarify PER2::LUC SCN phase analyses and better define the limits of our conclusions regarding central clock strengthening.

      In addition, we will improve the Methods and reporting throughout the manuscript, including clearer information about light conditions, behavioral testing timing, pathology quantification, sample sizes, exclusions or missing data, exact p values, and sex balance. We will also revise the discussion to acknowledge the limitations of the sequential design, the incomplete dissection of individual LiFE components, and the possibility that control wheel access may have reduced the dynamic range for detecting disease-model effects.

      Finally, we will correct and update the references noted by the reviewers and make the requested figure and terminology clarifications.

      Overall, we are encouraged that the reviewers found the study creative, interesting, and potentially important. We believe these revisions will sharpen the claims, improve statistical transparency, and more clearly separate the robust wild-type findings from the preliminary AD-model observations.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This useful study presents an improved protocol for long-term in vitro culture of Schistosoma mansoni that enables progression toward sexually dimorphic stages, representing a meaningful advance for studying parasite development and reducing reliance on animal models. The findings show that host-specific culture conditions support essential developmental and metabolic functions required for parasite maturation, although development remains delayed compared to in vivo conditions. The evidence is solid overall, but limited pairing efficiency and the absence of egg production indicate that the system does not yet fully recapitulate complete reproductive development.

      On behalf of the co-authors, we thank the three reviewers and the editors for their complimentary remarks as well as the major and minor comments/ concerns. Addressing these concerns have led to revisions that improved the manuscript. In particular, further analyses have generated an updated Figures 3 and 4, and Supplementary Tables S1, and S4-S6.

      Public Reviews:

      Reviewer #1 (Public review):

      Pichon, Rémi et al. describe an in vitro method for transforming Schistosoma cercariae into mature adult worms. The authors show that human serum (HS) supports parasite growth and differentiation more effectively than fetal bovine serum (FBS). They also observed differences in parasite growth and activity, with worms cultured in HS efficiently digesting human red blood cells (hRBC). Cultured worms were able to pair with ex vivo adult worms and produce eggs, indicating functional maturation suitable for downstream applications such as drug screening. While the experimental approach is comprehensive and supports the advantage of HS culture conditions, the pairing efficiency was low (≈7%) and required long culture periods (70-80 days), highlighting limitations that may affect reproducibility.

      We acknowledge the reviewer for the positive highlights. Regarding the low in vitro pairing efficiency, we have now edited the manuscript to clarify a misleading statement related to 7%. We decided to remove the value of 7% — which corresponds to the percentage of experiments in which couples were observed, as it does not accurately represent the actual number of observed worm pairs and it is probably misleading. We have updated the text as follows:

      Results, lines 230 ff.:

      “While the establishment of sexual dimorphism was robust and reproducible across more than 15 independent experiments, pairing between male and female parasites was rare. Pairing was observed only in experiments lasting more than 80 days in which we were only able to observe a few couples. In addition, these pairings were temporary (Figures 6A, B; Supplementary Video S4).”

      We also agree with the reviewer that the extended culture periods required to obtain fully sexually dimorphic parasites remain a limitation. As elaborated in Discussion (see below), key factors, probably derived from the host, are missing in the in vitro system explaining both the slow in vitro development and low rate of spontaneous pairing between in vitro developed, sexually dimorphic male and female worms. This was discussed as follows (lines 340-343): “That said, while our system was highly efficient in producing sexually dimorphic worms, spontaneous pairing between male and female parasites was extremely rare, mainly in aged in vitro cultures (from 80 to 100 days in culture) indicating that other factors, e.g., cholesterol, may be missing [35].”

      A major strength of the study, in particular, is that the authors clearly differentiate the effects of FBS versus HS on developmental progression. The conversion rate observed in HS cultures is significant and consistent with previously published data.

      While the study has several strengths, some aspects of the work are not fully explored. In particular, the role of hRBC supplementation requires further clarification. Although HScultured worms were shown to digest hRBC more readily, the implications of this observation remain unclear. Specifically, it would be useful to understand whether hRBC supplementation influences (1) long-term culture stability, (2) molecular pathways associated with development and differentiation, or (3) the pairing capacity of the worms. While addressing these questions may not be the main objective of the study, further discussion of these points would strengthen the manuscript.

      We agree that deciphering the role of the human Red Blood Cells (hRBCs) supplementation is critical. Regarding the influence of hRBCs on the long-term culture stability in parasite development it has been well established for more than four decades that schistosomes do need red blood cells to grow in culture [Basch, P. F. Cultivation of Schistosoma mansoni in vitro. II. production of infertile eggs by worm pairs cultured from cercariae. J Parasitol 67, 186-190 (1981); Basch, P. F. Cultivation of Schistosoma mansoni in vitro. I. Establishment of cultures from cercariae and development until pairing. J. Parasitol. 67, 179-185 (1981)]. The molecular pathways underlying development, sexual differentiation and pairing and modulated by hRBCs in culture is currently being investigated by our team. We decided not to include these data and analyses in the current manuscript, as they fall outside its scope.

      The manuscript is clearly written and represents a valuable contribution to the field. Overall, the experimental approach is sound, and the results support a useful methodological framework for the in vitro culture of Schistosoma worms and the attainment of sexual maturity, particularly for adult male worms.

      We thank the reviewer for highlighting the manuscript’s strengths.

      Reviewer #2 (Public review):

      Summary:

      The authors perform confirmation studies of Paul Basch's seminal schistosome work from 1981, demonstrating the development of transformed schistosomules into sexually dimorphic adult parasites, albeit without successful egg production. In addition to the findings from Basch's earlier work, the authors add some new molecular data in the form of an analysis of proliferative cells in in-vitro-derived animals.

      Strengths:

      The authors successfully confirm experimental results from earlier schistosome researchers, providing a potential new tool for studying schistosome biology without the need for vertebrate hosts.

      We thank the reviewer for highlighting the manuscript’s strengths.

      Weaknesses:

      The display of data from the authors is sometimes difficult to follow/understand where it comes from. For example:

      (1) Line 136: The authors claim that parasites in HS and FBS conditions have substantially different mortality rates (11.3 +/- 2.7 vs 5 +/- 2.3) but a quite high p-value (0.8). Analyzing the raw data myself, I obtained a mean of 8.2 +/- 1.7% vs 4.8% +/- 4.3% with a p-value of 0.15. Either the data are not clearly presented, and I did not follow them, or the data presented in the text do not match the raw data in the supplemental files.

      We thank the reviewer for pointing this out; we have now edited Supplementary Tables S1 and S6 by turning them into a long format for the sake of clarity. Accordingly, Results, Methods sections, and indicated supplementary tables were edited as follows:

      Results, lines 142 ff.:

      “No morphological differences were observed between parasites cultured either in FBS or HS within the first week in culture; in both conditions most parasites were classified as early schistosomula [category 1: 76% ± 30 (average ± SD) in FBS and 73% ± 29 (average ± SD) in HS] with few lung (category 2) and early liver schistosomula (category 3) (Figure 1B, week 1; Supplementary Figure S1). The mean mortality (category 0) at week 1 was slightly higher, but not statistically significant (P= 0.42), in worms cultured in HS [9.75% ± 2.76 (average ± SD)] compared to the mortality registered in FBS-cultured parasites [5.52% ± 5.18 (average ± SD), Supplementary Table S6], consistent with previous findings [39].”

      Methods, lines 463-465:

      “To evaluate differences in mortality between HS- and FBS-cultured parasites, data from 5 experiments were combined and analysed using a Shapiro-Wilk normality test to test normality of the data and a non-parametric Wilcoxon rank sum exact test (Supplementary Tables S1 and S6).”

      Supplementary Tables:

      Supplementary Table S1. “Raw counts of parasites within each developmental stage category. Each row corresponds to a picture of parasites in culture medium containing FBS or HS. Each column corresponds to the raw parasite counts at indicated stage development (categories 0 to 5), time in culture (Time in days - D), and experimental condition.”

      Supplementary Table S6. “Summary of all statistical tests employed in this study. 1. Statistical tests of parasite mortality and the raw data table used for this test. 2. Statistical tests for worm size comparisons (correspond to Figure 2). 3. Statistical tests for worm black gut comparisons (correspond to Figure 3). BG: Black gut. 4. Statistical tests for EdU positive cells comparisons (correspond to Figure 4). Replicate code: E, M and L correspond to day 2, 8 and 15 respectively; R and W correspond to the presence (R) or absence (W) of RBCs added 13 days after transformation.”

      For clarity, below we provide the R script used to perform the statistical tests on the data shown in Supplementary Table S6 (column ‘Raw count of parasite developmental category per image and experiment’)

      Author response image 1.

      (2) Line 187/Figure 4: Though it is not clearly stated, it appears that the authors treat their EdU counts as an ordinal data set of 61 steps (from 0 to >60) rather than a continuous measure of EdU+ cells per animal. In this author's opinion, the graph strongly suggests a continuous data set, and the fact that this reviewer had to dig through poorly-labeled raw data to discover the nature of the data is problematic. The authors should either switch to a continuous data set or make it explicit that the data shown are ordinal. If counting EdU+ cells is too arduous, the authors could consider comparing the amount of EdU+ area to the amount of DAPI+ area in maximum intensity projections of their confocal images, as this would roughly approximate the amount of proliferative cells in the animals.

      As the reviewer correctly pointed out, the data were treated as ordinal because counting worms with more than 60 Edu+ cells became extremely difficult and highly inaccurate. Therefore, we decided to group in a single category, “60 EdU+ cells”, all worms showing more than 60 EdU+ cells. We have now updated Figure 4 where medians are shown instead of media values, Supplementary Table S5 to provide more comprehensive access to the raw counts, and Supplementary Table S6 to indicate the data for EdU+ cells per worm were considered ordinal. Accordingly, we have revised the corresponding sections as follows:

      Results, lines 211 ff:

      “HS-cultured schistosomula showed higher numbers of proliferating stem cells, with a median of >48 and >60 EdU+ cells per worm at days 8 and 15, respectively (Figure 4). On the other hand, most FBS-cultured parasites displayed no more than an average of 20 EdU+ cells per worm (Figure 4).”

      Methods, lines 520 ff:

      “EdU+ cells per parasite were counted for an average of 100 parasites across three independent experiments (Supplementary Table S5). Worms were grouped based on the number of cells per individual, but all those showing ⪰ 60 EdU+ cells were counted in the same group named ‘60 EdU+ cells'. Therefore, the data were considered ordinal data. Statistical analysis was performed by Kruskal-Wallis test with Dunn multiple comparison post-hoc test, with P≤0.05 considered significant (Supplementary Table S6).”

      Figure 4 legend, lines 830 ff:

      “A. Violin plots showing the number of Edu+ cells per worm at indicated time points (2, 8, and 15 days post cercarial transformation) in parasites cultured either in Foetal Bovine Serum (FBS, blue) or Human Serum (HS, light brown). Human Red Blood Cells (hRBCs) were added in the culture at day 13 post cercarial transformation. The small black dots indicate individual worms, and the big black point indicates the median of EdU+ cells per worm. All worms showing ⪰ 60 EdU+ cells were counted and clustered together in the group named ‘60 EdU+ cells’. Hence, the data were treated as ordinal and statistical analysis performed by Kruskal-Wallis test with Dunn multiple comparison post-hoc test, with P≤0.05 (*) considered significant (Supplementary Tables S5 and S6).”

      We thank the reviewer for the very interesting suggestion to quantify cell proliferation by calculating the ratio between EdU+ area to DAPI+ area in maximum intensity projections images. Measuring the fluorescence area for each worm in maximum projection is an excellent idea; however, due to the number of EdU+ cells present in some samples, we think this technique would not provide additional information or produce more detailed data compared with our analysis when the number of Edu+ cells exceeds 60 per worm. We will certainly consider this approximation for future studies.

      There are some minor issues as well:

      (1) Line 122: It is perhaps incorrect to refer to humans as "the" definitive host of schistosomes, as S. japonicum is primarily considered a zoonotic infection with water buffalo/cows being the primary definitive host.

      We thank the reviewer for pointing this out; we have now replaced ‘schistosomes’ with ‘Schistosoma mansoni’ (current line 131)

      (2) Line 185/298: The authors refer to EdU pulse-chase experiments, but the experiments described here are EdU pulse experiments.

      This is a very good point, we thank the reviewer for bringing this up and have accordingly edited by replacing ‘EdU pulse-chase’ with ‘EdU pulse’ experiments in lines 37, 204, and 321.

      Reviewer #3 (Public review):

      Summary:

      This study is significant as it established a protocol for the long-term culture of Schistosoma mansoni newly transformed cercariae, which developed in vitro into sexually dimorphic forms. The impact of two different sera, Fetal Bovine Serum (FBS) and Human Serum (HS), added to the culture medium supplemented with human red blood cells was evaluated. The authors demonstrated that HS-cultured parasites were able to digest red blood cells, a critical step for long-term parasite development. Furthermore, while most FBS-cultured parasites did not progress beyond an early liver stage, sexual dimorphism was clearly evident in the HS-cultured worms, albeit delayed compared to in vivo development.

      Strengths:

      This study could contribute to further in vitro studies for a better understanding of the unique sexual biology of Schistosoma mansoni and for screening novel schistosomicidal compounds. By increasing parasite development in in vitro studies, this protocol could have a positive impact on the principles of the 3Rs (Replacement, Reduction and Refinement) for animal research.

      We thank the reviewer for highlighting the manuscript’s strengths.

      Weaknesses:

      As the authors mentioned, "pairing between male and female parasites was rare. Pairing was observed in approximately ~7% of the experiments, usually after day ~ 80 in culture. Egg production was also not achieved with this protocol.

      Following the reviewer’s point and to clarify a misleading point, we have now decided to remove the value of 7% - which corresponds to the percentage of experiments in which couples were observed. However, this value does not accurately reflect the actual number of observed worm pairs, and it is probably misleading. We have updated the text as follows:

      Results, lines 230 ff:

      “While the establishment of sexual dimorphism was robust and reproducible across more than 15 independent experiments, pairing between male and female parasites was rare. Pairing was observed only in experiments lasting more than 80 days in which we were only able to observe a few couples. In addition, these pairings were temporary (Figures 6A, B; Supplementary Video S4).”

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The manuscript is well-written overall. However, there are some minor revisions that would further improve the clarity and presentation of the data.

      (1) At the beginning of the manuscript, it would be helpful to clearly state three to four specific aims or objectives. This would help readers better understand the expected outcomes and the broader methodological contribution of the study.

      We agree with the reviewer and accordingly have stated the overall goals of the study, as follows:

      Introduction, lines 106 ff:

      “We aimed at optimising a platform to study intra-mammalian schistosomes that supports in vitro sexual dimorphism establishment, consequently leading to an overall positive impact in the 3Rs (Reduction, Replacement, Refinement) for animal research (https://nc3rs.org.uk/) [42]”.

      (2) In the abstract, you highlighted the relevance of the work according to the 3R principles of reduction in animal experimentation. However, this point is not clearly introduced in the Introduction section. Including a short discussion of this aspect would improve continuity and context.

      Following this and previous item raised by the reviewer, we have now clarified the potential impact in the 3Rs by our research outcomes and included that link to the NC3Rs website and a representative reference [Louis-Maerten E, Rodriguez Perez C, Cajiga RM, Persson K and Elger BS (2024). Conceptual foundations for a clarified meaning of the 3Rs principles in animal experimentation. Animal Welfare, 33, e37, 1–11)].

      (3) In line 43, please italicize Schistosoma spp.

      Edited accordingly.

      (4) When discussing the importance of "interfering with sexual development," in line 52, please specify the life cycle stages being referred to.

      Revised accordingly as follows:

      Introduction, lines 54-56:

      “This suggests that interfering with the sexual development of schistosome intra-mammalian stages could potentially restrict human pathology.”

      (5) Between lines 56-58, please rephrase this sentence for clarity.

      We thank the reviewer for this editorial suggestion. The text has been revised as follows:

      Introduction, lines 58 ff :

      “Therefore, novel control strategies are urgently needed, and new targets for drug/ vaccine development became a priority. A better understanding of the mechanisms underlying schistosome development, including sexual dimorphism establishment, will pave the wave to achieve this goal.”

      (6) In lines 66-68 & line 88, please clarify whether the transcriptomic studies cited were performed in vivo, in vitro, or ex vivo, and indicate the developmental stages analyzed.

      We have now included the information suggested by the reviewer as follows:

      Introduction, lines 69-70:

      “Transcriptomic studies, at both bulk [7-11] and single cell [12-1]4 levels for intra mammalian stages in vivo and ex vivo,...”

      (7) Please indicate, in line 110, the day of culture for reference. Without this information, the conversion rates per life cycle stage are difficult to interpret and reproduce. Overall, please try to give an overview in the text of these rates of conversion for context, wherever possible.

      Following the reviewer’s question, we have clearly indicated the in vitro and in vivo timings for ‘conversion’ (understood as sexual dimorphism establishment.) We have written:

      Introduction, lines 117-120:

      “Finally, while most of the FBS-cultured parasites did not progress beyond lung and early liver stage, HS-cultured parasites reached sexually dimorphic stages by week 6, albeit at a slightly delayed rate compared to in vivo development. In the mouse model, parasites become dimorphic by day 21 post-infection (~3 weeks) [12].”

      (8) The section beginning with "Furthermore, phenotypic...cell proliferation" (line 110) may be easier to follow if moved earlier in the Introduction.

      Following the reviewer’s suggestion, we have moved and slightly rewritten the sentence to current line 112, as follows: “First, phenotypic differences between FBS- and HS- cultured parasites became evident as early as 48 hours in culture, with HS-cultured parasites exhibiting higher rates of cell proliferation resulting in larger worms in the HS condition.”

      (9) In line 126, please remove the DOI and add the citation.

      Edited accordingly.

      (10) When referring to 10-week-old parasites, in line 130, please indicate the developmental stage at which they stalled and relate this to the phenotypic scoring shown in Figure 1.

      Based on this suggestion, we have now revised the third paragraph of Results section (‘Sexually dimorphic schistosomes developed entirely in vitro from cercariae’), as follows:

      Results, lines 137 ff.:

      “The development of schistosomula derived from mechanically transformed cercariae was assessed in at least 15 independent experiments, five of which were maintained over a period of at least 10 weeks to assess parasite survival and ability to mate and produce fertile eggs (Figure 1A; Supplementary Table S1).”

      Lines 151 ff.:

      “Differences in parasite development between the two conditions became apparent by week 2 (Figure 1B). At this time point, 14.8% ± 24.9 (average ± SD, excluding dead worms) or 36% ± 33.6 (average ± SD, excluding dead worms) of the parasites cultured in FBS or HS, respectively, have reached category 3, i.e., early liver schistosomulum. Parasites in FBS rarely progressed beyond this stage during the 10-week experiment, with very few parasites (<0.1% ± 0.2, average ± SD) reaching category 4, i.e., late liver schistosomulum. In contrast, worms cultured in HS developed over time across all categories, achieving marked sexual dimorphism by week 6 (13.4% ± 18.6, average ± SD) (Figure 1B; Supplementary Figure S3A), as confirmed by PCR (Supplementary Figure S3B; Supplementary Table S2). No differences in the timing for sexual dimorphism establishment were observed between male and female parasites. The mortality rate of FBS-cultured parasites reached an average of 76.24% ± 23.46 (average ± SD) by week 10, after which the experiments under this condition were stopped as most parasites were dead (Supplementary Figure S2). From that time point onwards only parasites in HS were kept in culture. As previously described for the in vivo development of schistosomes [12], in vitro cultured parasites showed developmental asynchrony in agreement with Basch’s observations [33]; however, by week 10 most of the worms in HS (73.7% ± 25.4, average ± SD) acquired an evident sexual dimorphism (Figure 1B).”

      (11) In line 142, please provide a standard deviation value for the reported average of 14.8%, if available. As well as the absolute numbers of these parasites or indicate them in the supplementary. Otherwise, it is difficult to understand the true conversion rate.

      We followed the reviewer’s suggestions and have now rewritten the text (see above, item 10). In addition, Supplementary Table S1 was edited in long format (see answer for item 1, reviewer #2)

      (12) Please explain, IN line 144, why all cultures were maintained for 10 weeks and provide the rationale for this experimental design.

      We thank the reviewer for this opportunity to clarify this point and hence improve the manuscript. The experimental condition stopped at week 10 included only FBS-cultured worms, not HS-cultured parasites. This is relevant as most of the parasites in FBS were dead by this time, unlike the HS-developed schistosomes. Indeed, some experimental groups consisting of parasites cultured in HS were maintained for up to 22 weeks. We have now updated the text to clarify this point, as follows:

      Results, lines 160 ff.:

      “The mortality rate of FBS-cultured parasites reached an average of 76.24% ± 23.46 (average ± SD) by week 10, after which the experiments under this condition were stopped as most parasites were dead (Supplementary Figure S2). From that time point onwards only parasites in HS were kept in culture.”

      (13) In lines 146-151, please streamline the timelines of culture conditions and observed outcomes in FBS versus HS media. As the current wording makes interpretation difficult.

      Following the reviewer’s suggestion we have streamlined the culture timelines and observed outcomes, as follows:

      Results, lines 137 ff.:

      “The development of schistosomula derived from mechanically transformed cercariae was assessed in at least 15 independent experiments, five of which were maintained over a period of at least 10 weeks to assess parasite survival and ability to mate and produce fertile eggs (Figure 1A; Supplementary Table S1).”

      Results, lines 151 ff.:

      “Differences in parasite development between the two conditions became apparent by week 2 (Figure 1B). At this time point, 14.8% ± 24.9 (average ± SD, excluding dead worms) or 36% ± 33.6 (average ± SD, excluding dead worms) of the parasites cultured in FBS or HS, respectively, have reached category 3, i.e., early liver schistosomulum. Parasites in FBS rarely progressed beyond this stage during the 10-week experiment, with very few parasites (<0.1% ± 0.2, average ± SD) reaching category 4, i.e., late liver schistosomulum. In contrast, worms cultured in HS developed over time across all categories, achieving marked sexual dimorphism by week 6 (13.4% ± 18.6, average ± SD) (Figure 1B; Supplementary Figure S3A), as confirmed by PCR (Supplementary Figure S3B; Supplementary Table S2). No differences in the timing for sexual dimorphism establishment were observed between male and female parasites. The mortality rate of FBS-cultured parasites reached an average of 76.24% ± 23.46 (average ± SD) by week 10, after which the experiments under this condition were stopped as most parasites were dead (Supplementary Figure S2). From that time point onwards only parasites in HS were kept in culture. As previously described for the in vivo development of schistosomes [12], in vitro cultured parasites showed developmental asynchrony in agreement with Basch’s observations [33]; however, by week 10 most of the worms in HS (73.7% ± 25.4, average ± SD) acquired an evident sexual dimorphism (Figure 1B).”

      (14) In lines 153-159, please clarify comparisons between worms cultured in FBS and HS at equivalent time points (e.g., 2 weeks FBS vs 2 weeks HS), rather than comparing only 10 week cultures.

      Following the reviewer’s comment, we have now rewritten the whole third paragraph in Results, under the heading “Sexually dimorphic schistosomes developed entirely in vitro from cercariae” - changes detailed in answers to items 10 and 13 (above).

      (15) It would also be helpful to include information on male versus female development in the context of sexual dimorphism.

      This is a relevant point that we have not clarified in the original submission - we have now indicated in the text that no differences were detected in the timing for male and female dimorphism establishment. New text included as follows:

      Results, lines 159-160:

      “No differences in the timing for sexual dimorphism establishment were observed between male and female parasites.”

      (16) In line 163, please resolve the editing marks and punctuation.

      Resolved accordingly.

      (17) In lines 169 and 172, when referring to stages such as "early liver stage," please indicate the corresponding time in culture (e.g., 3 weeks, 7 weeks + 3 days), or define these stage classifications earlier in the manuscript.

      Following the reviewer’s suggestion we have now included the developmental category after stating ‘early liver stage’, as follows:

      Results, line 187:

      “Even though few parasites in FBS reached the early liver stage (category 3)…”

      (18) Please indicate, in line 173, the developmental stage of worms used when assessing hRBC digestion in HS and FBS cultures. Additionally, here, it would be useful to discuss how hRBC supplementation may influence worm development beyond culture conditions, including possible molecular mechanisms. As a revision, that way maybe you can include data, if already performed or conduct it, to show the effect of adding or not adding hRBC even in HS cultured worms.

      We thank the reviewer for highlighting this important item that warrants further clarification. As stated in Results washed human red blood cells (hRBCs) were added to the culture at day 13. Pilot experiments in which hRBCs were added at different time points had been previously performed; no hemoglobin digestion was apparent when hRBCs were added at days 4, 5 and 6 consistent with previous findings (Correnti JM, Jung E, Freitas TC, Pearce EJ. Transfection of Schistosoma mansoni by electroporation and the description of a new promoter sequence for transgene expression. Int J Parasitol. 2007 Aug;37(10):1107-15. doi: 10.1016/j.ijpara.2007.02.011. Epub 2007 Mar 18. PMID: 17482194.).

      Following this observation, we have added a line to clarify this point, as follows (lines 181187): “Based on both previous reports [45], and pilot experiments in which adding human Red Blood Cells (hRBCs) to the culture before day ~10 did not show obvious haemoglobin digestion, we decided to supplement the culture media with hRBCs at day 13. The addition of hRBCs allowed the parasites to feed and thus continue their development [19]. At this point, they began to swallow and degrade erythrocytes, producing hemozoin, a black pigment derived from host haemoglobin degradation and visible in the worms' intestines.”

      Regarding the specific effect of adding hRBCs in the culture, this is a very good point. First, it has been well established for more than four decades that schistosomes need red blood cells in culture to grow, as example see (Basch, P. F. Cultivation of Schistosoma mansoni in vitro. II. production of infertile eggs by worm pairs cultured from cercariae. J Parasitol 67, 186-190 (1981); Basch, P. F. Cultivation of Schistosoma mansoni in vitro. I. Establishment of cultures from cercariae and development until pairing. J. Parasitol. 67, 179-185 (1981). Second, we are currently analysing transcriptomic data from parasites cultured in different conditions, including in the presence or absence of hRBCs. We decided not to include these data and analyses in the current manuscript, as they fall outside its scope.

      (19) In line 183, please clarify whether the referenced single-cell transcriptomic data were obtained from adult worms.

      We have now clarified this point in the manuscript as follows:

      Results, lines 199 ff:

      “In schistosomes, a complex stem cell system consisting of both somatic and germline stem cells has been described by leveraging recent single cell transcriptomic data across different developmental stages, including schistosomula and adult worms [47].”

      (20) In lines 210 and 213, please indicate the absolute number of worms used for these observations, rather than only percentages. If possible, also report any sex bias in pairing.

      Following this and a similar item raised by reviewer #3 (public review), we decided to remove the mention of 7% given it is misleading. This percentage corresponds to the percentage of experiments in which couples were observed. However, this value does not accurately reflect the actual number of observed worm pairs, and it is probably misleading. We have updated the text as follows:

      Results, lines 230 ff.:

      “While the establishment of sexual dimorphism was robust and reproducible across more than 15 independent experiments, pairing between male and female parasites was rare. Pairing was observed only in experiments lasting more than 80 days in which we were only able to observe a few couples. In addition, these pairings were temporary (Figures 6A, B; Supplementary Video S4).”

      (21) In the final results section, please clarify whether pairing enhances sexual maturation of already mature worms or whether maturation occurs primarily after pairing.

      This is a very relevant point, and we thank the reviewer for giving us the opportunity to clarify it in the manuscript. As described in the manuscript the parasite sexual dimorphism was established in vitro and developed male and female parasites were capable of pairing. Moreover, enlarged oocytes in the ovary’s posterior section of in vitro developed female parasites became apparent after pairing. This observation (Figure 6E, F and Supplementary Video S6) suggests that these female parasites, fully developed in HS-supplemented culture media, were not only capable of pairing, but of starting to fully maturate. We have clarified this aspect in the manuscript as follows:

      Results, lines 243 ff.:

      “Moreover, in vitro developed females coupled with ex vivo collected mature males displayed signs of primordial ovary maturation with larger oocytes towards the posterior region of the ovary (Figure 6E, F; Supplementary Video S6). On the other hand, females developed in vitro but not paired with ex vivo collected males remained immature.”

      (22) Further in the Materials and methods sections, please clarify, isn't 8000 schistosomula/well of a 6-well plate really a confluent culture condition, and does it contribute to NTS mortality in that way, as shown in previous in vitro transformation publications? Please clarify, at least with relative values, percentages of parasite transformation in such a concentrated system.

      No formal titration experiments were carried out but based on empirical observations during pilot experiments we decided to add no more than 8,000 schistosomula per well. This is something to further investigate in the future. We have now added the following sentence in Methods:

      Methods, lines 423-426:

      “The number of parasites cultured per well (~8,000 schistosomula) was determined empirically, as no formal titration experiments were performed. At higher densities (>10,000 per well), more frequent media changes were required, and parasite development appeared to be impaired.”

      (23) Also, what was the rationale of adding hRBCs as early as 13 days post-transformation, when the parasites are in the lung and early liver stage, just forming the guts? Therefore, is it possible that this would have contributed to the observation of lesser parasites disgesting hRBCs? Also, were the hRBC supplemented each time with the media change? This was not clear.

      We thank the reviewer for these questions. The rationale of adding hRBCs at day 13 has been elaborated above (question 18). In addition, in the mouse model, parasites have already migrated through and left the lungs by day 13 post-infection, as described by Nation et al [Nation CS, Da’dara AA, Marchant JK, Skelly PJ (2020) Schistosome migration in the definitive host. PLoS Negl Trop Dis 14(4): e0007951] as follows: “In the mouse, S. mansoni schistosomula begin to arrive in the lungs between 2 and 3 days post-infection, peaking at around day 7 and lasting until around day 11”. Hence, we do not think that adding hRBCs at day 13 contributed to the observation of fewer parasites digesting hemoglobin, because this was only seen in parasites cultured in FBS, not in HS.

      The hRBCs were replaced every two weeks, or sooner if their numbers decreased due to consumption. We have now clarified this point in Methods as follows (lines 427-430): “LTC medium was replaced twice a week and washed human red blood cells (hRBCs) added to a final concentration of 0.02% v/v at 13 days after transformation. Washed hRBCs were replaced every two weeks, or sooner if their numbers decreased due to consumption.”

      (24) In the Discussion, please address the limitations related to the relatively late onset and low frequency of pairing in vitro.

      Following the reviewer’s suggestion and comments from reviewer #1, we have now included a section in Discussion highlighting the limitations of the study and avenues to overcome these in the future.

      Discussion, line 360 ff.:

      “Considering these elements in future experiments will help overcome the limitations encountered in this study, including the low rate of spontaneous pairing between in vitro– developed male and female worms and the requirement for extended culture periods (>70 days). In addition, further research is needed to assess the role of host- and parasite-derived cues in schistosome development.”

      (25) Figure 1: Please consider adding arrows or markers indicating which parasites correspond to the representative developmental stages used for classification.

      We acknowledge the reviewer for the suggestion; however, we respectfully consider this may not be necessary as (1) the images shown in Figure are representative pictures of each time point included for illustrative purposes; (2) Supplementary Figure S1 clearly depicts representative images of worms in each developmental category associated with specific morphological descriptions. For greater clarity we have now added the following text at the end of Figure 1 legend:

      Figure 1 legend, line 810-811:

      “A detailed description of the developmental categories and representative images are provided in Supplementary Figure S1.”

      (26) Figure 2: This plot is somewhat misleading in showing that the HS cultured worms grew significantly more than the FBS worms, where the latter did not grow at all, as also shown by the blue bars all over the plot.

      We appreciate the reviewer’s observation; critically, the data shown in Figure 2 represent measurements of the worm's area, which means that some worms may have become longer but thinner maintaining the same area. Most of the FBS-cultured worms did not develop beyond lung or early liver stages, in which the parasites were long/ thin or shorter/wide, respectively. Therefore, the overall area of these FBS-cultured worms almost did not change (please see the raw data and statistical analyses in Supplementary Tables S3 and S6. We believe that, as presented, Figure 2 is sufficiently clear and self-explanatory. However, we would be happy to consider any suggestions to further clarify this point in the manuscript.

      (27) Figure 3: For panel A, what is the worm percentage corresponding to? The context is missing. Please clarify in the text.

      Following the reviewer’s question and for clarity, we have now (1) modified the axis-legend in Figure 3 as “Percentage of worms displaying or not Black Guts - BG (%)”, and (2) slightly edited the legend as follows:

      Figure 3 legend, lines 820-823:

      “Bar Plot representing the percentage of Human Serum (HS)- or Foetal Bovine Serum (FBS)-cultured schistosomula with (blue bar) or without (light brown bar) black guts (BG) due to the presence of intestinal hemozoin.”

      Reviewer #2 (Recommendations for the authors):

      The authors need to clarify their presentation of data. The raw data needs to be more clearly labeled/explained, and the representation of the data in Figure 4A needs to be explicitly described or changed.

      We acknowledge the reviewer for highlighting this issue related with the data presentation and have decided to follow their advice by editing Figures 3 and 4, and improving the data presentation in Supplementary Tables S1, and S4-S6. In particular:

      Figure 3. We have now modified the axis-legend as “Percentage of worms displaying or not Black Gut - BG (%)”, and slightly edited the legend as follows:

      Figure 3 legend, lines 820-823:

      “Bar Plot representing the percentage of Human Serum (HS)- or Foetal Bovine Serum (FBS)-cultured schistosomula with (blue bar) or without (light brown bar) black guts (BG) due to the presence of intestinal hemozoin.”

      Figure 4. We have edited this figure to show medians instead of media values, and updated the legend as follows: lines 830 ff.:

      “A. Violin plots showing the number of Edu+ cells per worm at indicated time points (2, 8, and 15 days post cercarial transformation) in parasites cultured either in Foetal Bovine Serum (FBS, blue) or Human Serum (HS, light brown). Human Red Blood Cells (hRBCs) were added in the culture at day 13 post cercarial transformation. The small black dots indicate individual worms, and the big black point indicates the median of EdU+ cells per worm. All worms showing ⪰ 60 EdU+ cells were counted and clustered together in the group named ‘60 EdU+ cells’. Hence, the data were treated as ordinal and statistical analysis performed by Kruskal-Wallis test with Dunn multiple comparison post-hoc test, with P≤0.05 (*) considered significant (Supplementary Tables S5 and S6).”

      Supplementary Table S1. We have clarified the data presentation by turning it into a long format and updated the legend accordingly as follows (lines 864-867): “Raw counts of parasites within each developmental stage category. Each row corresponds to a picture of parasites in culture medium containing FBS or HS. Each column corresponds to the raw parasite counts at indicated stage development (categories 0 to 5), time in culture (Time in days - D), and experimental condition.”

      Supplementary Table S4. We have clarified the table by turning it into a long format, simplified the data presentation, and updated the legend accordingly as follows (lines 873874): “Percentage of parasites displaying either black positive (hemozoin) or black negative (no hemozoin) intestine.”

      Supplementary Table S5. We have simplified the table by turning it into a long format, and explained the naming for elements in columns C (‘Group’) and D (‘Replicate’). We have updated the legend accordingly as follows (line 876 ff.): “Raw counting of EdU positive cells per parasite for indicated experimental group, replicate and experiment in long format. The worms were classified by group (column C) and replicate (column D), using the following code: E (‘early’), M (‘medium’) and L (‘late’), corresponding to days 2, 8 and 15, respectively. R and W correspond to conditions with (R) or without (W) human red blood cells, and HS and FBS to culture medium employed.”

      Supplementary Table S6. We have incorporated a new section with the statistical analyses for parasite mortality estimation and updated the legend accordingly as follows (lines 882887): “Summary of all statistical tests employed in this study. 1. Statistical tests of parasite mortality and the raw data table used for this test. 2. Statistical tests for worm size comparisons (correspond to Figure 2). 3. Statistical tests for worm black gut comparisons (correspond to Figure 3). BG: Black gut. 4. Statistical tests for EdU positive cells comparisons (correspond to Figure 4). Replicate code: E, M and L correspond to day 2, 8 and 15 respectively; R and W correspond to the presence (R) or absence (W) of RBCs added 13 days after transformation.”

      Reviewer #3 (Recommendations for the authors):

      The study was well conducted, and the data presented clearly support the conclusions. The protocol is well described, making it reproducible. The pairing experiments could be improved.

      Specific Questions.

      (1) "Male and female adult worms that developed in vivo and recovered from mice by portal perfusion on day 42 post-infection were sorted by sex and placed in culture with worms of the opposite sex developed in vitro (>70 days). Within 24 hours of initiating the co-culturing of in vitro developed worms with ex vivo collected worms, couples were observed".

      In the interest of clarity, and considering that stating ‘worms developed in vivo were collected from infected mice’ is redundant, we have now shortened and edited these lines as follows (lines 238- 242): “Male and female adult worms were recovered from mice by portal perfusion on day 42 post-infection, sorted by sex and placed in culture with worms of the opposite sex developed in vitro. Within 24 hours of initiating the co-culturing of in vitrodeveloped worms with ex vivo collected worms, couples were observed (Figures 6C, D; Supplementary Video S5).”

      (2) Have the authors conducted experiments with in vitro female and male parasites under the same experimental conditions as the in vitro/ex vivo pairing experiments? Is it possible that the tissue culture medium used for the development of sexually dimorphic forms is inhibiting pairing?

      The reviewer raises an interesting point that warrants clarification. First, the experimental conditions tested for in vitro developed parasites were the same as for the pairing experiments, as the ex vivo collected worms were washed and placed in HS-supplemented media. Second, as the culture conditions were the same (same culture protocol and medium) between in vitro pairing and in vitro / ex vivo pairing experiments, we do not think that the tissue culture medium used for developing sexually dimorphic parasites inhibited the pairing. As elaborated in Discussion (see below), key factors, probably derived from the host, are missing in the in vitro system explaining the low rate of spontaneous pairing between in vitro developed, sexually dimorphic male and female worms. This was discussed as follows (lines 340-343): “That said, while our system was highly efficient in producing sexually dimorphic worms, spontaneous pairing between male and female parasites was extremely rare, mainly in aged in vitro cultures (from 80 to 100 days in culture) indicating that other factors, e.g., cholesterol, may be missing [35].”

    1. Author response:

      Reviewer #1 (Public Review):

      Zeng et al.’s work links several key issues in Cryo Electron Tomography in ways that reinforce each other, inspired by the cycleGAN model, leading to very positive results across several benchmark datasets. The related topics include tomogram cleaning and simulations (two crucial areas in the field), with ”spin-off” outcomes in automatic annotation and the completion of the missing wedge. The manuscript covers nearly all essential topics in Tomography, making it very comprehensive and potentially critical in the field. The generalization capabilities on the SHREC 2021 data set are very interesting, although difficult to quantify. I appreciate the approach, but I have serious concerns about some of the limitations of the results presented by the authors.

      We thank the reviewer for the encouraging assessment of our work and for recognizing the potential importance of integrating tomogram denoising and simulation within a unified unsupervised framework. We appreciate the reviewer’s thoughtful evaluation and the concerns raised regarding the limitations of the current results. We address these concerns in detail below and have revised the manuscript to clarify the scope, evaluation strategy, and practical applicability of DUAL.

      (1) Simplified data versus nowadays challenging tomography data. It is acknowledged the difficulty inmaking general tests. In this work, the method shows excellent results on potentially simple data sets (the SHREC 2021, which was used for a benchmark in ET several years ago, but not much used since then) and, even more, the old Relion data set for picking).

      We appreciate the reviewer raising this important point regarding dataset difficulty and relevance. The SHREC 2021 dataset was selected because it is currently the most widely used benchmark simulated dataset for cryo-electron tomography and originates from the last SHREC contest specifically designed for evaluating cryo-ET analysis methods. It provides standardized simulated tomograms with known ground truth structures, which enables objective and reproducible quantitative comparison between different methods. The RELION ribosome dataset is also a commonly used experimental benchmark for evaluating particle detection performance. Nevertheless, we agree that demonstrating performance on additional recent and challenging datasets will further strengthen the evaluation of the method. In response to this comment, we have expanded the experimental evaluation in the revised manuscript by applying DUAL to additional recent cryo-ET datasets to further demonstrate its effectiveness on recent tomograms with more complex biological structures and imaging conditions.

      Specifically, we added an evaluation on the CZII Cryo-ET Object Identification dataset, a popular competition in 2025 with more than 1,000 participants. This experiment complements the original SHREC 2021 and RELION ribosome benchmark results and shows that DUAL can also be successfully applied to more recent cryo-ET data. The quantitative results and representative visual comparisons (shown above in Figure 1 and 2) are provided in the new section 2.6.

      (2) Reproducibility by the average user. I have found many cases in which a specific software producesexcellent results when run by the authors. Still, the average user is lost with the parameters and cannot reproduce these promising results. I propose that the authors address this issue by involving some experimental colleagues and ask them to repeat the work. This is a general concern that applies not only to this work but to many others. I think this consideration is crucial for a field that is growing very quickly and where method development happens at an extraordinary pace... but are all of them generally useful?

      We fully agree with the reviewer that reproducibility and usability are critically important for computational methods in cryo-ET. In response to this concern, we substantially improved the accessibility and reproducibility of the DUAL framework and revised the accompanying documentation to make the implementation easier to inspect and use, as two experimental colleagues have used and reproduced the results. The updated software repository now includes improved documentation, a clearer README, practical tutorials, a method-to-implementation description, a code reference, and example workflows demonstrating how to reproduce the experiments described in the manuscript. We also provide pretrained models together with the configuration files used to generate the results reported in the paper. In addition, the revised documentation clarifies the data interface, domain convention, training workflow, model outputs, and the interpretation of the trained translators. We believe that these improvements will significantly facilitate reproducibility and make it easier for users to apply the method to their own datasets.

      Reviewer #2 (Public Review):

      This study introduces DUAL (Deep Unsupervised simultAneous denoising and simuLation), an unsupervised deep learning framework that jointly addresses denoising and realistic data simulation for cryo-electron tomography (cryo-ET). By leveraging a cyclic, unpaired learning strategy, DUAL avoids reliance on paired clean ground-truth tomograms, which represents a practical advantage over many existing supervised approaches.

      We thank the reviewer for the positive summary of our work and for recognizing the advantages of the unsupervised framework in avoiding reliance on paired ground-truth data.

      Through extensive quantitative evaluations on benchmark datasets, together with qualitative and downstream analyses on diverse experimental tomograms, the authors show that DUAL performs robustly across both denoising and simulation tasks.

      We appreciate the reviewer’s recognition of the robustness of the framework and the evaluation strategy presented in the manuscript.

      If feasible, a limited quantitative or qualitative comparison with one or more recently published deep learning approaches for cryo-ET denoising or simulation, such as CryoSamba, or DeepDeWedge, would further strengthen the evaluation and help contextualize DUAL’s performance.

      We thank the reviewer for this helpful suggestion. As also recommended by the editor, we extended the experiments to include comparisons with recently proposed methods CryoSamba and DeepDeWedge. These comparisons were performed using the same evaluation metrics used in the current experiments so that the results remain directly comparable. The additional comparisons are added into section 2.6.

      Specifically, DUAL was compared with CryoSamba for denoising and with DeepDeWedge for missing wedge compensation on the CZII Cryo-ET Object Identification dataset, a popular competition in 2025 with more than 1,000 participants. The results are shown above in Figure 1 and 2.

      Reviewer #3 (Public Review):

      The paper is titled “DUAL: Deep Unsupervised Simultaneous Simulation and Denoising for Cryo-Electron Tomography.” The authors provided two closely related code branches: one for denoising and one for missingwedge correction. However, I did not find the simulation component. This is important, as the authors state that “the simulation branch provides learning-based cryo-ET simulation to generate synthetic tomograms indistinguishable from experimental ones.”

      We thank the reviewer for carefully examining the released code and for pointing out this source of confusion. We would like to clarify that, in the DUAL framework, simulation and denoising are the two simultaneous branches that are trained jointly, rather than separate sequential modules. The simulation branch learns the transformation from clean/simulated tomograms to realistic experimental cryo-ET tomograms, while the denoising branch learns the reverse transformation from experimental tomograms to the clean domain. Together, these two translators form the cyclic unsupervised learning framework described in the manuscript.

      In the original repository release, the organization of the code may not have made this relationship sufficiently clear, which likely led to the impression that only denoising and missing-wedge correction components were provided. To address this issue, we have substantially revised the repository structure and documentation. The updated repository now explicitly documents the two simultaneous branches of DUAL, explains how the simulation and denoising translators interact during training, and provides clear instructions for reproducing both functionalities. We have also added a dedicated method-to-implementation guide, code reference, and tutorial examples that describe the usage of the simulation component and its role in generating realistic synthetic tomograms that are statistically and visually consistent with experimental cryo-ET data.

      We believe these revisions clarify the implementation of the simulation branch and make the correspondence between the manuscript and the released code substantially easier to understand and reproduce.

      In addition, no pre-trained models were provided. Given that the authors indicate that all training data are publicly available, sharing trained models together with references to the corresponding datasets would significantly facilitate evaluation of the reported performance.

      We agree with the reviewer that providing pretrained models will greatly facilitate reproducibility and evaluation by other researchers. In the revised release of the repository, we have provided pretrained models corresponding to the experiments described in the manuscript together with clear references to the datasets used for training.

      The provided instructions are quite minimal and do not currently support reproduction of the reported findings.

      We appreciate the reviewer highlighting this issue. We have expanded the documentation substantially and provided detailed instructions describing the full workflow required to reproduce the experiments presented in the manuscript. In the revised repository, we added documentation that more explicitly connects the method described in the manuscript with the released implementation. The README summarizes the repository scope and data interface, the tutorial describes the practical workflow for preparing data and running training, and the method and code reference documents describe the mapping between the DUAL formulation and the main implementation files. We believe these additions will make the workflow clearer for users who wish to reproduce or adapt the experiments.

      After many hours of trial, debugging, and experimentation, I was able to train a model for missing-wedge correction using the default parameters, although the process was slow and memory-intensive.

      We thank the reviewer for investing significant effort to test the software and for reporting this observation. Training large 3D deep learning models on cryo-ET volumes can indeed be computationally demanding. We have clarified the computational requirements in the revised manuscript and provide guidance for efficient training and inference.

      Once these points are addressed, I would return to my original request that the authors provide: 3. A fully solved and functional tutorial based on their updated notebooks with all the intermediate results.

      We agree that a comprehensive tutorial will be extremely helpful for users. In the revised repository we have provided a complete end-to-end tutorial demonstrating the workflow from raw tomograms to the final outputs including simulated tomograms, denoised tomograms, and missing-wedge-corrected tomograms.

      We once again thank the editor and reviewers for their insightful comments and suggestions, which have helped us significantly improve the manuscript and the accompanying software.

  3. Jun 2026
    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors provide a resource to the systems neuroscience community by offering their Python-based CLoPy platform for closed-loop feedback training. In addition to using neural feedback, as is common in these experiments, they include a capability to use real-time movement extracted from DeepLabCut as the control signal. The methods and repository are detailed for those who wish to use this resource. Furthermore, they demonstrate the efficacy of their system through a series of mesoscale calcium imaging experiments. These experiments use a large number of cortical regions for the control signal in the neural feedback setup, while the movement feedback experiments are analyzed more extensively. The revised preprint has improved substantially upon the previous submission.

      Strengths:

      The primary strength of the paper is the availability of their CLoPy platform. Currently, most closed-loop operant conditioning experiments are custom built by each lab, and carry a relatively large startup cost to get running. This platform lowers the barrier to entry for closed-loop operant conditioning experiments, in addition to making the experiments more accessible to those with less technical expertise.

      Another strength of the paper is the use of many different cortical regions as control signals for the neurofeedback experiments. Rodent operant conditioning experiments typically record from the motor cortex, and maybe one other region. Here, the authors demonstrate that mice can volitionally control many different cortical regions not limited to those previously studied, recording across many regions in the same experiment. This demonstrates the relative flexibility of modulating neural dynamics, including in non-motor regions.

      Finally, adapting the closed-loop platform to use real-time movement as a control signal is a nice addition. Incorporating movement kinematics into operant conditioning experiments has been a challenge due to the increased technical difficulties of extracting real-time kinematic data from video data at a latency where it can be used as a control signal for operant conditioning. In this paper, they demonstrate that the mice can learn the task using their forelimb position, at a rate that is quicker than the neurofeedback experiments.

      Weaknesses:

      Many of the original weaknesses have been addressed in the revised preprint.

      While the dataset contains an impressive amount of animals and cortical regions for the neurofeedback experiment, my excitement for these experiments is tempered by the relative incompleteness of the dataset.

      As we have responded earlier, we acknowledge that some of the neurofeedback experiments include data from only a single mouse for some cortical regions, while for some cortical regions, there are several animals. This was due to practical constraints during the study, and we understand the limitations this poses for drawing broad conclusions. We felt it was still important to include these data sets with smaller sample sizes, as they might be useful for others pursuing this direction in the future. To address this, we have revised the text to explicitly acknowledge these limitations and clarify that the results for some regions are exploratory in nature. We believe our flexible tool will provide a means for our lab and others to include more animals representing additional cortical regions in future studies. Importantly, we have included all raw and processed data as well as code for future analysis.

      Additionally, adoption of the platform may be hindered by the absence of a tutorial on how to run a session.

      We thank the reviewer for this valuable suggestion. We agree that the absence of clear documentation and tutorials could limit the accessibility and broader adoption of the platform. In response, we have significantly improved the available resources by adding a comprehensive tutorial. Specifically, we have created a dedicated “Wiki” section on the GitHub repository, along with detailed documentation hosted on ReadTheDocs (https://clopy-docs.readthedocs.io). These resources now provide step-by-step guidance on setting up and running a session, along with additional usage examples to facilitate ease of use for new users.

      Reviewer #2 (Public review):

      Summary:

      In this work, Gupta & Murphy present several parallel efforts. On one side, they present the hardware and software they use to build a head-fixed mouse experimental setup that they use to track in "real-time" the calcium activity in one or two spots at the surface of the cortex. On the other side, they present another setup that they use to take advantage of the "real-time" version of DeepLabCut with their mice. The hardware and software that they used/develop is described at length, both in the article and in a companion GitHub repository. Next, they present experimental work that they have done with these two setups, training mice to max out a virtual cursor to obtain a reward, by taking advantage of auditory tone feedback that is provided to the mice as they modulate either (1) their local cortical calcium activity, or (2) their limb position.

      Strengths:

      This work illustrates the fact that thanks to readily available experimental building blocks, body movement and calcium imaging can be carried out using readily available components, including imaging the brain using an incredibly cheap consumer electronics RGB camera (RGB Raspberry Pi Camera). It is a useful source of information for researchers that may be interested in building a similar setup, given the highly detailed overview of the system. Finally, it further confirms previous findings regarding the operant conditioning of the calcium dynamics at the surface of the cortex (Clancy et al. 2020) and suggests an alternative based on deeplabcut to the motor tasks that aim to image the brain at the mesoscale during forelimb movements (Quarta et al. 2022).

      Weaknesses:

      This work covers 3 separate research endeavors: (1) The development of two separate setups, their corresponding software. (2) A study that is highly inspired from the Clancy et al. 2021 paper on the modulation of the local cortical activity measured through a mesoscale calcium imaging setup. (3) A study of the mesoscale dynamics of the cortex during forelimb movements learning. Sadly, the analyses of the physiological data appears incomplete, and more generally, the paper shows weaknesses regarding several points:

      The behavioral setups that are presented are representative of the state of the art in the field of mesoscale imaging/head fixed behavior community, rather than a highly innovative design. Still, they definitely have value as a starting point for laboratories interested in implementing such approaches.

      We agree with the reviewer that the behavioral setup presented here reflects current state-of-the-art approaches in the mesoscale imaging and head-fixed behavior community, and that similar systems have been implemented in other laboratories. However, the primary contribution of our work lies not in introducing a fundamentally new design but in providing a fully open-source, modular, and accessible implementation of such a system. By detailing both the hardware and software components, along with protocols for assembly and use, we aim to lower the barrier to entry for laboratories that may lack the specialized expertise or resources required to develop these systems independently. We hope this accessibility and ease of adoption will facilitate broader use of closed-loop and mesoscale imaging approaches across the field.

      Throughout the paper, there are several statements that point out how important it is to carry out this work in a closed-loop setting with an auditory feedback. Still, sadly there is no "no feedback" control in cortical conditioning experiments. At the same time, there is a no-feedback condition in the forelimb movement study, which shows that learning of the task can be achieved in the absence of feedback.

      We appreciate the reviewer’s insightful comment. We acknowledge that a no-feedback control group was not included in the neurofeedback experiments. This was due in part to the extensive exploration of multiple ROI combinations, as well as preliminary pilot experiments with a no-feedback condition that did not show consistent evidence of learning. Based on these initial results, we chose to prioritize conditions with feedback and did not pursue the no-feedback experiments further. We agree that including such a control would strengthen the study and consider this an important direction for future work.

      The analysis of the closed-loop neuronal data behavior lacks controls. Increased performance can be achieved by modulating actively only one of the two ROIs, this is not really analyzed, while this finding which does not match previous reports (Clancy et al. 2020) would be important to further examine.

      We agree that further analysis of this aspect would strengthen the interpretation of the dataset, and we encourage the community to explore this question using the publicly released data. In our 2-ROI paradigm, we observed that mice often adopt a strategy of predominantly modulating a single ROI to achieve task success, rather than dynamically balancing both regions. This behavior is noted in the manuscript. Importantly, our task design did not impose explicit constraints on the directionality of modulation across ROIs (i.e., increasing one while decreasing the other), in contrast to the paradigm used in Clancy et al. (2020). This difference in task structure may account for the observed divergence in strategies and outcomes.

      Reviewer #3 (Public review):

      Summary:

      The study demonstrates the effectiveness of a cost-effective closed-loop feedback system for modulating brain activity and behavior in head-fixed mice. Authors have tested real-time closed-loop feedback system in head-fixed mice two types of graded feedback: 1) Closed-loop neurofeedback (CLNF), where feedback is derived from neuronal activity (calcium imaging), and 2) Closed-loop movement feedback (CLMF), where feedback is based on observed body movement. It is a python based opensource system, and the authors call it CLoPy. Authors also claim to provide all software, hardware schematics, and protocols to adapt it to various experimental scenarios. This system is capable and can be adapted for a wide use case scenarios.

      Authors have shown that their system can control both positive (water drop) and negative reinforcement (buzzer-vibrator). This study also shows that using the closed-loop system, mice have shown to better performance, learnt arbitrary tasks and can adapt to changes in the rules as well. By integrating real-time feedback based on cortical GCaMP imaging and behavior tracking authors have provided strong evidence that such closed-loop systems can be instrumental in exploring the dynamic interplay between brain activity and behavior.

      Strengths:

      Simplicity of feedback systems design. Simplicity of implementation and potential adoption.

      Weaknesses:

      Long latencies, due to slow Ca2+ dynamics and slow imaging (15 FPS), may limit the application of the system.

      We agree that the latency introduced by calcium dynamics and imaging frame rates is an inherent limitation of calcium imaging–based approaches. Future improvements, including faster calcium indicators, higher frame-rate imaging systems, and more efficient computational pipelines, are expected to mitigate these constraints and enhance temporal precision.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      This version is a substantial improvement from the previous version. My main recommendation is to add a tutorial, with visualizations of some sort, to show how to run a session with the platform. The tutorials for the probe trajectory planner PinPoint is a good example for reference (https://virtualbrainlab.org/pinpoint/tutorial.html).

      We thank the reviewer for this valuable suggestion. We agree that the absence of clear documentation and tutorials could limit the accessibility and broader adoption of the platform. In response, we have significantly improved the available resources by adding a comprehensive tutorial. Specifically, we have created a dedicated “Wiki” section on the GitHub repository, along with detailed documentation hosted on ReadTheDocs (https://clopy-docs.readthedocs.io). These resources now provide step-by-step guidance on setting up and running a session, along with additional usage examples to facilitate ease of use for new users.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Reviewer #1 (Public review):

      Strengths:

      This is an ambitious study that provides a quantitative dissociation of the roles of phasic and tonic pain in adaptive behavior, by integrating ecological neuroscience, motivational theory, and computational modeling. The use of immersive VR combined with a freeoperant foraging task offers a more ecologically valid context to study pain-related behavior compared to traditional paradigms. Furthermore, the study employs a multimodal approach by combining behavioral data, computational frameworks, physiological signals, and EEG. In particular, one of the main strengths of the study is the use of sophisticated computational modeling to capture phasic and tonic pain effects. The experiment codes are available on GitHub, increasing reproducibility.

      We appreciate the reviewers’ recognition of the study’s ambition, the integration of ecological and computational approaches, and our efforts to support reproducibility through open code.

      Weaknesses:

      The main limitations of this article are that it provides insufficient detail on VR implementation. The design of the VR environment is, at this stage, under-described. Crucial information is missing, such as the number of pineapples per block, timing precision, details on how motion is mapped to the virtual movement, etc. This aspect strongly limits the reproducibility of the experiments.

      We thank the reviewer for highlighting the importance of detailed reporting to ensure reproducibility. In response to this valuable feedback, we have taken the following steps:

      (1) Open Access to Software and Data: We have now uploaded the full software and hardware specifications used in our study to a public GitHub repository: https://github.com/ShuangyiTong/PineappleStudy2025ReplicationSoftware. This includes the complete VR implementation, allowing readers to directly experience the task using a commercially available VR headset. The repository also contains the raw data and analysis scripts to facilitate full replication of our results. These links have been updated in “Data and Code Availability” section.

      (2) Expanded Methodological Details: We have revised the Methods section to include the specific details requested, such as:

      (a) The number of pineapples presented per block,

      (b) The temporal resolution and precision of the data collection,

      (c) The mapping between physical motion and virtual movement within the VR environment.

      Specifically, the paragraph containing the changes is following: “At the beginning of each one-minute block, a total number of 150 virtual pineapples of varying heights from 0.33 to 1 m were randomly generated in a circle centred around the participant with a diameter of 6.67 m. Five identical baskets were placed within the space. Spatial locations of trees and vegetation were generated using the game engine's default tree painting tool (Unity Technologies, San Francisco, US).”

      We hope these updates address the reviewer’s concerns and significantly improve the transparency and reproducibility of our experimental design.

      A second limitation lies in the lack of clarity regarding the study hypotheses. Although two overarching hypotheses can be inferred, they are not explicitly formulated. To this end, it is unclear which analyses were merely exploratory, especially for physiological and EEG outcomes.

      We thank the reviewer for this constructive feedback. We agree that making the hypotheses more explicit—particularly regarding the computational framework and the role of physiological measures—strengthens the manuscript. We have significantly revised the final section of the Introduction to explicitly formulate our two primary hypotheses and operationalise the associated behavioural and neurophysiological measures.

      (1) Phasic Pain Hypothesis: We hypothesised that phasic pain serves as a discrete valuation signal that updates the state-action value of specific actions. We predicted this would be evidenced behaviourally by reduced choice probability and increased ‘distance bias’ for pain-associated targets. Neurally and physiologically, we predicted that these aversive values would be tracked by skin conductance responses (SCRs) and the amplitude of pain event-related potentials (ERPs), which serve as established markers for the encoding of aversive magnitude and salience.

      (2) Tonic Pain Hypothesis: We hypothesised that tonic pain acts as a coefficient modulating the trade-off between opportunity cost and vigour cost. This was tested by applying tonic pain to the non-dominant (non-task) limb to ensure that any observed changes were motivational rather than mechanical. We predicted a global reduction in motivational vigour, operationalised as decreased movement velocities and foraging rates.

      By framing the study this way, we clarify that the physiological and EEG outcomes were used to quantitatively test whether the brain and body implement the computations (valuation and vigour-regulation) defined by our model. We have updated the text in the Introduction (see below) to reflect these explicit formulations.

      Updated paragraphs: “Our first hypothesis was that phasic pain provides a distinct valuation signal that updates the value of specific actions within complex environments. In our task, this was implemented by associating specific fruit (distinguishable by colour) with a brief electrical stimulus to the grasping hand, emulating thorns. In our computational model, this was defined as an aversive utility term incorporated into the state-action value evaluation process. We predicted that this computational mechanism would manifest behaviourally as a reduction in choice probability for pain-associated targets and an increase in ‘choice distance bias’ (the willingness to travel further for pain-free options). Neurally and physiologically, we predicted that these aversive values would be tracked by skin conductance responses (SCRs) and the amplitude of nociceptive event-related potentials (ERPs), specifically the N1-P2 complex (Favero et al., 2023).

      Second, we hypothesised that tonic pain acts as a coefficient modulating the tradeoff between opportunity cost and vigour cost, thereby serving a recuperative function. To test this in Experiment 2, we delivered continuous tonic pressure to the non-dominant arm via an inflated cuff to emulate a background state of injury. Within our free-operant framework, tonic pain was modelled as a weighting factor that shifts the optimal balance toward reduced energy expenditure. Because the stimulus was applied to the non-task limb, we specifically predicted a global reduction in motivational vigour—operationalised as decreased movement velocities and foraging rates—rather than a direct mechanical impairment. By applying this formal computational approach, we move beyond exploratory observations to provide a rigorous, mechanism-based explanation for how distinct pain states adaptively govern choice and action.”

      In Experiment 2, the reduction in vigor during tonic pain could plausibly reflect attentional load rather than pain per se. As recognized by the authors, there is no control condition involving an innocuous salient stimulus to rule out non-specific effects of distraction. Perhaps a tonic non-painful but salient somatosensory stimulus (e.g., a strong vibrotactile stimulus applied on the same arm) could have been used as a control stimulus.

      We agree that examining the potential role of attentional load on the interaction between tonic and phasic pain is an important area of future investigation. The inclusion of additional control conditions matched for attentional salience with additional experiments is possible but introduces other confounds related to their different qualities (e.g. a salient vibrotactile stimulus might invigorate behaviour). More fundamentally, attentional processes are a core part of pain function, and should not necessarily be viewed as a confound (i.e. the way that pain mediates some of its core functional effects may directly be through its salient attentional nature). This view is formalised in Wall and Melzack’s classical tripartite model of pain, and distinguishes pain from purely sensory systems such as somatosensation, vision and so on.

      Reviewer #1 (Recommendations for the authors):

      (1) Computational models may be difficult to follow without prior familiarity. Including simplified explanations could make the approach more accessible.

      We thank the reviewer for this constructive suggestion. To make the computational framework more accessible to a broader audience, we have added two new schematic diagrams (Figure 2 and Figure 8) that provide a visual overview of the models used in Experiment 1 and Experiment 2, respectively. These figures illustrate the state-action transitions and provide a clear decomposition of the payoff components—including reward, pain, and temporal costs. We believe these additions significantly clarify the modelling logic and help ground the mathematical descriptions in a more intuitive visual context.

      (2) Lines 220-222: I don't think it is possible to talk about "objective measures of pain" as pain is, by definition, subjective. I suggest rephrasing the sentence.

      We thank the reviewer for this thoughtful observation regarding our terminology. We recognise that the phrase ‘objective measures of pain’ may be misintepreted. Our intention was to highlight the distinction between the internal, reported experience and the behavioural manifestations of pain that our computational method reveals.

      To avoid ambiguity and to better align the text with the core focus of our study, which is the motivational function of pain, we have rephrased the sentence as suggested. We have shifted the emphasis from ‘measuring pain’ to quantifying its specific impact on behaviour.

      Original lines 220-222 have been revised as follows:

      "Taken together, this indicates the composite nature of overall aversiveness and highlights the benefit of combining subjective ratings with model-based measures of its motivational impact on behaviour."

      We believe this revision more accurately reflects our approach of using choice and movement as objective indices of the motivational value of pain.

      (3) The explanation for choosing the foraging task is very interesting, but should be provided in the Introduction rather than in the Methods section. In contrast, the Methods section should include the details of the VR implementation.

      We thank the reviewer for these constructive suggestions regarding the manuscript structure.

      Regarding the rationale for the foraging task: We agree that providing the theoretical justification for the task earlier in the manuscript improves the narrative flow. We have revised the Introduction to explicitly outline why a foraging paradigm was chosen by added the following sentences:

      “A foraging paradigm provides a robust, free-operant framework that captures the core components of adaptive behaviour: it is goal-directed, involves complex movement, and requires the learning of an optimal strategy to maximise rewards. This allows us to computationally dissociate how different types of pain influence the control of action.”

      We believe this addition clarifies the link between our computational hypotheses and the experimental design.

      Regarding the VR implementation: We have updated the Methods section to include the specific experimental parameters requested in the reviewer's previous comments (e.g., timing precision, stimulus counts, and motion mapping) to ensure full reproducibility. However, we have opted not to include the exhaustive engineering details of the underlying software architecture and communication protocols. To ensure complete transparency, the full software and firmware source code, which allows for the exact replication of the environment, is available in our public GitHub repository shown in the code and data availability section.

      (4) It is unclear how the sample size was determined. This information should be included.

      We thank the Reviewer for this comment. For the present study, an a priori power analysis was not conducted due to the novelty of the investigation and the complexity of the analyses. Standard power analyses are not commonly conducted for studies where computational modelling is the primary focus, as results would be potentially misleading. Instead, we based our sample size estimate of N ≈ 30 participants on previous studies using computational modelling of neurophysiological data [6], as well as EEG, SCR and pain studies [7, 8] and studies in our group using combined neurophysiological recordings and VR [9]. This approach represented a pragmatic balance which ensured the credibility of our results and the stability of our model estimates while accounting for the high persubject cost and the depth of the data collected from each individual. This has now been described more accurately in the Method section:

      “An a priori power analysis was not conducted due to the novelty of the investigation and the complexity of the analyses. Instead, we based our target sample size (N ≈ 30 per experiment) on previous studies using computational modelling of neurophysiological data (Mahajan et al., 2025), as well as EEG, SCR, and pain studies (Schulz351 et al., 2015; Zhang et al., 2018), and studies from our group using combined neurophysiological recordings and VR (Hewitt et al., 2026). This approach represents a pragmatic balance that ensures the credibility of the results and the stability of model estimates while accounting for the high per-subject cost and depth of data collected from each individual.”

      (5) Please clarify how / when the monetary performance incentive was provided.

      We thank the reviewer for the opportunity to clarify the incentive structure. The monetary performance incentive is detailed below:

      Participants were informed at the start of the study that they would earn a performance-based bonus of up to £10, determined by the points they collected during the foraging task. To ensure that motivation remained consistent across the entire session for all individuals—regardless of their baseline foraging speed—the specific exchange rate between points and currency was not disclosed. This prevented potential 'ceiling effects', where a high-performing subject might stop exertive effort after reaching the maximum bonus early, or 'floor effects', where a subject might perceive the reward for an individual action as too small to be motivating.

      Following the completion of the experimental session, all participants were compensated with the full £10 bonus in addition to their base payment for participation.

      We have updated the Methods section to reflect these details:

      “Participants were informed at the start of the experiment that their total points would be rewarded with a monetary incentive of up to £10. To maintain a constant level of motivation throughout the task, the exact point-to-currency exchange rate was not specified. Upon completion of the session, all participants were awarded the maximum bonus of £10.”

      Reviewer #2 (Public review):

      Strengths:

      Overall, this study aims to address an important topic and is generally well written.

      We thank the Reviewer for the generally positive evaluation of our work.

      Weaknesses:

      First, phasic pain was induced using electrical stimulation, which typically elicits somatosensory evoked potentials (SEPs). These responses may not reflect pain-specific processes and thus complicate interpretation. This issue bears directly on the study's conclusions, especially when discussing interactions between phasic and tonic pain. For example, tonic pain is known to reduce perceived intensity or cortical responses to phasic pain stimuli delivered elsewhere on the body - an effect not expected for SEPs elicited by electrical stimuli.

      We acknowledge the reviewer’s concern regarding the specificity of evoked potentials elicited by electrical stimulation. We agree that traditional SEPs— particularly those evoked by large surface electrodes—primarily reflect activation of non-nociceptive A-beta fibres and thus may not reliably index pain-specific processes or be modulated by tonic pain via descending nociceptive control. However, we would like to clarify that phasic pain was administered in the present study using small-diameter concentric ‘Wasp’ electrodes. These are comparable to intraepidermal electrodes shown to preferentially activate nociceptive A-delta fibres, thereby eliciting ERPs more closely associated with nociceptive processing rather than mixed somatosensory input [1, 2]. Accordingly, our ERP results demonstrated a reliable increase in N1-P2 amplitude with higher phasic pain intensity, suggesting that the evoked responses captured stimulus-evoked nociceptive processing.

      We acknowledge that these ERPs may still reflect mixed sensory processing and thus may not be fully modulated by tonic pain. Previous studies have shown that ERPs elicited by nociceptive electrical stimulation can be attenuated during tonic pain using cold-water immersion in CPM paradigms [3, 4]. However, these studies typically employ passive tasks, whereas our paradigm involved continuous voluntary behaviour during sustained tonic pressure pain. This difference in task context may engage distinct modulatory systems, possibly prioritising behavioural adaptation over sensory gating.

      We have revised the Discussion and Methods sections to explicitly clarify the electrode design and address the lack of ERP modulation by tonic pain in the context of active behaviour:

      Discussion: “Although we utilised concentric ‘Wasp’ electrodes designed to selectively activate nociceptive A-delta fibres, and confirmed that the resulting ERPs (N1-P2) were significantly modulated by phasic intensity (Figure 6E, F), we observed no such attenuation by tonic pain (Fig. 6G, H).”

      Methods: “These electrodes preferentially activate nociceptive A-delta fibres, thereby eliciting ERPs that more accurately reflect nociceptive processing compared to standard bipolar stimulation (Inui et al., 2002; Mørch et al., 2011).”

      Second, additional control experiments are necessary to rule out alternative explanations. For instance, the authors are suggested to deliver phasic pain to the contralateral arm (e.g., at 1-2 Hz), which might also reduce action velocity. Similarly, tonic pain applied to the grasping hand should be tested to disentangle hand-specific effects.

      We thank the reviewer for these suggestions regarding the spatial configuration of stimuli. The decision to deliver phasic pain to the grasping hand and tonic pain to the contralateral arm was a deliberate feature of our experimental design.

      First, delivering phasic pain to the grasping hand ensured spatial congruency between the virtual stimulus (the fruit) and the physical consequence (the pain). This congruency is essential for subjects to form a coherent representation of the 'painful' object; a contralateral delivery would have introduced a sensory-motor mismatch that could complicate the interpretation of the learning and choice data.

      Second, tonic pain was applied to the contralateral arm specifically to avoid mechanical interference with the grasping action. Applying sustained pressure to the ipsilateral limb would likely have impeded the manual dexterity and fine motor control required to operate the controller buttons. This would have introduced a physical confound, making it difficult to determine if changes in behaviour were due to motivational vigour or simply the mechanical difficulty of performing the grasp while the arm was under pressure.

      We agree that exploring the spatial generalisation of these effects is an important future direction, and we have added a paragraph to the Discussion to clarify these design choices:

      “It is also important to consider the spatial configuration of the stimuli used in this study. Phasic pain was delivered to the grasping hand to maintain spatial congruency with the virtual fruit, ensuring a coherent nociceptive feedback signal for the interactive task. Additionally, tonic pain was applied to the contralateral arm to prevent mechanical interference with motor execution, which would have occurred if pressure were applied to the ipsilateral limb used for grasping the controller. Whilst this design promotes spatial congruency and avoids mechanical confounds, future studies might explore how these effects generalise across different body parts, for which VR experiments serve as a promising tool to test relevant hypotheses (Hewitt et al., 2026).”

      Reviewer #2 (Recommendations for the authors):

      (1) First, the abstract mentions only EEG, yet Experiment 1 employed skin conductance response (SCR) measures while Experiment 2 utilized EEG. Also, the rationale for using SCR in Experiment 1 and EEG in Experiment 2 is not provided and should be explicitly stated.

      We thank the reviewer for identifying the discrepancy between the physiological signals reported in Experiment 1 and Experiment 2. We have revised the Abstract and Methods section to clarify the rationale for these measures.

      In Abstract, the following sentence has been revised: This could be explained by a free-operant computational framework that formalises and quantifies the function of tonic and phasic pain in terms of motivational vigour and decision value, and model parameters correlated with EEG “physiological and neural responses.”

      Regarding the rationale for the measurements, the following sentences were inserted into the Methods section: “Experiment 1 was designed to establish the robust behavioural effects of the foraging task while ensuring the collection of reliable physiological data. We chose SCR as it is a well-validated index of autonomic arousal that we were confident would provide a clear peripheral measure of pain-related processing in this novel VR paradigm.”

      For Experiment 2, we aimed to build on these findings by adding EEG. This was intended as a complementary piece of neural evidence to provide insights into the underlying central neural mechanisms of phasic and tonic pain interactions.

      (2) Second, the quality of both SCR (Figure 3A) and EEG/ERP data (Figure 5A-D) appears compromised by low SNR. For instance, ERP signals show baseline drift at low frequencies, potentially due to movement-related artifacts. The authors are encouraged to enhance data quality and provide cleaner, more interpretable results.

      We thank the reviewer for this observation. We acknowledge that our recordings exhibit a lower SNR compared to conventional, stationary EEG studies. This is a recognized characteristic of Mobile Brain-Body Imaging (MoBI), particularly in immersive VR experiments where participants are physically active [10]. However, previous research has demonstrated that it is possible to recover valid, interpretable neural signals in active settings using modern cleaning methods including trained ICA labels which we have adopted for artefacts cleaning [11]. We also believe we should be restrained from over cleaning the EEG data as pointed out by Delorme in the paper ‘EEG is better left alone’ [12]. Therefore, we have added a new paragraph in the Discussion:

      “It is important to acknowledge that the signal-to-noise ratio in both our physiological and neural recordings is lower than that typically observed in conventional, stationary laboratory experiments (Gramann et al., 2011). This is primarily due to the motion artefacts inherent in an immersive and active virtual reality environment. Whilst we utilised robust cleaning and artefact-correction methods (Klug and Gramann, 2021), the elevated noise floor may limit our capacity to detect more subtle neural effects or interactions. These challenges highlight a critical area for future methodological research, particularly in the development of hardware and signal-processing tools designed to isolate neural signals during complex, mobile behavioural tasks.”

      Another factor contributing to the appearance of the raw signal is the "free-operant" nature of our task. Unlike conventional neurophysiological study paradigms with fixed, sufficient intervals between trials, our participants were free to move and interact with fruit at their own pace. This means that neurophysiological signals from successive actions (e.g., picking up one fruit followed quickly by another) can overlap. For the SCR analysis, we addressed this by using a canonical response function (CRF) to model and "unfold" the overlapping signals with GLM to produce our final results [13]. While we did not perform a similar deconvolution for the EEG data, we focused our analysis on the early, salient components (N1-P2 and early time-frequency changes < 500ms) which are less susceptible to overlap from subsequent actions than the much slower SCR.

      In summary, while significant efforts representing the state-of-the-art approach for MoBI analyses have been taken to minimise the contributions of noise to the dataset, residual noise does remain in the final data. We have employed a combination of robust preprocessing and model-based analytical methods to account for the complexities of a free-operant task. We believe these results represent the best possible balance between signal clarity and the ecological validity of an active foraging task, and we have called for future research to continue improving these tools for immersive VR environments.

      (3) Third, although the authors state that time-frequency analysis was conducted on the EEG data, no corresponding results are presented in Figure 8 or elsewhere. Furthermore, the statistical maps shown appear noisy and require further clarification and possible denoising.

      We thank the reviewer for pointing this out. The time-frequency results are indeed presented in Figure 8 (now Figure 10); however, they are depicted as topographic maps of the t-statistics derived from our LMM rather than raw power change plots.

      The application of EEG to a novel, free-operant task represents a significant methodological development in this study. Unlike conventional EEG experiments where variables are strictly controlled and a "clean" pre-stimulus baseline is easily obtained, our task involves continuous participant engagement and movement. In this context, for the decision-making event, a stable baseline is unattainable as multiple variables, most notably head movements, are constantly in effect.

      Therefore, we believe that presenting the LMM statistical maps in the main text is the most appropriate and rigorous interpretation of the time-frequency results, as these maps represent the signal after accounting for these complex fixed and random effects. This approach was also adopted in previous pain studies [7]. We also updated the figure legend and caption specifically saying that the figure represented correlation between band power and variables we were investigating to improve clarity.

      Second, for more salient stimuli like phasic pain stimulation, we can indeed obtain a highly interpretable time-frequency analysis without further LMM analysis. We have added induced oscillatory responses to phasic pain stimuli to the Supplementary Material (section: Induced oscillatory responses to phasic pain stimuli). The results showed that, consistent with our ERP findings, the intensity of phasic pain significantly modulated induced responses, while the background tonic pain state did not significantly alter the induced oscillatory response to the phasic pain stimulus.

      Regarding the SNR and Denoising Strategy, we acknowledge that the statistical maps appear noisier than those from stationary studies. This is a direct consequence of the lower signal-to-noise ratio (SNR) inherent in mobile VR. Moving EEG from strictly controlled laboratory settings to ecologically valid, "real-world" VR scenarios introduces higher levels of noise, which we believe represents a key frontier for future methodology research. Regarding the denoising process, the maps in the main text represent the data after our full pipeline (including ICA-based artifact rejection and high-pass filtering). Regarding further denoising, we have deliberately chosen not to apply excessive spatial or temporal smoothing [12]. Also, it is important to note that the LMM framework itself serves as a powerful statistical "filter." By including head movement velocity as a regressor and accounting for random intercepts across subjects, the model effectively "cleans" the signal by partitioning out noise components not related to the task conditions.

      Reviewer #3 (Public review):

      Strengths:

      The experimental paradigm is highly innovative. Assessing human behaviour in a naturalistic yet highly controlled setting represents a promising approach to pain research. Notably, assessing pain magnitude implicitly, via its motivational value, offers insights about the overall pain experience that are not usually accessible via common pain ratings.

      Weaknesses:

      Despite these strengths, the manuscript would benefit significantly from more precise definitions of key concepts and an overall clearer, more coherent presentation of its main arguments. The writing, in its current form, often presents claims that are too vague or insufficiently connected with the experimental findings. Moreover, certain aspects of the computational modeling and statistical analysis appear flawed or inadequately justified.

      We thank the Reviewer for the generally positive evaluation of the manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) The analyses presented in the section

      "Results/Additional cost of effort associated with movement" require clearer explanations. The intention here appears to be to assess the association between moving distances and pain intensity to test the hypothesis that the higher the average pain ratings within blocks, the longer the distances moved (i.e., the higher the effort to avoid pain). It is unclear why and how exactly "egocentric distance differences between painful and non-painful fruits" were computed.

      We thank the reviewer for pointing out the need for a clearer definition of the egocentric distance calculation. As the reviewer correctly identified, this analysis tests the hypothesis that subjects would trade off physical effort (distance) for pain avoidance. To compute this, we used a blockwise approach: for each one-minute block, we calculated the average egocentric distance travelled to pick up non-painful fruits and subtracted the average distance travelled to pick up painful fruits. This difference (labelled as "Choice Distance Bias" in Figure 3B) represents the additional effort subjects were willing to exert to reach a pain-free option. We have clarified the computation method and our motivation for using it in the revised text:

      “As shown in Figure 3B, the vertical axis represents the 'choice distance bias', calculated as the difference between the average egocentric distance to non-painful fruits and the average egocentric distance to painful fruits within each block. The egocentric distance is the fruit distance relative to the participant. This metric was computed to test whether subjects would trade off physical effort for pain avoidance; specifically, a positive bias indicates that subjects were willing to bypass closer painful fruits to reach more distant pain-free ones. As hypothesised, we found that as the pain intensity (VAS) of the aversive fruits increased, this distance bias grew significantly, confirming that subjects exerted greater movement effort to avoid higher levels of pain.”

      We have also updated the text in the beginning of " Avoidance increases with increasing phasic pain intensity" section to emphasize the calculation is analysed at the block level to clarify the computation procedure:

      “For this analysis, both aversive choice probabilities and subjective pain ratings were estimated at the block level.”

      (2) In its current form, the explanation of the first optimality equation lacks precision and transparency. Consider the following improvements:

      (a) Precisely define the features that characterize a state/decision point: e.g., i) memory of available options (= set of 7 fruits that were seen but not picked up) and ii) subject's current position, iii) pain intensity associated with green fruit in the current block.

      (b) Precisely define the set of values the action variable a can assume.

      (c) Precisely define the function u(a) in mathematical notation, including its hyperparameters. The fact that a is likely a categorical variable, while u(a) is later described as a sigmoid function (i.e., as a function of a continuous variable), is confusing. In my understanding (see Figure 2F), u is actually a function of the stimulus intensity associated with a given fruit. Since the stimulus intensity depends on the current state s (and varies from block to block), the phasic pain utility function technically also depends on s.

      (d) Precisely define the function d(a) in mathematical notation, including its hyperparameters.

      (e) Precisely describe how the separate horizontal and vertical components of C_m enter the equation.

      (f) Provide a summary of all parameters and hyperparameters being optimized. Are parameters and hyperparameters optimized jointly? What distinguishes parameters and hyperparameters practically?

      We thank the reviewer for this insightful critique. We agree that the original presentation of the optimality equation was insufficiently formal. We have now added a dedicated subsection, "Experiment 1 model summary", which includes a comprehensive table (Table 2) and supporting text to address these points with mathematical precision.

      Specifically, we have implemented the following clarifications in the revised manuscript:

      State and Action Space (a, b): We have formally defined the state s as an ordered memory list M_s of up to 7 items, governed by a FIFO principle. The action a is now explicitly defined as a one-to-one mapping from these memory items to physical reach trajectories.

      Utility and Cost Functions (c, d, e): We have provided the full mathematical notation for the phasic pain utility u(a) and the effort cost d(a). We have clarified that while the choice of fruit (a) is categorical, it serves as an indicator variable that determines the application of a continuous sigmoid utility function based on the block-level pain intensity (x_stim). We have also explicitly decomposed the effort cost into its horizontal (C_h) and vertical (C_v) egocentric components.

      Parameters and Hyperparameters (f): We have clarified that because our model focuses on steady-state motivational trade-offs rather than online learning, the hyperparameters listed are the only variables subject to optimisation. These are fixed for each subject across the duration of the experiment.

      We believe these additions, centred around the new Table 2, provide the transparency and precision requested.

      Furthermore, we would like to clarify a subtle caveat regarding the assumption of a fixed x_stim for the entirety of a block. While participants were aware that green pineapples were aversive, the specific stimulation intensity for a given block was only fully revealed upon picking up the first green pineapple.

      To ensure our model-fitting remains robust despite this 'information lag', we considered several computational alternatives:

      (1) Prior Estimation Modelling: Modelling a participant’s prior estimation of pain stimulation based on previous blocks. We found this unsuitable due to the independent block design and the limited number of trials available to establish a stable prior.

      (2) Data Trimming: Excluding all decisions made before the first green pineapple pickup. While theoretically 'cleaner', this approach introduces significant data imbalance and ignores blocks where a participant—dissuaded by high pain— only picked up a single green fruit before ceasing (approx. 8.75% of blocks).

      Crucially, we performed a sensitivity analysis by re-running the model-fitting procedure using only the data collected after the first green pineapple was harvested in each block. This analysis yielded the same qualitative statistical results as the full-block model presented in the main text. We have added a detailed discussion of this caveat and the alternative study designs we explored (such as pre-block stimulation or stochastic choice paradigms) to the Supplementary Material (Section Discussion of pain intensity information and model robustness). We believe this confirms that our current approach provides a faithful representation of the underlying motivational trade-offs.

      (3) The statistical method selected for assessing the association between decision values and pain ratings is problematic (Figure 2G): Since there are multiple data points from multiple subjects, which introduces dependence between data points, a multilevel instead of a single-level linear regression should be employed.

      We appreciate the reviewer’s suggestion to utilise a multilevel modelling approach. We agree that a single-level regression does not fully account for the nested structure of our data.

      In response, we re-analysed the association using a linear mixed-effects model with a maximal random effects structure. Specifically, we included both random intercepts and random slopes for Ratings grouped by Subject (in R syntax: PainFunc ~ Ratings + (1 + Ratings | Subject)).

      The results of this mixed effect model are consistent with our original findings, showing a significant relationship between decision values and pain ratings (p = .001). We have updated the Figure caption (now Figure 3G) to reflect these multilevel model statistics. We believe this addition addresses the concern regarding data dependence and provides a more rigorous validation of our conclusions.

      (4) The statistical method selected for assessing how decision values/pain ratings relate to SCR coefficients is problematic (Figures 3B and C): Again, a multilevel regression method should be used.

      We thank the reviewer for this important point. We agree that a multilevel approach is more appropriate for our nested data structure, and that the interpretation of the SCR data required more explicit justification in the context of the divergence between decision values and ratings.

      We have now re-analysed the relationship between SCR coefficients (both fixationevoked and shock-evoked), decision values, and subjective ratings using a multilevel (mixed-effects) regression model. This model included random intercepts and random slopes for each participant to account for individual variability. We have updated Figure 4 (previously Figure 3) caption and the corresponding Results and Discussion sections to reflect these findings (revised text are copied to the response to next comment (5) below. This more rigorous approach provided a clearer and more nuanced picture of the data. Specifically, while the simple regression previously suggested that both measures correlated with fixation-evoked SCR, the multilevel model reveals a dissociation: fixationevoked SCR is significantly associated with decision values, but not with subjective ratings.

      (5) The interpretation of the skin conductance analysis results as evidence of "dissociation between expected and experienced utility" is vague and not well-supported given the presented data and statistical shortcomings. The low R2 in Figure 2G already indicates divergence between decision values and pain ratings. It is unclear what the decision values' differential association with shock-evoked SCR coefficients adds to this insight.

      The reviewer correctly notes that the low R^2 in the correlation between decision values and pain ratings (Figure 3G) already suggests a divergence between these two measures. We agree that this is one of the key findings, as it highlights that decision values provide a dimension of pain assessment that is not fully captured by subjective report. However, we believe the SCR results add crucial physiological evidence to explain why and how these measures diverge. The updated multilevel results provide a more concrete double dissociation that aligns with the distinction between decision utility and experienced utility:

      Experienced Utility (Shock-evoked SCR): This measure of physiological arousal during the painful event was significantly predicted by subjective pain ratings (beta = 0.0154, p = .006) but not by decision values (p = .672). This suggests that ratings are more closely tied to the immediate, experienced aversiveness of the stimulus.

      Decision Utility (Fixation-evoked SCR): In contrast, arousal during the period of evaluation/fixation was a significant predictor of decision values (beta = -0.0739, p = .009) but was not significantly associated with subjective ratings (p = .105).

      By using a more rigorous statistical method, we found that decision values are actually a more robust predictor of anticipatory/evaluative arousal (fixation) than subjective ratings are. This supports our interpretation that decision values and ratings capture different temporal and functional aspects of pain processing— specifically, the evaluation of potential outcomes (decision utility) versus the reaction to the outcome itself (experienced utility). We have revised the Discussion to be more conservative regarding the strength of this evidence while clearly articulating how these physiological results provide a mechanistic grounding for the divergence observed in the behavioural data.

      Summary of changes in the manuscript:

      Figure 4 Caption: Updated to report multilevel regression statistics (beta, 95% CI, t, and p-values) instead of R^2 from simple linear regression.

      Results Section: Updated the text to describe the mixed-effects model results, highlighting the dissociation between fixation-evoked and shock-evoked SCRs. Revised text:

      “Analysis using a multilevel linear mixed-effects model revealed a clear dissociation in the relationship between physiological responses and motivational parameters. Fixation-evoked SCR coefficients were significantly associated with decision values, but not with subjective pain ratings (Fig. 4B). Conversely, shock-evoked SCR coefficients showed a significant association with subjective pain ratings, while the association with decision values was not significant (Fig. 4C). This double dissociation suggests a notable divergence between the physiological correlates of expected utility (at the decision level) and experienced utility (the actual pain experience). Taken together, these findings highlight the composite nature of the overall aversiveness of pain and underscore the benefit of combining subjective ratings with model-based measures to capture its distinct impacts on behaviour.”

      Discussion Section: Revised the paragraph discussing decision versus experienced utility to include the "further hint" provided by the divergent SCR correlations.

      Revised text:

      “In our task we get a further hint of this in the SCR measures in experiment 1, whereby a discrepancy exists between decision values and pain ratings in their respective associations with fixation-evoked SCRs and phasic pain-evoked (shock) SCRs. Taken together, this indicates the composite nature of overall aversiveness of pain, and highlights the benefit of combining subjective ratings with model-based measures of its motivational impact on behaviour.”

      (6) When investigating the effects of tonic pain on the neural processing of phasic pain (Figure 5), why were only ERPs analyzed and not induced oscillatory responses?

      We thank the reviewer for this insightful suggestion. We initially focused our analysis on Event-Related Potentials (ERPs) because the N1-P2 amplitude is an established and robust marker in pain research, providing a clear and reliable metric for comparing phasic pain processing across conditions.

      However, we agree that induced oscillatory responses provide a more comprehensive view of cortical dynamics. Following your suggestion, we have performed a Time-Frequency Representation (TFR) analysis at electrode Cz. These results, now included in the Supplementary Material (Figure S4, S5), are entirely consistent with our ERP findings. Specifically:

      Phasic Modulation: Both ERP amplitudes and induced oscillatory power (notably in the theta and gamma bands) were significantly modulated by the intensity of the phasic pain stimulus.

      Tonic Independence: Consistent with the ERP results, the presence of background tonic pain did not significantly modulate the induced oscillatory responses to phasic stimuli.

      We believe this additional analysis significantly strengthens the manuscript by demonstrating that the observed effects are consistent across both phase-locked and non-phase-locked neural domains. We have amended the ERP results section to reflect the addition of induced oscillatory responses in supplementary materials: “We focused our neural analysis of phasic pain on ERPs as phasic stimuli are well characterised by these time-locked evoked potentials. Nevertheless, to ensure a comprehensive assessment of the neural response, we also examined induced oscillatory responses. These results were consistent with the ERP findings and are detailed in the Supplementary Materials (Fig. S4, S5).”

      (7) The explanation of the second optimality equation (involving motivational vigour) requires substantial clarification. Besides the points mentioned for the previous optimality equation, specific opportunities to improve the explanations include the following:

      - In the provided formula, C_v and C_m appear indistinguishable given they are multiplied together, rendering this an ill-posed optimization problem. This should be clarified.

      - In my understanding, d(a)/V_speed corresponds to the temporal delay associated with picking fruit a. Then, what is tau, and why compute the sum tau + d(a)/V_speed?

      - V* is not introduced properly. Is V*(s') = Q*(s', a, tau)? If so, why introduce V*? Moreover, the notational similarity between V_speed and V* is confusing.

      - Gamma = 0 still holds?

      - Summarize all parameters and hyperparameters that are optimized to model the data and more precisely describe the method used for optimization.

      We thank the reviewer for these insightful comments. We agree that the transition from a standard reinforcement learning framework to one incorporating motivational vigour requires precise definitions to ensure the model is well-posed and interpretable. We have addressed these points as follows:

      (1) Clarification of C_v and C_m: We have clarified C_m and d(a) in the newly added Experiment 1 model summary table. Specifically, C_v is the scalar vigour constant and C_m is a unit vector representing the horizontal and vertical components. Because C_m is a unit vector, the optimization does not suffer from a collinearity issue from the scalar multiplication between C_v and C_m.

      (2) Bridging Theory to Practice (tau and Total Delay): In the theoretical framework of Niv et al. (2007), "delay" is an abstract sum encompassing both waiting and execution. In practice, when fitting to real-world VR data with variable execution times , we must distinguish between the waiting time tau (time spent stationary or searching) and the execution time (||d(a)|| / V_speed). This is necessary because participants take time to look around the forest to search for fruits before deciding to commit to an action. The sum tau + ||d(a)|| / V_speed represents the total delay between two actions, which directly aligns with the notion of opportunity cost of time. We have added a table (Table 3) and added a new Figure 8 to clarify these distinctions.

      (3) V*, Q*, and gamma: The reviewer is correct that V*(s') = max_{a’, tau’} Q*(s', a', tau'). We previously used V* for simplicity. Since the notation of V* and V_speed was confusing, we have updated the term to max_{a’, tau’} Q*(s', a', tau') in the optimality equation. We confirm that gamma = 0 (a greedy policy) still holds for the Experiment 2 framework to maintain focus on steady-state motivational trade-offs. We have added this statement to the method section.

      (4) Summary of Parameters and Optimization: We have summarized the hyperparameters {k, x_0, C_p, C_v, h, v} in the new summary table for Experiment 2.

      (8) It is not clear what the results of the modelling approach presented in Figure 7a+b concretely add to the comparison of movement velocities and collection rates in Figure 6.

      We appreciate the reviewer's comment regarding the relationship between the raw behavioral metrics and the computational results. While both sets of findings support the argument for reduced motivational vigour in the tonic pain condition, we believe the modeling approach provides distinct and essential value:

      (1) Finer-Grained Analysis Tool: The computational model acts as a more sophisticated analysis tool than simple velocity or rate averages. Unlike Figure 9a+b (in the revised manuscript, previously Figure 7), which summarizes overall performance, the model accounts for the trial-by-trial trade-off between opportunity costs, movement effort, and choice values. This allows us to isolate vigour from other confounding components.

      (2) Direct vs. Indirect Measurement: If we assume that motivational vigour in a free-operant task can be quantified through an RL framework, as established in animal studies, then the model's vigour constant (C_v) serves as a direct, concrete estimate of that internal state. In contrast, overall speed and collection rates are indirect markers that can be influenced by multiple factors, such as different choice sets available to the participants as the fruits locations are randomly generated.

      In summary, the computational approach provides a rigorous, parameterized bridge between observable behavior and the underlying neuro-computational mechanisms of recuperative pain. We have updated the Discussion section to more explicitly state how the computational approach provides a controlled measure that is isolated from the other confounders of the task. Added text to the Discussion:

      “Compared to overall speed and collection rate, which can be influenced by multiple factors, such as different choice sets available to participants as the fruit locations are randomly generated, the model's fitted parameters (e.g. vigour constant C_v) in theory serves as a direct, concrete estimate of that internal state.”

      (9) Claims made in the discussion should be more thoroughly and closely linked to the results presented previously. Specifically, experimental outcomes supporting the following claims should be directly referenced:

      - "tonic and phasic pain serve different motivational functions".

      - "phasic pain provides a punishment teaching signal that directs avoidance".

      - "tonic pain reduces motivational vigour".

      - "these two functions [punishment teaching signals and reduction of motivational vigour?] can be formally distinguished and quantified".

      - "We did not see interactions between tonic and phasic pain".

      We have revised the Discussion to more explicitly link these claims to our experimental results. Revised text:

      “The experiments show that tonic and phasic pain serve different motivational functions during adaptive behaviour, in line with ecological and evolutionary theories of pain (Bolles and Fanselow, 1980; Walters and Williams, 2019). Specifically, our findings point towards phasic pain providing a punishment teaching signal that directs avoidance through value-based learning, balancing the cost of future harm alongside potential reward. This is supported by the observation that increasing phasic pain intensity significantly reduced choice probability and increased distance bias between choices, whereby participants were willing to travel further to reach a pain-free fruit. In contrast, we found that tonic pain reduces motivational vigour, which supports energy conservation and recuperation in the context of bodily damage. This claim is directly evidenced by the reduction in taskrelated movement velocities and fruit collection rates during tonic pain blocks. The experiments are the first to show that these two functions can be formally distinguished and quantified during ongoing behaviour. By utilising a free-operant RL computational framework, we were able to dissociate these roles phasic pain was quantified as a generally negative utility term affecting choice values, while tonic pain was formalised as a change in vigour constants that were significantly higher (increasing delays between actions) in tonic pain condition. This illustrates how pain simultaneously acts in different ways to serve self-protection.”

      “One notable aspect of our results is that we did not see interactions between tonic and phasic pain at either the behavioural or neural level. Behaviourally, we observed that average aversive choice probabilities remained similar regardless of the presence of tonic pain, with no significant interaction effect on punishment sensitivity. Furthermore, our model-fitting confirmed that tonic pain did not significantly modulate the fitted phasic pain utility values. There are two contexts in which these might be predicted. First, in `conditioned pain modulation' paradigms (Kennedy et al., 2016), a tonic pain stimulus is sometimes seen to reduce both the perceived intensity and the cortical evoked responses to phasic pain stimuli delivered somewhere else on the body (Hoffken et al., 2017; Enax-Krumova et al., 2020). Although we utilised concentric ‘Wasp’ electrodes designed to selectively activate nociceptive A-delta fibres (Inui et al., 2002), and confirmed that the resulting ERPs (N1-P2) were significantly modulated by phasic intensity, we observed no such attenuation by tonic pain. Indeed, neither subjective pain ratings nor the N1-P2 amplitude showed a significant modulation by the tonic pressure pain stimulus. In contrast, our results were more compatible with a trend in the other direction.”

      (10) The paragraph in the discussion "A concern that is sometimes raised..." (lines 243 - 254) raises interesting points, but its particular relevance to the study at hand is unclear.

      We appreciate the reviewer's feedback. The motivation for including this discussion is to address a common critique we received for the study: whether the observed reduction in vigour under tonic pain is "simply" due to distraction or cognitive load, rather than being a specific functional output of the pain system. We have revised this paragraph to link the concern to our paper’s specific finding.

      Our central argument is that for tonic pain, distraction is not a confounding "sideeffect" but rather the primary mechanism of action. By being inherently "distracting," tonic pain successfully withdraws resources from ongoing tasks (like foraging) to promote the energy conservation required for recuperation.

      (11) The clinical perspective of the methodological framework presented at the end of the discussion is interesting and could be expanded.

      We thank the reviewer for this encouraging comment. We have expanded the final paragraph of the Discussion to more explicitly state the clinical utility of our framework. Specifically, we now contrast our approach with standard clinical assessments such as Quantitative Sensory Testing (QST). We highlight that while QST is a valuable tool, it can lack ecological validity; in contrast, our VR-based task allows for a more realistic, behaviourally sensitive assessment of how pain impacts a patient’s daily functional activities and motivational state. We believe this represents a significant step towards more objective and "real-world" clinical pain phenotyping.

      (12) The statistical analyses part in the methods section should provide a clear definition of dependent and independent variables and clearly state which test was used for which analysis, e.g., by referencing the corresponding subfigure in the main text.

      We agree that a more structured summary of the statistical approach would improve the clarity of the Methods section. We have now included a comprehensive summary table (Table 1) in the Statistical Analysis subsection. This table explicitly defines the dependent and independent variables for each analysis, identifies the specific statistical model used (e.g. Linear Mixed Models or repeated measures ANOVA), and directly maps these to the corresponding figures in the results section.

      Minor comments:

      (1) Introduction:

      (a) The introduction should elaborate more on the advantages of employing an "ecologically meaningful context".

      We thank the reviewer for suggesting further elaboration on the advantages of employing an "ecologically meaningful context". We have updated the introduction to provide additional reasoning of choosing an ecologically valid context for the study:

      “One of the challenges in studying adaptive functions of pain is the difficulty of embedding experiments within ecologically meaningful contexts. To solve this, we designed an immersive foraging task using virtual reality (VR), in which humans search a forest to collect fruits from the low-lying bushes at varying heights. A foraging paradigm provides a robust, free-operant framework that captures the core components of adaptive behaviour: it is goal-directed, involves complex movement, and requires the learning of an optimal strategy to maximise rewards. This allows us to computationally dissociate how different types of pain influence the control of action.”

      (b) It would be helpful to clarify why tonic pain applied to a limb not involved in the task is expected to influence the motivational vigour with respect to the task.

      We thank the reviewer for pointing out additional clarification for applying tonic pain to the non-dominant arm. We have added the following text to the introduction clarifying our hypothesis and why it was applied to the non-task limb:

      “Second, we hypothesised that tonic pain acts as a coefficient modulating the tradeoff between opportunity cost and vigour cost, thereby serving a recuperative function. To test this in Experiment 2, we delivered continuous tonic pressure to the non-dominant arm via an inflated cuff to emulate a background state of injury. Within our free-operant framework, tonic pain was modelled as a weighting factor that shifts the optimal balance toward reduced energy expenditure. Because the stimulus was applied to the non-task limb, we specifically predicted a global reduction in motivational vigour—operationalised as decreased movement velocities and foraging rates—rather than a direct mechanical impairment.”

      (2) Results/Experiment 1:

      (a) How were monetary rewards implemented exactly? How much money per fruit?

      We thank the reviewer for the opportunity to clarify the incentive structure. Participants were informed at the start of the study that they would earn a performance-based bonus of up to £10, determined by the points they collected during the foraging task. To ensure that motivation remained consistent across the entire session for all individuals—regardless of their baseline foraging speed—the specific exchange rate between points and currency was not disclosed. This prevented potential 'ceiling effects', where a high-performing subject might stop exertive effort after reaching the maximum bonus early, or 'floor effects', where a subject might perceive the reward for an individual action as too small to be motivating.

      Following the completion of the experimental session, all participants were compensated with the full £10 bonus in addition to their base payment for participation. We have updated the Methods section to reflect these details:

      “Participants were informed at the start of the experiment that their total points would be rewarded with a monetary incentive of up to £10. To maintain a constant level of motivation throughout the task, the exact point-to-currency exchange rate was not specified. Upon completion of the session, all participants were awarded the maximum bonus of £10.”

      (b) A green pine apple is not ripe and, in a naturalistic context, possesses some aversive value, even in the absence of phasic pain stimuli. Why was the color coding not counterbalanced across individuals? To what degree could this have confounded the results?

      We thank the reviewer for this insightful point. We acknowledge that the lack of counter-balancing for fruit colour (green vs. yellow) is a limitation of the current study design. However, we believe the potential confounding effect of "unripe" green pineapples on the final analysed data is minimal due to the principles of associative learning.

      While a naturalistic heuristic (green = unripe) might establish a weak prior bias, fundamental associative learning [14] and reinforcement learning models [15] demonstrate that extensive training with a highly salient unconditioned stimulus (such as pain) rapidly overrides mild initial priors. The task objective focused strictly on maximizing reward points, and participants underwent extensive training (10 blocks in Experiment 1; 6 blocks in Experiment 2) before the analysed sessions began. During this time, the strong, explicit contingencies (green = pain, yellow = safe) were learned and verbally verified. Therefore, by the time the main experimental data was collected, any weak baseline aversion to green had been overshadowed by the explicit task contingencies, making the learned associative value the primary driver of behaviour. We have added a statement acknowledging this limitation and outlining this theoretical rationale in the Methods section.

      “While the colour association (green for painful, yellow for pain-free) was not counter-balanced across subjects, any inherent aversive value of green pineapples (e.g., as 'unripe' fruit) is expected to have a minimal confounding effect on the analysed data. In associative learning frameworks, while mild prior biases may influence initial value estimations, extensive training with a highly salient unconditioned stimulus (e.g. phasic pain) rapidly updates these values, driving them toward an asymptote determined entirely by the explicit task contingencies (Rescorla & Wagner, 1972; Sutton & Barto, 2018). Because participants underwent extensive training (10 blocks in Experiment 1 and 6 blocks in Experiment 2) to establish the explicit pain associations prior to the analysed sessions, the observed avoidance behaviour was predominantly driven by the learned phasic pain contingencies rather than baseline colour preferences.”

      (c) In the "Avoidance increases with increasing phasic pain intensity" section, clarify upfront that pain ratings and choice probabilities were estimated at the block level. This information is provided only in a later section.

      We agree with the reviewer that this information should be stated earlier for clarity. We have updated the beginning of the "Avoidance increases with increasing phasic pain intensity" section to specify that these metrics were estimated at the block level:

      “For this analysis, both aversive choice probabilities and subjective pain ratings were estimated at the block level.”

      (3) Results/Experiment 2:

      (a) ERP visualizations (Figure 5) should include standard error indicators.

      We have updated Figure 5 (now Figure 6) to include 95% confidence intervals for standard error of the mean across subjects for all ERP traces. This provides a clearer visualization of the variance in the neural response.

      (b) In the section "A unified model...", clarify what is meant by saying that the unified model is "validated by the behavioural data", since behavioral data is what is being modeled in the first place.

      We clarify that "validation" in this context refers to the consistency between the parameters estimated by our generative unified model and the results obtained from the independent, model-free regression analysis of the raw behavioural data. While both approaches use the same source data, the unified model provides a finer-grained analysis of latent internal states (like motivational vigour), whereas the regression provides a direct empirical benchmark (more details were discussed in the response to major comment (8)). We have rephrased this section to better describe this as a consistency check against empirical regression results.

      (c) In the context of Figure 8a, the term "correlations" is misleading if referring to pairwise comparisons.

      We appreciate the opportunity to clarify our terminology. The results presented in Figure 8a (and the associated text) are derived from a Linear Mixed Model (LMM) where the tonic pain condition was treated as a binary independent variable. The term "correlation" was used to describe the statistical association (represented by the t-values) between the presence of tonic pain and EEG band power, accounting for subject-level random effects. It does not refer to simple pairwise comparisons (like t-tests). However, we agree that "correlation" can be ambiguous when applied to a binary predictor. We have revised the text and figure legends to use the terms "associated with" or "predicted by" to more accurately reflect the LMM framework.

      (d) Based on the presented data, there is no evidence for the section headings claim "Neural activities link to vigour".

      We agree with the reviewer that our results primarily provide evidence for a significant neural association with the tonic pain condition rather than a direct, statistically robust correlation with the vigour parameter itself (after Bonferroni correction). While tonic pain is associated with reduced vigour behaviourally, the EEG markers we identified are more accurately described as signatures of the pain state. We have revised the section heading and the corresponding text to focus on the characterisation of the tonic pain state to ensure our claims are strictly supported by the statistical evidence.

      (4) Methods:

      In the supplementary materials, the headings pertaining to different LMMs are confusing and not consistent with the Figure labeling in the manuscript (e.g., 4(ii)b likely corresponds to Figure 4d).

      We thank the reviewer for identifying these inconsistencies in the supplementary material. We apologize for the confusion caused by the labelling errors during reformatting the manuscript. We have now thoroughly audited the supplementary headings and updated them to ensure they correspond directly and consistently with the figure labels in the main manuscript.

      References

      (1) Inui, K., Tran, T. D., Hoshiyama, M., & Kakigi, R. (2002). Preferential stimulation of Adelta fibers by intra-epidermal needle electrode in humans. Pain, 96(3), 247–252. https://doi.org/10.1016/S0304-3959(01)00453-5

      (2) Mørch, C.D., Hennings, K. & Andersen, O.K. Estimating nerve excitation thresholds to cutaneous electrical stimulation by finite element modeling combined with a stochastic branching nerve fiber model. Med Biol Eng Comput 49, 385–395 (2011). https://doi.org/10.1007/s11517-010-0725-8

      (3) Höffken, O., Özgül, Ö.S., Enax-Krumova, E.K. et al. Evoked potentials after painful cutaneous electrical stimulation depict pain relief during a conditioned pain modulation. BMC Neurol 17, 167 (2017). https://doi.org/10.1186/s12883-017-0946-7

      (4) Enax-Krumova, E., Plaga, A.-C., Schmidt, K., Özgül, Ö. S., Eitner, L. B., Tegenthoff, M., & Höffken, O. (2020). Painful Cutaneous Electrical Stimulation vs. Heat Pain as Test Stimuli in Conditioned Pain Modulation . Brain Sciences, 10(10), 684. https://doi.org/10.3390/brainsci10100684

      (5) Enrico Schulz, Elisabeth S. May, Martina Postorino, Laura Tiemann, Moritz M. Nickel, Viktor Witkovsky, Paul Schmidt, Joachim Gross, Markus Ploner, Prefrontal Gamma Oscillations Encode Tonic Pain in Humans, Cerebral Cortex, Volume 25, Issue 11, November 2015, Pages 4407–4414, https://doi.org/10.1093/cercor/bhv043

      (6) Mahajan Pranav, Tong Shuangyi, Lee Sang Wan, Seymour Ben (2024) Balancing safety and efficiency in human decision making eLife 13:RP101371 https://doi.org/10.7554/eLife.101371.2

      (7) Enrico Schulz, Elisabeth S. May, Martina Postorino, Laura Tiemann, Moritz M. Nickel, Viktor Witkovsky, Paul Schmidt, Joachim Gross, Markus Ploner, Prefrontal Gamma Oscillations Encode Tonic Pain in Humans, Cerebral Cortex, Volume 25, Issue 11, November 2015, Pages 4407–4414

      (8) Suyi Zhang, Hiroaki Mano, Michael Lee, Wako Yoshida, Mitsuo Kawato, Trevor W Robbins, Ben Seymour (2018) The control of tonic pain by active relief learning eLife 7:e31949

      (9) Hewitt, D., Tong, S., Schreiber, S., & Seymour, B. (2026). Tonic pain modulates neural correlates of associative phasic pain memories. PAIN. DOI: 10.1097/j.pain.0000000000003917

      (10) Gramann, K., Gwin, J. T., Ferris, D. P., Oie, K., Jung, T.-P., Lin, C.-T., Liao, L.-D., and Makeig, S. (2011). Cognition in action: imaging brain/body dynamics in mobile humans. Reviews in the Neurosciences, 22(6):593–582.

      (11) Klug, M. and Gramann, K. (2021). Identifying key factors for improving ica-based decomposition of eeg data in mobile and stationary experiments. European Journal of Neuroscience, 54(12):8406–8420.

      (12) Delorme, A. EEG is better left alone. Sci Rep 13, 2372 (2023). https://doi.org/10.1038/s41598-023-27528-0

      (13) Bach, D. R., Flandin, G., Friston, K. J., and Dolan, R. J. (2010). Modelling event-related skin conductance responses. International Journal of Psychophysiology, 75(3):349–356.

      (14) Rescorla, R. and Wagner, A. (1972). A theory of Pavlovian conditioning: Variations in the effectiveness of reinforcement and nonreinforcement, volume Vol. 2

      (15) Sutton, R. S. and Barto, A. G. (2018). Reinforcement learning: An introduction, 2nd ed. Adaptive computation and machine learning. The MIT Press, Cambridge, MA, US.

    1. Author response:

      (1) Clarification of the scope of the present study and future mechanistic analyses

      We agree that the downstream molecular mechanisms by which SOX17 regulates Sertoli valve formation remain to be elucidated. Our findings are consistent with a model in which SOX17 regulates Sertoli valve formation through paracrine signaling; however, the downstream effectors have not yet been identified. Despite extensive analyses of Sox17 conditional knockout and wild-type mice, including single-cell RNA sequencing, identifying the downstream molecular targets of SOX17 has remained challenging (Uchida et al., 2022). The transgenic mouse model generated in the present study now provides a valuable experimental platform for investigating SOX17-dependent molecular pathways. We are currently performing transcriptomic analyses using this model to identify candidate downstream pathways and genes regulated by SOX17. However, further investigation will be required to determine whether these candidates represent direct transcriptional targets of SOX17 and whether they function specifically within the rete testis during Sertoli valve formation.

      Accordingly, we will avoid overinterpreting the molecular mechanisms in the present study and will revise the Discussion to more clearly acknowledge these limitations while emphasizing that elucidation of these mechanisms represents an important direction for future research. We therefore believe that a comprehensive mechanistic analysis is beyond the scope of the present study.

      (2) Clarification of the quantitative methodology

      We will provide a more detailed description of the methodology used for Sertoli cell quantification. Specifically, Sertoli cells were counted within the SV region extending 100 μm from the rete testis (RT) boundary, and Sertoli cells protruding into the RT lumen were also included in the analysis. The sampling procedure for sagittal RT-SV-seminiferous tubule (ST) sections will be described more explicitly in the revised Methods to improve reproducibility.

      (3) Clarification regarding expression levels

      We appreciate the reviewer's comment regarding the quantitative assessment of SOX17 and other SV-associated molecules.

      The Sertoli valve (SV) is an extremely small transitional structure, with only approximately 20 SVs present in each mouse testis. In addition, Sertoli cells within the SV are tightly interconnected. Consequently, selectively isolating the SV without contamination from adjacent tissues while obtaining sufficient material for quantitative molecular analyses, such as quantitative PCR, remains technically challenging.These technical limitations partly explain why the Sertoli valve has remained an understudied structure in testicular biology. Therefore, in the present study, the expression of SV-associated molecules was primarily evaluated by histological and immunohistochemical analyses. We will clarify these technical limitations in the revised manuscript and revise the relevant text accordingly.

      (4) Additional revisions

      We will address the remaining comments, including clarification of the phenotypic differences between Tg26 (established line) and Tg27 (F0), standardization of gene nomenclature, correction of methodological descriptions, and improvements to the Discussion and figure presentation where appropriate.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      This study by Li and colleagues examines how defensive responses to visual threats during foraging are modulated by both reward level and social hierarchy. Using a naturalistic paradigm, the authors test how the availability of water or sucrose, with sucrose being more rewarding than water, shapes escape behavior in mice exposed to looming stimuli of different intensities, which are used to probe perceived threat level and defensive responses. In parallel, the study compares dominant and subordinate animals to assess how social rank biases the trade off between reward seeking and threat avoidance. By combining detailed behavioral analyses with computational modeling, the work addresses how reward level and social context jointly influence escape decisions in an ethologically relevant setting.

      Across the different experimental conditions, perceived threat level is the main determinant of behavior. The authors show that looming stimuli associated with higher threat (contrast) consistently elicit faster and more robust escape responses than lower threat stimuli. This effect is particularly evident during early exposures, when animals are highly vigilant and have not yet habituated to the looming stimulus (learned that it is not dangerous). Later they described that as animals gain experience and habituate, behavior becomes more flexible, and reward level begins to exert a graded modulation of the escape response. Importantly, the authors show that under high threat conditions increasing reward value leads to more frequent and faster escape rather than greater reward pursuit. This finding is particularly relevant, as it suggests that highly valued rewards can heighten vigilance and thereby enhance responsiveness to threat, highlighting that reward does not simply compete with defensive behavior but can also reshape it depending on the perceived level of danger, in contrast to low threat conditions, where threat can be more easily outweighed by reward. Thus, an important conceptual contribution of the study is the introduction of vigilance as a useful framework to interpret these effects. Vigilance is treated as a behavioral state reflecting heightened attention to potential danger. In line with what is known from natural foraging, mice initially maintain high vigilance when confronted with an innate threat. This perspective helps clarify a finding that might otherwise appear counterintuitive. One might expect higher rewards to motivate animals to tolerate risk, explore more, and habituate faster in any scenario. Instead, the data suggest that highly rewarding outcomes can elevate vigilance, making animals more responsive to threat and leading to faster or more frequent escape under high threat conditions. In this sense, reward does not simply compete with threat but can also amplify sensitivity to it, depending on the internal state of the animal.

      The social results are particularly interesting in this context as well. Dominant mice consistently prioritize avoidance over reward, showing stronger escape responses and slower habituation than subordinates. This behavior is well captured by the vigilance framework proposed by the authors: dominant animals appear to maintain higher vigilance, which biases decisions toward threat avoidance. The authors further suggest that stable social relationships sustain high vigilance and slow habituation, framing this as an evolutionarily conserved strategy that may enhance survival. This interpretation provides a valuable perspective on how social structure shapes defensive behavior beyond immediate physical interactions. At the same time, there are important limitations to this interpretation. All experiments were conducted in male mice, and it is possible that the relationship between social hierarchy, vigilance, and defensive behavior would differ substantially in females. In addition, the idea that stable social relationships maintain elevated vigilance does not straightforwardly align with broader views of social stability as protective for mental health and as a buffer against anxiety and stress. These points do not undermine the findings but suggest that the social effects described here should be interpreted with caution and within the specific context of the task and sex studied.

      We thank the reviewer for raising this important point. In the context of repeated looming exposure, slower habituation reflects more sustained vigilance over time. Compared to individually housed mice, group-housed mice exhibit slower habituation (Lenz et al., 2022), and pair-housed mice showed even slower habituation in our current work. Importantly, this pattern does not indicate that pair-housed mice have higher overall vigilance than individually housed animals. Although individually housed mice habituate more quickly, they display higher initial vigilance, as reflected by their increased probability of escaping in response to looming stimuli (Lenz et al., 2022). Thus, pairhoused mice exhibited reduced defensive responses compared to individually housed animals, consistent with a social buffering effect.

      Furthermore, in a separate study (Rank- and Threat-Dependent Social Modulation of Innate Defensive Behaviors; Li, Gao, Li, 2026, eLife 15:RP109571), we directly compared responses to looming stimuli when mice were tested alone versus in the presence of a social partner and observed clear evidence of social buffering.

      Another important limitation is that the neural mechanisms underlying these effects remain speculative. The manuscript includes an extensive discussion of candidate circuits, particularly involving the superior colliculus and downstream structures, but this section is necessarily based on prior literature rather than on data presented in the study. Given the complexity of the circuits involved in integrating internal state, reward, social context, and vigilance, the current work should be viewed as providing a strong behavioral and conceptual framework rather than direct insight into underlying neural mechanisms.

      We fully agree that the proposed neural mechanisms remain speculative and that the circuits involved in integrating internal state, reward, and social context are likely far more complex. We have revised the manuscript to acknowledge this limitation.

      Methodologically, the behavioral paradigm is well suited for studying escape decisions in socially housed animals, and the machine learning based classification of defensive responses is a clear strength. The computational model provides a useful formalization of how threat level, reward level, and vigilance interact and may be valuable for other laboratories studying escape, approach avoidance, or conflict situations, particularly as a way to classify behavioral outcomes after pose estimation. More generally, the work will be of interest to the neuroethology community for its detailed characterization of escape behavior under naturalistic conditions.

      Given the ethological nature of the study and the high inter individual variability reported by the authors, clarity and precision in the methods are especially important for reproducibility. While the revised manuscript addresses many earlier concerns, some aspects remain slightly difficult to follow. For example, the main text states that animals were not water deprived to avoid differences in internal state, whereas parts of the methods describe conditions in which animals were water deprived, suggesting that internal state manipulation may differ across experiments. Clearer separation and explanation of these conditions would further strengthen confidence in the work.

      To improve clarity, we have revised the Methods section to clearly distinguish between experimental conditions that involved water deprivation and those that did not.

      Overall, this study provides a rich and thoughtful analysis of how reward level and social hierarchy modulate defensive behavior through changes in vigilance. It offers a useful conceptual advance for thinking about escape behavior in naturalistic settings and lays a solid foundation for future work aimed at linking these behavioral states to underlying neural circuits.

      Reviewer #2 (Public review):

      Zhe Li and colleagues investigate how mice exposed to visual threats and rewards balance their decisions in favour of consuming rewards or engaging in defensive actions. By varying threat intensity and reward value, they first confirm previous findings showing that defensive responses increase with threat intensity and that there is habituation to the threat stimulus. They then find that water-deprived mice have a reduced probability of escaping from low contrast visual looming stimuli when water or sucrose are offered in the environment, but that when the stimulus contrast is high, the presence of sucrose or water increases the probability of escape. By analysing behaviour metrics such as the latency to flee from the threat stimulus, they suggest that this increase in threat sensitivity is due to increased vigilance. Analysis of this behaviour as a function of social hierarchy shows that dominant mice have higher threat sensitivity, which is also interpreted as being due to increased vigilance. These results are captured by a drift diffusion model variant that incorporates threat intensity and reward value.

      The main contribution of this work is quantifying how the presence of water or sucrose in water-deprived mice affects escape behaviour. The differential effects of reward between the low and high contrast conditions are intriguing, but I find the interpretation that vigilance plays a major in this process not supported by the data. The idea that reward value exerts some form of graded modulation of the escape response is also not supported by the data. In addition, there is very limited methodological information, which makes assessing the quality of some of the analyses difficult, and there is no quantification on the quality of the model fits.

      (1) The main measure of vigilance in this work is reaction time. While reaction time can indeed be affected by vigilance, reaction times can vary as a function of many variables, and be different for the same level of vigilance. For example, a primate performing the random dot motion task exhibits differences in reaction times that can be explained entirely by the stimulus strength. Reaction time is therefore not a sound measure of vigilance, and if a goal of this work is to investigate this parameter, then it should be measured. There is some attempt at doing this for a subset of the data in Figure 3H, by looking at differences in the action of monitoring the visual field (presumably a rearing motion, though this is not described) between the first and second trials in the presence of sucrose. I find this an extremely contrived measure. What is the rationale for analysing only the difference between the first and second trials? Also, the results are only statistically significant because the first trial in the sucrose condition happens to have zero up action bouts, in contrast to all other conditions. I am afraid that the statistics are not solid here. When analysing the effects of dominance, a vigilance metric is the time spent in the reward zone. Why is this a measure of vigilance? More generally, measuring vigilance of threats in mice requires monitoring the position of the eyes, which previous work has shown is biased to the upper visual field, consistent with the threat ecology of rodents.

      (2) In both low and high contrast conditions, there are differences in escape behaviour between no reward and water or sucrose presence, but no statistically significant differences between water and sucrose (eg: Figure 3B). I therefore find that statements about reward value are not supported by the data, which only show differences between the presence or absence of reward. Furthermore, there is a confound in these experiments, because according to the methods, mice in the no-reward condition were not water-deprived. It is thus possible that the differences in behaviour arise from differences in the underlying state.

      (3) There is very little methodological information on behavioural quantification. For example, what is hiding latency?

      Is this the same are reaction time? Time to reach the safe zone? What exactly is distance fled? I don't understand how this can vary between 20 and 100cm. Presumably, the 20cm flights don't reach the safe place, since the threat is roughly at the same location for each trial? How is the end of a flight determined? How is duration measured in reward zone measures, e.g., from when to when? How is fleeing onset determined?

      (4) There is little methodological information on how the model was fit (for example, it is surprising that in the no reward condition, the r parameter is exactly 0. What this constrained in any way), and none of the fit parameters have uncertainty measures so it is not possible to assess whether there are actually any differences in parameters that are statistically significant.

      These are the public reviews for the original submission. The corresponding authors responses are provided below.

      (1) We agree that reaction time can be influenced by multiple factors, including stimulus strength. Consistent with this, reaction times (i.e. latencies to flee) were substantially shorter under high-contrast conditions (Figure 3E). However, even under the same high-contrast condition, reaction times were significantly shorter in the water condition compared to the no-reward condition, suggesting that other factors such as vigilance may contribute.

      Upward-directed attention includes rearing, up-stretching, and upward head orientation, which will be clarified in the Method section. To address concerns about statistical validity, we will quantify these behaviors across the first 10 trials rather than limiting the analysis to the first two.

      As for the dominance-related results, we interpret them as reflecting both enhanced vigilance and reduced reward-seeking behavior. Time spent in the reward zone is not a measure of vigilance but an indicator of reward-seeking motivation. We will clarify this in the revised manuscript.

      (2) In Figure 3B, the difference between water and sucrose conditions did not reach statistical significance (p = 0.08). We plan to collect additional data to determine whether this is due to limited statistical power. It is also possible that some behavioral readouts are more sensitive to the differences between water and sucrose conditions. For example, Figure 3F shows that escape speed was significantly higher in the sucrose than in the water condition under high-contrast stimulation.

      Thank you for pointing this out. To control for the potential confounds related to internal state, mice were not water-deprived under any of the three conditions in Figures 3A-3H. We will clarify this in the main text and Methods. For Figures 3I-3M, which compare decision-making under no-reward and water conditions, we will conduct additional experiments using non-deprived mice in the water condition.

      (3) Hiding latency was defined as the time from stimulus onset to the animal’s arrival at the safe zone. Reaction time was quantified as the latency to flee, measured from stimulus onset to the initiation of the first flight state. The flight state was defined as locomotion exceeding 10 cm at a speed greater than 10 cm/s. Distance fled was defined as the distance covered between stimulus onset and offset for all trials. However, in trials classified as no reaction or freezing, this measure does not accurately reflect escape behavior. We will therefore rename it as distance under threat to better capture its meaning. The reward zone was defined as the region within 15 cm of the reward port at the end of the arena. Duration in the reward zone was measured as the time spent within this region during the 20 seconds following stimulus onset. In Figure 4E, the percentage of time spent in the reward zone was calculated relative to the total time the mouse remained in the arena during the 2-hour social session.

      All definitions and additional details on behavioral quantification will be included in the revised Methods section.

      (4) We appreciate the comment and agree that further clarification is needed. We will provide a more detailed description of the model fitting procedure in the revised Methods section. Specifically, the drift rate parameter (r), which reflects the perceived reward value, was constrained to zero in the no-reward condition. To enable statistical comparison across conditions, we will report uncertainty measures for all fit parameters.

      Comments on the revised manuscript:

      The manuscript has been revised and improved significantly by the addition of methodological details and new analysis. I remain, however, unconvinced by the argument that increased vigilance in the presence of reward leads to heightened escape behaviour.

      In response to my criticism that the work does not measure vigilance directly, the authors have included measures of foraging interval and foraging speed, which they state are "two direct behavioral analyses of vigilance". I disagree - like reaction time, foraging speed and foraging interval can be modulated, for example, by changes in threat sensitivity. Increased threat sensitivity comes with diverse behavioral changes that may well include increased vigilance, but foraging interval and foraging speed can certainly change without the animal expressing increased vigilance behaviors. A bigger issue I still have though, is with the conclusion that the presence of reward increases "direct escape behaviors". Comparing the no reward, water and sucrose groups indeed shows a difference (which is now clear after the split into early and late phases), but the issue is that these are different mice. As the text is written, is sounds like introducing reward will acutely increase escape. But if we look at the raw data show in Figure 2C, what I think is happening is that the presence of reward is decreasing habituation to the stimulus. The data for trials 1 and 10 in the three conditions show this - there is habituation with no reward (reaction times are all shifting to the right), a bit less with water and very little with sucrose. This is interesting in its own right and we can speculate why it might be happening, but I think this is conceptually different from what the authors are proposing.

      We agree that vigilance is not directly observable as a single variable. Our intent was not to claim that foraging speed and foraging interval provide a direct measure of vigilance, but rather to suggest that they may serve as indirect behavioral correlates.

      We also considered an alternative interpretation: these two measures could reflect perceived reward value under high-threat conditions across distinct reward types. If that were the case, animals would be expected to exhibit shorter intervals and faster speeds across no reward, water, and sucrose conditions. However, our data do not support this interpretation (Figures 3L and 3M), suggesting that these measures are more likely correlated with vigilance.

      Furthermore, it is unlikely that changes in foraging interval and speed are driven by altered threat sensitivity, as animals could not see the threat during most of the foraging bout and only encountered it at the end.

      Regarding the conclusion that the presence of reward increases direct escape behaviors, our interpretation is that increased reward value reduces habituation, thereby maintaining higher vigilance during the late phase. This was discussed in the second-to-last paragraph of the "Economic and social modulations of innate decision-making under threat" subsection in the Discussion.

      Reviewer #3 (Public review):

      Male mice were tested in a classic behavioral "flee the looming stimulus" paradigm. This is a purely behavioral study; no neural analyses were done. Mice were housed socially, but faced the looming stimulus individually, using an elegant automated tunnel (see videos for clarity).

      The additional changes made to the paper clarify the work done. While there are some limitations (male mice, weird stimulus), the general results are interesting and a valuable addition to the experimental literature. The main claim of the paper is that the different rewards (none, water, sucrose) did not change the escape properties early in learning, but did late, particularly that in the late (already experienced) conditions, reward value (assuming sucrose > water > no reward) interacted with the salience of the looming stimulus (light gray, dark gray). (Panels 3D, 3G, 3K, 3N).

      For readers, I want to note that one of the most interesting results is actually in Figure S2, where they find that a looming stimulus behind the mouse still makes a mouse run to the nest. In these conditions, the mouse runs past the looming stimulus to get to safety! (I also do love the video of the mouse running around the barriers like a snake to get home.)

      I have a few minor clarification questions and a few notes that I think would be useful additions for authors and readers to think about.

      Dominance: What does the mouse social science literature say about the "test tube" test? What can we conclude from this test? This would be useful when trying to understand what is causing the dominance/submissive difference in responses. Figure 4 shows that the dominant mice are more risk-averse than the submissive mice. Is "dominance" in the test-tube actually a measure of risk-seeking? Is the issue that the submissive mice don't think they can get back to the food-site easily, so they are less willing to sacrifice the current (if dangerous) foraging opportunity? Is the issue that the submissive mice can't get back to the nest? As I understand it, the nest was always available to all the mice, so I suspect inability to get to the nest is an unlikely hypotheses. Is the issue that the submissive mice also don't feel safe in the nest?

      The tube test is a widely used assay in the rodent social behavior literature to assess dominance hierarchies, operationally defined by the ability of one animal to force its opponent to retreat from a narrow tube. Importantly, this assay does not directly measure risk-seeking or anxiety-related traits, but rather competitive outcomes during social conflict. Furthermore, our data indicate that the behavioral responses of subordinate mice to looming stimuli are primarily driven by the visual threat itself rather than by social avoidance. This point was elaborated in the second paragraph of the “Social modulation of innate decision-making” subsection in the Results section.

      Limitations of the study: There is an acknowledged limitation to male mice, and the limitations of the small data sets that are typical of such experiments. In addition, however, it is also worth noting the strangeness of the looming stimulus, which is revealed clearly in the videos. The stimulus is a repeating growing circle, growing in a single location within the environment. The stimulus repeats 10 times, once per second. This is not what an attacking hawk or owl would look like. (I now have this image of an owl diving down, and then teleporting up and diving down again.) Note - I am fine with this stimulus. It produces an interesting experiment and interesting results. I do not think the authors need to change anything in their paper, but readers need to recognize that this is not a "looming predator".

      These "limitations" are better seen as "caveats" when folding these results in with the rest of the literature that has gone before and the literature to come. (Generally, I do not believe that science works by studies making discoveries that change how we think about problems - instead, science works by studies adding to the literature that we integrate in with the rest of the literature.) Thus, these caveats should not be taken as problems with the study or as fixes that need to be done. Instead, they are notes for future researchers to notice if differences are found in any future studies.

      Thus, my only suggestion is that I think authors could write a more careful paper by using the past and subjunctive tense appropriately. Experimental observations should be in past tense, as in "the influence of reward was contextdependent and emerged in the late phase" instead of "the influence of reward is context-dependent and emerges in the late phase" - it emerged in the late phase this once - it might not in future experiments, not due to any fault in this experiment nor due to replicability problems, but rather due to unexpected differences between this and those future experiments. At which point, it will be up to those future experiments to determine the difference. Similarly, large conclusions should be in the subjunctive tense, as in "these data suggest that threat intensity is likely to be the primary determinant of decision making" rather than "threat intensity is the primary determinant of decision making", because those are hypotheses not facts.

      We thank the reviewer for the helpful suggestions and have revised the Abstract accordingly.

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Figure 5: The points in panel 5G and 5H are unreadable. What are these stars and symbols supposed to mean? They are also too small to see without zooming way in.

      We have increased the symbol size.

      Figure 5: What is the final panel of 5J? I did not understand this panel at all. The first three panels of 5J (threat-based detection, reward-based detection, vigilance-based detection) are, I believe, three patterns we should look for in the data. But then what is the "experimental results" section? It contains all three, but they don't overlap? Shouldn't we have an experimental results section for each condition?

      Panel 5J was to compare three hypothesized decision patterns with the experimentally observed data. To make this distinction explicit, we have revised the panel titles to: “H1: Threat-based decisions,” “H2: Reward-based decisions,” “H3: Vigilance-based decisions,” and “Experimental results.”

      Thank you for including the videos. They made the task construction and the stimulus much clearer.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Weaknesses:

      While the evidence in favour of the two gradients largely supports the claims, the evidence for a new visual field map cluster in the anterior temporal lobe falls short of the level used historically when identifying visual field maps in the visual cortex and is, at present, not convincing. More specifically, the progressions of polar angle within the putative anterior lobe cluster are highly variable across subjects. Few subjects have convincing polar angle reversals at either the horizontal or vertical meridians. In other cases, a putative border is shown that spans different polar angles, which does not align with the accepted definitions for visual field maps in the cortex.

      We agree with the reviewer that more evidence could be provided in support of retinotopic representations within the anterior temporal lobe. We have performed a number of new analyses to further explicate the receptive field properties of this anterior temporal lobe visual representation. We have pasted updated Figure 2e-i. We have added additional participants, increasing the total number from N=12 to N=21. In panel g, we show that in this larger group, we can still observe pRFs that are about 3x larger than those in early visual cortex, and that the relationship between their size and eccentricity shows the expected steeper slope compared to these early representations. In this new participant group, we also illustrate the visual field coverage of the left and right anterior temporal lobe representations (panel h). As expected, the left hemisphere pRFs largely sample the right visual field, and right hemisphere pRFs largely sample left visual space. One can also see that both the upper and lower visual fields are sample quite evenly, consistent with the hemi-field representation of visual field maps observed in earlier visual cortex. To quantify whether there is a left-right contralateral bias in the sampling of visual space (and to test whether such a bias is significantly different in each hemisphere), we calculated for each pRF a laterality index as previously defined by Sheremata and Silver (2015) according to the equation below:

      Where resulting values of 1 mean the pRF is contralateral, 0.5 is no laterality bias, and 0 is ipsilateral bias. Additionally, we input pRF sigma values that were adjusted for the non-linearity exponent as defined by Kay et al. (2013). For the purposes of visual comparison, we subtracted 0.5 from index values so that resulting laterality scores were relative to 0 to represent the center of the visual field, and then values were inverted with a -1 scalar so that left hemisphere pRF laterality index values are plotted on the right side of space, and the right hemisphere on the left as shown in panel i. The laterality index was calculated for each pRF for a given participant and then averaged within that participant to result in a single mean laterality index for the left hemisphere pRFs and a single index for their right hemisphere pRFs. The histograms illustrated in panel i depict density of participants (kernel smoothed). We find a significant difference between laterality indices with left AT pRFs showing significantly rightward index values compared to right AT pRFs (paired-samples t-test, t(20) = 7.6, p = 2.7 x10<sup>-7</sup>). These data thus offer stronger evidence of a hemifield representation with a contralateral bias, and it should also be noted that there is stronger ipsilateral coverage in these high-level visual pRFs compared to earlier visual field maps like V1, which is consistent visual field maps in latera stages of the visual processing hierarchy as quantified by Mackey et al. (2017).

      Lastly, we note that the progression of polar angle values on the cortical surface is certainly not as strikingly topographic as in visual field maps V1 through hV4. This is perhaps a result of the strong ipsilateral visual field coverage in which pRFs whose centers were near or within the ipsilateral field (especially those near the fovea) are not visualized appropriately when using a contralateral colormap. It is also possible that at this very late stage of visually-responsive cortex within entorhinal cortex that retinotopic topography becomes less clear as is the case in higher stages of the dorsal visual stream. To improve visualization, we have created a new Supplemental Figure 6 using a binary color map that colors lower and upper visual field in separate colors and extends into the ipsilateral visual field (pasted below for convenience). We hope that this color map helps to show the upper and lower visual field coverage. While there is a clear radial eccentricity gradient within these AT pRF clusters, and while most participants do show a polar angle gradient that runs perpendicular to this radial eccentricity gradient as expected for a visual field map, we do agree that it is difficult to observe polar angle traversals as clearly as in earlier visual cortex. Nonetheless, the presence of these pRF clusters which show their own distinct eccentricity representation (i.e., a foveal confluence) and a full sampling of the contralateral visual space is still consistent with our anatomical model’s prediction in which PC2 anchor points predict foveal representations shared by visual field map clusters. While the topographic clarity of these representations on the cortical surface is less than earlier visual cortex, the existence of contralateral representations of visual space with a full eccentricity gradient that spans the upper and lower visual field is strongly supported by the data and consistent with our anatomical model’s prediction that there should have been a distinct eccentricity gradient. These findings are also consistent with work showing that the human hippocampus also shows sensitivity to contralateral visual space (Silson et al., 2021) and suggests the hippocampus may inherit this contralateral bias from this entorhinal visual representation. We have updated the manuscript to incorporate these new findings, and refer to these AT clusters as contralateral visual representations, remaining agnostic to whether or not they can be fully defined as topographic maps which can be the focus of future work using smaller voxel sizes to better capture small topographic gradients.

      We have revised the manuscript to incorporate these points in the following sections.

      Line 466: “We performed pRF mapping on 21 participants with high-contrast, …”

      Line 601-625: “To produce maps of visual field coverage (Figure 2h) similar to previous work, … The histograms illustrated in Figure 2i depict density of participants (kernel smoothed).”

      Line 236-246: “We find that consistent with its high position within the processing hierarchy, … We find a significant difference in laterality indices between left and right AT pRF’s (pairedsamples t-test, t(20) = 7.6, p = 2.7 × 10-7).”

      Line 373-383: “The organization of polar angle in anterior temporal cortex was not as orderly as earlier visual cortex, … in more posterior portions of ventral occipitotemporal cortex.”

      Reviewer #2 (Public review):

      Weaknesses:

      (1) The neurobiological model does not take into consideration present knowledge about the microstructural organization of the visual system. This limits the way the results are interpreted correctly. Critical information on the layer-specific myeloarchitecture and cytoarchitecture (and their relation to cortical thickness), as explored for example by Sereno et al. 2013 Cereb Cortex, is missing. There is no information given with respect to how different visual areas differ in their microstructural profile. It is also not mentioned that cortical parcellation is indeed characterized by sharp boundaries between areas, rather than structural gradients, so it remains unclear why focusing on a gradient is of interest. The authors cite the parcellation atlas by Glasser et al. 2016, but do not discuss the rationale of this publication, which was not the definition of gradients, but the definition of sharp boundaries for cortex parcellation. Indeed (as explained below), the results of the authors seem to a large extent to be driven by cortex parcellation, but instead of acknowledging this fact, the authors write (line 179) that "we hypothesize that these local deviations from the canonical thickness and density of cortex underlie the finer-scale division of visual cortex into categorically distinct regions. That is, does the realization of the cortex into distinct regions involve these regions becoming more distinct from a prototypical cortical sheet (i.e., gradient 1)?" - While the first sentence is reasonable, the second sentence is pure speculation ignoring present knowledge on cortical parcellation of this area according to which there is no "prototypical cortical sheet", but each area has its distinct microstructural profile.

      We thank the reviewer for this important comment. We first want to point out that we believe there is a conceptual misunderstanding on the part of the reviewer, as we address in our lengthy response below. In this response, we explain that our findings capture what we believe is a novel finding—that variation across participants in the cortical sheet is not random across the spatial expanse of cortex but respects its functional boundaries—which we view as a finding that is complimentary to the current knowledge about the microstructure of visual cortex. It was not our intention to ignore or gloss over this present knowledge, but instead show that variation in these cortical microstructures across brains is not random.

      We agree that incorporating current knowledge about the microstructural organization of visual cortex, including its laminar architecture and sharp areal boundaries, is critical for situating our findings within the broader literature. In response, we have added key background information on the relationships among cytoarchitecture, myeloarchitecture, and cortical thickness, as described in previous studies (for example, Maingault et al., 2021; Sereno et al., 2013; Shafee et al., 2015). While our study does not aim to capture layer-specific properties per se, which would require different imaging modalities and higher-resolution data, we focus on spatial properties tangential to the cortical surface.

      We first address a concern that the particular parcellation might be driving effects with an analysis showing that we believe our finding is robust to this concern. As suggested by the overall negative covariance observed between cortical thickness and tissue density, we further confirmed this relationship not only across larger visual ROIs, which could potentially reflect effects of arealization, but also within individual ROIs at a finer spatial scale. To avoid potential circularity in ROI definition, we used a visual ROI atlas derived from population-level retinotopy based on independent datasets (Abdollahi et al., 2014). We found that at the global level, cortical thickness and T1w/T2w ratio showed a strong negative correlation across visual ROIs (Fig. 3, revised Supp. Fig. 3a & b). Although only a portion of the visual cortex is clearly delineated in this atlas, we replicated similar results across the entire visual cortex using the MMP atlas (Glasser et al., 2016). At the within-ROI level, we found robust negative correlations between cortical thickness and T1w/T2w ratio across most visual ROIs in both hemispheres, with the notable exception of V1, V2 and VO1, which exhibited a positive relationship, consistent with prior work (for example, Maingault et al., 2021; Sereno et al., 2013; Shafee et al., 2015). These results highlight both common and distinct microstructural profiles across the visual cortex and provide important context for interpreting our data-driven findings.

      We also want to address what we think is a conceptual misunderstanding by the reviewer, which likely resulted from a lack of clarity on our part. The reviewer’s confusion likely results from the fact that we theoretically “transposed” the typical PCA analysis such that we get a subject-wise contribution (PC loadings) per participant (also see response to next point), which is how we’re able to relate inter-participant variability in their loadings to behavior in Figure 3. This is also why we refer to a “typical” cortex/cortical sheet because the surface maps being visualized for PC2 can be thought of as a map explaining variance of deviation orthogonal to PC1 (which captures the primary relationship between thickness and T1/T2). Thus, because PC2 is orthogonal to PC1, it captures the spatial pattern in which participants deviate from the primary relationship (e.g., the typical relationship). Therefore, if a given participant is far from the PC1 vector and has high PC2 loading, their cortical sheet is either thicker or more myelinated than predicted by the PC1 relationship and is therefore more distinct from the “typical” or “average” cortical sheet values captured by PC1. We want to emphasize that PCA is agnostic to spatial structure across the cortex. Thus, the fact that deviation from the primary thickness-myelination relationship (i.e. PC2) captured by PC1 had any spatial structure at all is interesting. Furthermore, the fact that the spatial structure of PC2 across the cortical sheet seems to separate visual cortex into its constituent processing streams is also interesting. Therefore, we are not speculating but rather describing the PCA model itself whereby a participant’s loading on PC2 describes their deviation or distinctness from the PC1 relationship. The fact that PC2 has spatial structure on the cortical sheet (which did not have to be true) and the fact that this structure seems to capture broad borders between visual processing streams and field maps is what we find interesting and quantify within the paper. We hope this additional explanation clarifies the broader theoretical thrust of the paper. We view these findings as complimentary to the present knowledge of the microstructural organization of the visual system. Our findings suggest that variability in these microstructural features across participants (PC2) don’t occur randomly across cortex but seem to respect the functional borders of the neural populations of the underlying cortical sheet.

      Regarding the concern that our gradient approach may contradict established knowledge of cortical arealization, we would like to clarify that the primary goal of our gradient analysis is not to redefine visual areas, or to go against cortical arealization, but to explore the continuous variation in cortical architecture across brains that may co-exist alongside sharp boundaries which is phenomenon complementary to the arealization. In our study, cortical thickness maps were regressed for curvature before entering any analyses, given the covariance between cortical folding and area borders (Fischl et al., 2008). We acknowledge that cortical parcellation is traditionally characterized by discrete transitions between areas. However, our results suggest that gradients of cortical properties—particularly those shared across participants—may capture supra-areal organizing principles that reflect how distinct regions relate to one another within a broader cortical sheet.

      Finally, we agree with the reviewer that the phrase “prototypical cortical sheet” was speculative and potentially misleading. We have removed this language from the manuscript and revised the corresponding discussion.

      We have revised the manuscript to incorporate these points in the following sections.

      Line 92-94: “Thickness and density maps showed a robust anti-correlation both at the coarse across-area level based on an independent parcellation and at the finer within-area level, except in primary regions (Figure S3a, b).”

      Line 350-353: “The convergence pattern, arising from the negative correlation between thickness and density, is consistent with previous findings and may support the balloon model, whereby cortical thinning is associated with tangential stretching due to myelination.”

      Line 188-189: “That is, does the arealization of cortex into distinct regions involve these regions becoming more distinct from a typical cortical sheet (i.e., gradient 1)?”

      (2) Instead of building on present, detailed knowledge of brain anatomy and in-vivo cortex parcellation of the visual system and its known relation to visual maps, the authors focus on two metrics of cortex architecture (mean T1/T1 over depth and cortical thickness), and conduct a PCA to explore their shared variance. It needs to be clarified if the PCA was conducted correctly. There is no mention of standardizing the variables, which could bias the results. In addition, in a PCA, all possible features are categorized as vector components, and those are scanned through the samples, hence, one such analysis per vertex. But the authors write "in which participants are features and cortical vertices are samples" and "the thickness and tissue density maps were concatenated". This needs clarification. The architecture of the PCA should be visualized better.

      We thank the reviewer for pointing out the need to clarify the PCA methodology. In response, we have revised the Methods section to provide a clearer and more accurate description of our approach.

      We also would like to point the reviewer’s attention to Figure 1a, in which the PCA was illustrated graphically. The reviewer’s confusion likely results from the fact that we theoretically “transposed” the typical PCA analysis such that we get a subject-wise contributions (PC loadings) per participant, which is how we’re able to relate inter-participant variability in their loadings to behavior in Figure 3. This is also why we refer to a “typical” cortex/cortical sheet because the surface maps being visualized for PC2 can be thought of as a map explaining variance of deviation orthogonal to PC1 (which captures the primary relationship between thickness and T1/T2). Thus, because PC2 is orthogonal to PC1, it captures the spatial pattern in which participants deviate from the primary relationship (e.g., the typical relationship).

      We have revised the manuscript in the following sections.

      Line 493-502: “For each hemisphere, individual cortical thickness and T1/T2-weighted ratio maps from all HCP-YA participants—each represented as an M × N matrix, … corresponding participant-wise contributions (i.e., PC loading or individual weights) in pairs.”

      (3) Because the PCA only contains two features, PC1 is driven by the positive relationship between cortical thickness and mean T1/T2, whereas PC2 is driven by their negative relationship. Because in the early visual cortex, cortical thickness and mean T1/T2 correlate positively, it naturally follows that PC1 relates to pRF size (but mediated by the actual cortex parcellation). However, it is unclear why this insight is interesting. I also do not share the view that "these findings demonstrate that gradient 1 acts as a global gradient enveloping the entire visual cortex (...) while gradient 2 acts as a local gradient specific to individual visual streams". I think this relationship between cortical thickness and T1/T2 ratio does not have much to do with local and global gradients. But if so, stronger arguments as to why this should be the case should be presented. What the authors make of this result (particularly the discussion starting line 366) is not clear to me. I cannot follow the line of argumentation, which in my view is too far away from the data.

      We appreciate the reviewer’s thoughtful comments and agree that, in general, cortical thickness and T1w/T2w ratio tend to be negatively correlated, with early visual areas (i.e., V1 and V2) representing a notable exception—an observation we highlight and support with evidence in R2. Given this overall pattern of correlation, it may seem intuitive to interpret PC1 as capturing a convergent relationship across the two metrics, and PC2 as reflecting their divergence. Alternatively, one can think of PC2 as the orthogonal residuals from the linear relationship between thickness and myelin captured by PC1. In this framework, PC2 is not necessarily the inverse correlation, but instead what is left unexplained through a simple linear model. However, it is important to note that PCA is inherently agnostic to spatial structure, as our PCA operates solely on inter-subject variance. As such, the spatial patterns observed in the resulting component maps are not direct or trivial consequences of the input correlations.

      Upon examining the spatial properties of the PCA-derived maps (Fig. 1d), we found that PC1 manifests as a large-scale, low-frequency gradient spanning broad portions of the visual cortex, whereas PC2 exhibits a fine-scale, high-frequency pattern confined to subregions of the visual cortex (quantified in Fig. 1f, g). Our initial use of the terms “global” and “local” may have inadvertently implied functional interpretations beyond our intent. We have revised the manuscript to clarify that these descriptors were intended purely to convey differences in spatial scale based on the observed frequency content of the gradients.

      Motivated by the reviewer’s comment, we performed additional analyses to explicitly test whether the PCA components reflect consistent (i.e., global) or variable (i.e., local) relationships across visual ROIs. Specifically, we examined whether the direction and magnitude of PC1 and PC2 scores within each ROI align with the global relationships between cortical thickness and tissue density. As shown in the revised Supp. Fig. 3e, we found that in most ROIs, vertices with high PC1 scores consistently exhibit high cortical thickness and low T1w/T2w ratios, while those with low PC1 scores show the opposite pattern. This within-ROI consistency mirrors the largescale cross-ROI correlation structure (see Supp. Fig. 3a), supporting the interpretation of PC1 as reflecting a large-scale, cortex-wide organizational principle. In contrast, PC2 shows more heterogeneous profiles across ROIs, with peaks and troughs that differ in the two metrics. This variability suggests that PC2 captures more localized, region-specific features.

      We have incorporated the results of these new analyses into the Results section to strengthen our argument regarding the spatial scale and cross-regional consistency of the PCA-derived gradients:

      Line 102-107: “Within-area analyses further confirmed that PC1/2 represent the consistent/deviating components … while PC2 represents the spatial divergence from this commonality.”

      Recommendations for the authors:

      Reviewing Editor Comments:

      Through collaborative discussions among the reviewers, we first summarised the key recommendations for enhancing the significance and strengthening the evidence of the work - integrating public reviews and recommendations to authors by each reviewer individually. The individual reviewer recommendations can be found below this.

      (1) Modelling component 2

      The geodesic model for component 2 is interesting but we can recommend ways to improve the evidence and interpretation (see Reviewer 1 comments). As the polar angle reversals are inconsistent and boundaries ambiguous, the OTS maps do not meet the standard of evidence required for showing a new map. The 181 pRF maps available for these HCP data would provide an independent more powerful test of the OTS map cluster. To further strengthen the evidence for the proposed correspondence of foveal confluences and gradient 2, why not define the geodesic model anchoring points based on retinotopic measures, e.g., using HCP pRF data? About the current anchoring points for the geodesic model, what were the criteria - were they objective to avoid circularity?

      We appreciate the reviewer’s suggestion to incorporate the HCP 7T retinotopy dataset as an independent test of the proposed geodesic model and its relation to foveal confluences and gradient 2. We agree in principle that such data could provide a valuable validation resource. However, as detailed in the publication accompanying the HCP 7T retinotopy dataset (Benson et al., 2018), the authors recommend a threshold of 9.8% variance explained to distinguish reliable pRF estimates from noise. As illustrated in their Figure 4, this thresholded pRF data shows poor signal coverage in higher-order visual regions, particularly those along the occipitotemporal sulcus (OTS), where gradient 2 effects are most prominent in our data. This lack of reliable pRF signal in these regions limits the utility of the HCP retinotopy data for anchoring the geodesic model or validating the observed spatial gradients.

      To address this limitation, we relied on our in-house data collected using high-contrast, naturalistic images designed to robustly activate high-level visual areas. This approach allowed us to define more complete and consistent topographic patterns in the regions of interest. We have thus expanded the size of this in-house dataset to N=21. We also point the editor’s attention to the response to Reviewer 1’s first comment regarding the visual field maps for a more detailed response to this point. For convenience, we have pasted the Figure 2 e-i panels in which we conduct additional analyses showing that these anterior temporal pRF clusters tile contralateral visual space as one might expect (Fig 2h), and significantly differ across hemispheres in their laterality bias (Fig 2i). We have revised the manuscript accordingly.

      To mitigate the concern of circularity in defining the geodesic model’s anchor points, we conducted a split-half cross-validation. Anchors were defined on one half of the participants and used to predict the PC2 map in the other half. The PC2 maps across the two halves were highly similar (r = 1.00, p < 0.001), indicating strong reliability. Importantly, the cross-predicted geodesic model accounted for a significant portion of variance (r<sup>²</sup> = 0.23) in the held-out PC2 map, suggesting that the geodesic organization is not an artifact of overfitting or circular reasoning. We have revised the manuscript accordingly:

      Line 139-142: “A split-half cross-validation yielded similar results, … underlying the spatial organization of PC2.”

      (2) Speculation about prototypical cortical sheet

      You hypothesise that gradient 1 characterises a global "prototypical cortical sheet" characteristic, with gradient 2 reflecting that regions become more distinct from this prototype. There is an alternative simpler possibility: the data can be explained by the stronger relationship between cortical thickness and T1/T2 ratio in early compared to late sensory areas, as can for example be seen in Glasser et al. 2016 Nature, Figure 4. We recommend omitting or balancing the statement about a "prototypical" cortex, and integrating findings on cortex parcellation and the view that sharp boundaries characterize transitions between high and low T1/T2 and cortical thickness areas.

      Please see R2 for reviewer #2

      (3) Confounds

      We'd like to see more data to understand the contributions of data quality to these results. For the component 1 gradient specifically, could its features be influenced by spatial SNR inhomogeneities? Could the developmental effects for both gradients be explained by lower SNR and other data quality markers in younger and older participant data? We missed appropriate tests that gradients develop differently across age, controlling for such confounds (Reviewer 1 comments).

      Regarding the reviewer’s concern about the component 1 gradient, we believe it is unlikely to be merely a consequence of uneven spatial SNR. Our findings are consistent with previous histological studies demonstrating systematic variations in cortical architecture—specifically, thinner cortex (Wagstyl et al., 2020) and higher myelin content (Dinse et al., 2015) in occipital compared to ventral visual regions. This correspondence between in vivo MRI-derived measures and postmortem histology suggests that the large-scale organization captured by PC1 is grounded in biologically meaningful cortical architecture, and not an artifact of SNR variability.

      To statistically assess whether the two PCs show different developmental trajectories across age, we performed an ANOVA with age, LC, and their interaction as factors on LC’s similarity to PC (i.e., r ~ age + LC + age × LC). Significant age × LC interactions were observed in the developmental (HCPD: F<sub>1,118</sub> = 257.01, p < .001) and aging (HCPA: F<sub>1,132</sub> = 263.85, p < .001) cohorts, but not in the young adult cohort (HCPYA: F<sub>1,202</sub> = 0.02, p = 0.80). These findings indicate that the two gradients show distinct age-related changes during development and aging but remain stable in young adulthood. We have revised the manuscript accordingly:

      Line 313-327: “Examining the correlation between the young adult gradient and LC … F<sub>1,132</sub> = 263.85, p < 0.001).”

      (4) Implementation of PCA

      The manuscript raises questions about the correct implementation of the PCA - please clarify that the variables were first standardised to enable fair weightings, and visualise the PCA matrix in more detail than in Figure 1a to ensure the samples and features are correctly defined (Reviewer 2).

      Please see R3 for reviewer #2

      References

      Abdollahi, R. O., Kolster, H., Glasser, M. F., Robinson, E. C., Coalson, T. S., Dierker, D., Jenkinson, M., Van Essen, D. C., & Orban, G. A. (2014). Correspondences between retinotopic areas and myelin maps in human visual cortex. NeuroImage, 99, 509–524. https://doi.org/10.1016/j.neuroimage.2014.06.042

      Benson, N. C., Jamison, K. W., Arcaro, M. J., Vu, A., Glasser, M. F., Coalson, T. S., Van Essen, D. C., Yacoub, E., Ugurbil, K., Winawer, J., & Kay, K. (2018). The HCP 7T Retinotopy Dataset: Description and pRF Analysis. https://doi.org/10.1101/308247

      Dinse, J., Härtwich, N., Waehnert, M. D., Tardif, C. L., Schäfer, A., Geyer, S., Preim, B., Turner, R., & Bazin, P.-L. (2015). A cytoarchitecture-driven myelin model reveals area-specific signatures in human primary and secondary areas using ultra-high resolution in-vivo brain MRI. NeuroImage, 114, 71–87. https://doi.org/10.1016/j.neuroimage.2015.04.023

      Fischl, B., Rajendran, N., Busa, E., Augustinack, J., Hinds, O., Yeo, B. T. T., Mohlberg, H., Amunts, K., & Zilles, K. (2008). Cortical Folding Patterns and Predicting Cytoarchitecture. Cerebral Cortex, 18(8), 1973–1980. https://doi.org/10.1093/cercor/bhm225

      Glasser, M. F., Coalson, T. S., Robinson, E. C., Hacker, C. D., Harwell, J., Yacoub, E., Ugurbil, K., Andersson, J., Beckmann, C. F., Jenkinson, M., Smith, S. M., & Van Essen, D. C. (2016). A multimodal parcellation of human cerebral cortex. Nature, 536(7615), 171–178. https://doi.org/10.1038/nature18933

      Kay, K. N., Winawer, J., Mezer, A., & Wandell, B. A. (2013). Compressive spatial summation in human visual cortex. Journal of Neurophysiology, 110(2), 481–494. https://doi.org/10.1152/jn.00105.2013

      Mackey, W. E., Winawer, J., & Curtis, C. E. (2017). Visual field map clusters in human frontoparietal cortex. eLife, 6, e22974. https://doi.org/10.7554/eLife.22974

      Maingault, S., Pepe, A., Mazoyer, B., Tzourio-Mazoyer, N., & Crivello, F. (2021). Characterization of late structural maturation with a neuroanatomical marker that considers both cortical thickness and intracortical myelination. https://doi.org/10.1101/2021.02.24.432645

      Sereno, M. I., Lutti, A., Weiskopf, N., & Dick, F. (2013). Mapping the Human Cortical Surface by Combining Quantitative T1 with Retinotopy†. Cerebral Cortex, 23(9), 2261–2268. https://doi.org/10.1093/cercor/bhs213

      Shafee, R., Buckner, R. L., & Fischl, B. (2015). Gray matter myelination of 1555 human brains using partial volume corrected MRI images. NeuroImage, 105, 473–485. https://doi.org/10.1016/j.neuroimage.2014.10.054

      Sheremata, S. L., & Silver, M. A. (2015). Hemisphere-Dependent Attentional Modulation of Human Parietal Visual Field Representations. The Journal of Neuroscience, 35(2), 508–517. https://doi.org/10.1523/JNEUROSCI.2378-14.2015

      Silson, E. H., Zeidman, P., Knapen, T., & Baker, C. I. (2021). Representation of Contralateral Visual Space in the Human Hippocampus. The Journal of Neuroscience, 41(11), 2382–2392. https://doi.org/10.1523/JNEUROSCI.1990-20.2020

      Wagstyl, K., Larocque, S., Cucurull, G., Lepage, C., Cohen, J. P., Bludau, S., Palomero-Gallagher, N., Lewis, L. B., Funck, T., Spitzer, H., Dickscheid, T., Fletcher, P. C., Romero, A., Zilles, K., Amunts, K., Bengio, Y., & Evans, A. C. (2020). BigBrain 3D atlas of cortical layers: Cortical and laminar thickness gradients diverge in sensory and motor cortices. PLOS Biology, 18(4), e3000678. https://doi.org/10.1371/journal.pbio.3000678

    1. Author response:

      eLife Assessment

      In this valuable manuscript, the authors tackle a highly relevant question in biology: how cells integrate attractive and repulsive cues to achieve directed migration. They present solid data demonstrating that two wunen genes act as negative regulators of Hedgehog signalling, thereby enabling efficient primordial germ cell (PGC) migration in Drosophila embryos. Beyond its immediate scope, this work has broader implications, particularly for understanding key mechanisms underlying complex processes such as cancer metastasis, where the coordinated interpretation of guidance cues is critical.

      Thank you for the reviews and the overall assessment of our manuscript. It is our impression that both the reviewers and the senior editor find the study interesting and potentially of general relevance. The reviewers have made specific suggestions to improve the manuscript. They have also recommended ways to uncover the mechanistic basis to add to the broad appeal of the findings.

      To begin with, we would like to point out that since the discovery of Wunen in 1996 by Ken Howard and colleagues, a number of genetic and molecular studies have attempted to identify and characterize the putative target(s) of the two lipid phosphate phosphatase(s). We and others have shown that Hh acts as a guidance signal for the migrating PGCs. Our data demonstrating the ability of Wunen(s) to attenuate Hh signaling constitutes an important step in elucidating the molecular underpinnings of the repulsive activity of Wun(s) during PGC migration.

      Thus, we feel the need to share these findings with the scientific community at this juncture. In the following, we will summarize our response to the relevant points included in the individual public critiques of the reviewers without going into specific details.

      Public Reviews:

      Reviewer #1 (Public review):

      This manuscript addresses how PGCs migrate towards SGPs in the Drosophila embryo. It's been shown that Hh produced by SGPs acts as an attractive cue, and that Wunnen(s) act as repulsive cues. In this work, the authors propose that Wun and Wun2 refine PGC guidance by attenuating Hedgehog signalling coming from other tissues.

      Overall, the study is potentially interesting and could make an important contribution to the field. The data shown support the idea that Wun/Wun2 negatively regulate Hh signalling and produce PGC migration phenotypes associated with Hh. However, in my opinion, there are two major questions that should be addressed.

      (1) Which is the mechanism by which Wun/Wun2 attenuates Hh signalling? The authors propose that Wun/Wun2 block Hh ligand transmission, but their data could also be explained by other possibilities, such as altered Hh production, uptake, retention or degradation, among others. The authors should either show the effect of Wun/Wun2 in Hh transmission mechanistically or attenuate their claim.

      (2) How do Wun/Wun2 attenuate Hh signalling in PGCs? The authors propose that Wun/Wun2 function both in somatic tissues and in PGCs, but these two sites of action may have very different mechanistic implications. In the soma, Wun/Wun2 could affect Hh transmission, but a PGC-autonomous role cannot be explained simply by reduced Hh ligand transmission from producing cells; it would more likely involve ligand uptake, receptor trafficking, intracellular degradation or altered PGC responsiveness. This distinction should be central to the interpretation of the data.

      We thank the reviewer for recognizing the importance of the problem and we are sensitive to both the points of criticism regarding the mechanism(s) Wunen(s) may employ to downregulate Hh signalling.

      The reviewer correctly pointed out that we singled out Hh transmission as the putative target of Wunen(s) which need not be the case. We agree with this assessment and would like to thank the reviewer for pointing us in the right direction(s). Indeed, Wunen(s) could act at several different levels to regulate Hh signalling including “Hh production, uptake, retention or degradation”. We will modify the text to incorporate these possibilities in the appropriate sections of the manuscript.

      The only reason for the emphasis on the ‘Hh transmission’ in the text was to contrast it with Hmgcr which acts in a qualitatively opposite manner. Hmgcr potentiates Hh signalling by altering the range/strength of the Hh ligand in the embryonic context. This was also confirmed in the wing discs and adult wings as hmgcr mutants could dominantly suppress the wing duplications and abnormalities induced by the ‘gain of function’ allele of hh (hh<sup>MRT</sup>). Upon compromising hmgcr, Hh ligand was shown to be sequestered in the Hh producing cells in the ectoderm. However, we have not carried out similar experiments to either rule in or rule out the different possibilities suggested by the reviewer. We will ensure that the claims made in the manuscript will appropriately reflect the scope of the analysis and the related arguments will be suitably modified.

      The reviewer also makes a very critical point regarding cell autonomous v/s cell non autonomous activities of Wun(s). We have briefly mentioned the possible role of individual Wun(s) in the SGPs/mesoderm as well as within the PGCs. It has not escaped our notice that Wun(s) could regulate Hh internalization within the PGCs or its subcellular compartmentalization (within the ER, golgi or lysosomes). Wunen(s) could also act at the level of Hh reception by changing the activity/localization of Hh receptors, either Smoothened or Patched and could influence the outcome of the signaling pathway in a multi-pronged manner.

      We appreciate the thoughtful suggestions and as recommended, future analysis will focus on these aspects. In our view, data included in the present version of the manuscript are novel and sufficient to argue a functional relationship between Wun(s) and Hh signalling which is qualitatively antagonistic to Hmgcr.

      Reviewer #2 (Public review):

      Summary:

      In this submission, Roy et al. examine the process of Drosophila PGC migration. Directed cell migration requires the concerted activities of chemoattractants and repellents to guide cells to the correct locale. In their submission, the authors describe a role for regulated Hedgehog (Hh) signaling to inform PGC migration. In prior work, the authors reported that Hmgcr potentiates Hh signaling, providing a permissive axis. A gap in the field, however, was the identification of the repulsive cues that guide PGCs out of the midgut and toward the future gonad. In the current work, the authors report that two wunen genes (wunen and wunen 2) inhibit Hh signaling, thereby repressing Hh activity. The model is that Hmgcr and wunen(s) balance the transmission of Hh signals to enable effective PGC migration.

      Strengths:

      A strength of this work is the comprehensive genetic analysis performed by the authors. The authors examine zygotic versus maternal contributions, autonomous versus non-autonomous requirements, and use a variety of RNAi and mutant allele combinations to examine genetic requirements and interactions. Another strength is that the data presented are generally clear and well quantified. Insets are provided to enhance visualization, and relevant data are quantified through replicated experiments.

      Weaknesses:

      Weaknesses of the work include a lack of biochemical data to validate some of the proposed interactions. Although the authors do report lipidomics data, little is done with these findings to validate or place the results in the context of a mechanistic model. Despite these issues, the conclusions stated are generally well supported by the results.

      We would like to thank the reviewer for their positive feedback and a succinct description of the findings reported in the manuscript.

      We agree that the mechanistic basis of DAG accumulation was not explored in this manuscript. Prior work in the Ratnaparkhi and Kamat labs identified a Serine hydrolase that functions as a phospholipase C in biochemical assays (Kumar et al., 2024, Biochemistry 63:3000-3010). We have since conducted several genetic experiments, and preliminary data indicate that, in the embryonic context, mutations in the specific Phospholipase C display phenotypes analogous to wun(s). We hope to present these data along with the comparative molecular and biochemical analysis in the near future.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      In this manuscript, the authors present a method to detect natural selection on transcription factor binding sites (TFBSs), which is an upgraded version of a previously published method (Liu and Robinson-Rechavi, 2020). This upgraded version of the test implements more explicit models of evolution and is shown to outperform its predecessor in terms of both power and false positive rate. I think this method can be a valuable resource for the community and can be helpful not only to studies of TFBSs but also broader evolutionary questions related to genotype-phenotype maps or fitness landscapes.

      Major comments:

      (1) Questions related to Figure 1

      Figure 1, along with the first section of the Results, shows that the SVM score and its sensitivity to mutations are generally correlated with the strength of ChIP-seq signals. It is not very clear to me, however, what the motivation is behind this part of the paper. It seems that the model used to predict binding strength is a pre-existing one, and it is unclear what is new in this section. Was the prediction model retrained using different data? Was its validity confirmed using new data? I would appreciate some more elaboration on how these results differ from what was presented in the previous study of Liu and Robinson-Rechavi (2020).

      We agree that the current manuscript does not clearly distinguish which parts of Figure 1 are novel and which are foundational. The SVM itself is not new and is the same as in Lee et al. (2016), as used in Liu & Robinson-Rechavi (2020). In the revision, we will explicitly state that the SVM used in Figure 1 is the standard gapped-kmer SVM (ls-gkm) approach. We retrained all gkm-SVM models de novo for each species-TF dataset, ensuring consistency across all analysed ChIP-seq peaks. For this, we recalled all ChIP-seq peaks in a homogeneous and robust manner using the nf-core ChIP-seq pipeline v2.0 (Ewels et al. 2022). Figure 1A confirms that the predicted binding affinity from the SVM correlates with experimental ChIP peak height. In addition, examining scores per site rather than per peak is new compared with Liu and Robinson-Rechavi (2020). The correlations between the SVM-derived scores and other features had not been shown before to the best of our knowledge, thus Figure 1B-C is entirely novel. In other words, this analysis is meant to show that our phenotypic metric (SVM score per site) indeed tracks binding intensity, i.e. molecular phenotype.

      The existence of weak or negative correlations between SVM and coverage, which reportedly reflects low-quality peaks, seems applicable not only to this paper, but also to previous ones, so I would like to have it confirmed whether the question and the authors' answers apply to previous studies as well.

      Yes, this is a well-known issue in ChIP-seq studies. Low coverage often matches weak predicted binding affinity scores because noisy or unreliable peaks naturally have weaker signals. This is not specific to our work, and it has been observed in many other studies (e.g., Bailey et al. 2013 doi:10.1371/journal.pcbi.1003326; Nakato and Shirahige 2017 doi:10.1093/bib/bbw023). It is simply an expected property of the data.

      It is reported that SVM scores capture TF binding signals better than conservation-based statistics do. My intuitive interpretation is that both ChIP-seq peaks and SVM scores are supposed to reflect binding strength, whereas conservation is supposed to reflect selection (i.e., different definitions of "function" as mentioned above). It is not explicitly explained in the Results, however, what the difference indicates, leaving only an impression that the SVM score is "better" than the conservation statistics.

      While the reviewer is correct that there are different definitions of function, both conservation-based statistics and RegEvol seek to capture selected function. The difference is that RegEvol aims to measure functional change, whereas conservation-based statistics aim to detect sequences that retain the same function across species. In both cases, we expect a correlation with causal function (i.e., binding). We will clarify these concepts and how they apply to our results in the revised manuscript.

      (2) Lack of directional selection for low binding affinity

      In the analysis of Drosophila melanogaster ChIP-seq peaks, there were more cases of directional selection for higher binding affinity than directional selection for lower binding affinity. The authors suggested that this observation is "likely biological" because the same pattern was not seen in simulations (line 412-413). I wonder if this could have resulted from a difference in the distribution of ancestral binding affinity across TFBSs between real and simulated data. If binding affinity was generally low in the common ancestor of D. melanogaster and D. simulans, selection for low binding affinity would manifest mainly as purifying selection against mutations that increase affinity instead of directional selection. Ancestral sequences for simulations, if I understood correctly, are observed peaks in D. melanogaster (line 715-719), which would include high fraction sequences that could be rarer in the real ancestral sequences.

      The description of this particular result does not refer to a figure or table, nor is it revisited in the Discussion. Figure 5 treats peaks under directional selection as a single category. Taken together, it is hard to tell how this observation should be interpreted. If the authors consider this result as biologically meaningful, I would suggest adding more details (e.g., the number of each side).

      We appreciate this insight. We agree that the text was not clear, but in fact, the simulations were performed using the reconstructed ancestral sequences of ChIP-seq peaks themselves. Thus, simulated and empirical results should be directly comparable, and different results should be due to biology. We will revise the Manuscript to explicitly state that simulations are performed from reconstructed ancestral sequences and why. We will also add more descriptive statistics of the simulated and real data.

      (3) Selection in non-focal lineages

      Regarding the detected signals of directional selection for stronger binding in certain tissues (Figure 6), I wonder if it is the focal species or those very tissues that are "special": did the human lineage undergo more adaptive regulatory evolution than the chimpanzee lineage, or do nervous and male reproductive systems have a high "propensity" for adaptive regulatory evolution? Assuming that the binding preference of the same TF did not undergo a significant change since human-chimpanzee split (which, I believe, is a built-in assumption in both RegEvo and the permutation test), it should be possible to perform the same test using chimpanzee sequences that are homologous to the human ChIP-seq peak regions. In the case of coding sequences, for example, Bakewell et al. (2007) found that it was the chimpanzee that had more genes under positive selection than humans; I wonder if TFBSs show the same or a different pattern.

      This is an excellent suggestion. To compare in an unbiased manner, we would need transcription factor ChIP-seq from the same organs in chimpanzees and humans. We are not aware of such a dataset. If one is identified, we would be very interested in analysing it, and thus answer this question. As suggested by the reviewer, we will analyse the human homologous sequences. Although it should be clear that this will provide a biased estimate for comparing adaptation between the two species, as we will lack newly acquired binding sites in the chimpanzee.

      (4) Comments on terminology

      (a) Meaning of "function"

      The word "function" has had different meanings in the biology literature, with some authors using "functional" to refer to anything with a phenotypic effect and some using it only for targets of selection. A (putative) TFBS would be considered "functional" as long as it has TF binding affinity if we follow the effect-based definition, but only if its binding affinity is under selection if we follow the selection-based definition. In this manuscript, the term "function" appears to have been used to refer to TF binding but not selection, most notably in the first Results section. There are also places where it is less clear what "function" means exactly (e.g., "deeply conserved elements that are likely to be functionally important" of line 61). Since this paper is about evolution, it is likely that many readers prefer the selection-based definition or assume that the selection-based definition would be used. Thus, using "function" to refer to just TF binding could be confusing. To this end, I would suggest that the authors drop the word "function" or give an explicit definition early in this paper.

      We thank the reviewer for this precision and fully agree, we will revise our terminology for clarity. We will clarify the distinction between selected function and causal function, and we will pay attention to their use throughout the manuscript.

      (b) Directional selection in different directions

      In this paper, selection for increased TF binding affinity is referred to as "positive directional selection", and selection in the opposite direction is called "negative directional selection" (as exemplified in Figure 2). I understand that using such shorthand names would make the text less clumsy, but these two terms could potentially be confusing, as "positive selection" and "negative (purifying) selection" are also terms referring to specific types of selection and have some connection to directional and stabilizing selection. Therefore, I suggest that the authors use something like "selection for increased/decreased binding affinity" instead, or note explicitly in the text that "positive/negative directional selection" would be used as shorthand.

      We agree with this ambiguity in the current terminologies. We will replace the phrases “positive directional selection” and “negative directional selection” with, e.g., “selection for increased binding affinity” and “selection for decreased binding affinity” as suggested when presenting our biological result on ChIP-seq peaks. However, we will still use “positive/negative directional” for the general framework (genotype → phenotype →fitness map) and insert a note that we use “positive/negative directional” as shorthand to mean increasing/decreasing affinity in the case of CHIP-seq peaks.

      Reviewer #2 (Public review):

      Summary:

      The manuscript by Laverre et al. provides an interesting new test of selection on TF binding. Rather than focusing on sequence changes, this test is specifically for changes in predicted TF binding affinity. The authors report directional selection on 5.1% of tested regions in Drosophila, as well as a signal of selection on CTCF binding in the human CNS and male reproductive system.

      Strengths:

      Overall, I think this represents an important direction for the field of molecular evolution: now that TF binding can be predicted fairly well from sequence, it can be a very useful focus for tests of selection.

      Weaknesses:

      As mentioned several times in the manuscript, Jiang and Zhang (2024) pointed out some issues with a previous permutation-based version of this test. Foremost among these was the issue of ascertainment bias: when testing only experimentally supported TF binding sites from a focal species, and then asking what type of selection (or lack of selection) led to those sites, one is guaranteed to find more substitutions that increase affinity, simply because the sites were selected in the first place as those with maximum (empirically measured) affinity.

      To address this issue, the authors simulated Drosophila CTCF peaks evolving neutrally and then tested different ascertainment cutoffs in Figure 4D. It was not entirely clear to me what is shown in Figure 4D: the text says the bins were stratified by derived delta-SVM, whereas the figure says SVM, and the legend says derived SVM (both without the delta). I was unable to find any clarification of this in the Methods section. In any case, I am not really convinced by his, for two main reasons. First, when analyzing empirical ChIP-seq data, I would guess that only a tiny fraction of the genome is bound (far less than 1%, especially in mammalian genomes). However, the most extreme bin in Figure 4D is taking the top 10% of (delta?) SVM values. What would Figure 4D look like at bins of the highest 0.1%, 0.001%, etc? My guess is there would be a strong uptick in the FPR.

      We apologise for the confusion in Figure 4D, we will clarify the caption and text and specify that bins are stratified by derived SVM (post-simulation binding affinity proxy), not genome % or ΔSVM.

      We want to note that we used the same subsampling approach as Jiang and Zhang (2024) to evaluate ascertainment bias, and that Figure 4 both confirms the issue that they identified with Liu and Robinson-Rechavi (2020), and shows very clearly that RegEvol does not have the same issue (flat red lines). Following the reviewer's suggestion, we can extend the figure to 1% or 0.1% bins. We note that the % of the total genome is different from the % of peaks: while actual peaks cover a very small proportion of the genome, the subsampling in Figure 4 (and in Jiang and Zhang 2024) aims to estimate the impact of detecting only the strongest peaks.

      One difference between Jiang and Zhang (2024) and our study is that we simulated using whole empirical peaks, whereas they simulated 10-nucleotide transcription-binding sites, meaning that each substitution represented a 10% change. We will clarify these differences in the revised text.

      The second reason is actually more important and fundamental than the first. As long as this method is working as described, I cannot see any way that it would ‘not’ be impacted by ascertainment bias. As an extreme case, imagine that all TF binding sites tested had the maximum possible SVM scores; then none of them would have any chance of showing directional selection against binding, while even those that evolved neutrally would appear to have directional selection in favor of binding. Of course, real empirical data are not as extreme as this, but the same concept applies in less extreme scenarios.

      This bias could explain patterns observed in the real data. For example: "We observe much more positive than negative directional selection, a pattern likely biological rather than methodological, since it is absent from simulations." This is exactly the pattern predicted under ascertainment bias (in the extreme-scenario thought experiment above). I suspect it is absent from simulations simply because the authors did not properly account for this bias in their simulations.

      If the main result reported by the authors had been a lack of any directional selection in favor of binding, and instead only neutrality or directional selection against binding, then this ascertainment bias would not be an issue- it would only have made their results conservative. Unfortunately, this is not the case, and the directional selection in favor of binding, which is the main result emphasized from the empirical analysis, could be inflated by this bias.

      There is indeed a possible ascertainment bias, although we believe it concerns only the detection of negative directional selection, as long as we have only empirical peaks in the focal species and not the sister species. This is not so much a limitation of our method as an intrinsic limitation of asymmetrical sampling of species: to study both gain and loss of function, function must be studied experimentally in several species. We will revise the manuscript to highlight this limitation.

      Concerning positive directional selection, the mathematical foundation of RegEvol makes it inherently robust to ascertainment bias for positive directional selection. RegEvol calculates the likelihood of the entire sequence of observed substitutions accounting for the starting ancestral state and the mutational landscape. In other words, the model does not assume a uniform probability of phenotypic change; instead, it models the probability of each nucleotide mutation to result in a substitution (i.e., go to fixation) depending on its phenotype.

      In an extreme case where all tested TF binding sites had the maximum SVM score, detecting negative directional selection would indeed be impossible, as ancestral states would have had equivalent or lower scores. However, positive directional selection would be inferred only if the likelihood of observing the substitution pattern’s deltaSVM distribution significantly exceeded that expected under the mutational landscape. If a sequence evolved neutrally but reached a maximum SVM score, the likelihood of detecting directional selection would depend on: either the ancestral state being close to maximum with few substitutions increasing SVM (resulting in low statistical power), or the ancestral state being distant with many neutral substitutions and rare chance shifts to maximum (where the substitution distribution would be indistinguishable from neutrality). Then, even in such an extreme dataset, neutral evolution remains detectable, demonstrating RegEvol's strength beyond deltaSVM comparisons between two states.

      Minor point:

      The following statement: "In contrast, phastCons and phyloP scores lack such enrichment and have a lower dynamic range, suggesting that the conservation scores are less sensitive to fine-scale variation of TF occupancy and thus regulatory region function" is only true if one assumes that TF binding is the only function of this region. One could even turn this around and say the fact that the sites affecting TF binding are not the most conserved is actually evidence that TF binding is not a good indicator of these regions' entire function. I suggest the authors soften this claim that conservation scores are less sensitive to regulatory region function.

      We thank the reviewer for this comment, the text will be revised to soften this claim. We will explicitly state that sequence conservation reflects general functional constraints, whereas sequence-to-phenotype predictions capture highly specific and lineage-specific TF-DNA interactions.

    1. Author response:

      Reviewer #1 (Public review):

      Summary:

      The authors aim to understand, in the context of leaf shape, how the constraints imposed by development inform evolution. Leaf shape is a good place to study the influence of development on evolution because it is a trait that exhibits a lot of diversity, and the developmental mechanisms that give rise to leaf shapes are apparently rather conserved across angiosperms.

      As part of the motivation for their work, the authors cite a previous study (Geeta et al), which found that in angiosperm phylogenies, transitions from complex to simple leaf shapes occur through evolution more often than transitions in the opposite direction. Is this due to developmental constraints or adaptation?

      The authors undertake two parallel lines of work:

      (1) Extending the study of Geeta et al with more data, consisting of both phylogenies and a shape classification dataset. The conclusion from this line of inquiry is that transitions from lobed to unlobed leaves are more common than transitions away from unlobed leaves.

      (2) The authors conduct evolution simulations in a computational model of leaf development. Here, they look at {\it neutral} mutations and whether simply neutral evolution is sufficient to drive the observed trend.

      The conclusion of the second part of the work is that the driver of the evolution toward simple leaf shape is entropy: there are more ways to make unlobed leaves than to make lobed leaves (at least in terms of gene regulation parameters that will produce the two leaf types). The argument is that random gene regulatory networks are more likely to produce unlobed leaves than lobed leaves; therefore, neutral evolution drives this trend.

      Data Analysis

      Roughly $9000$ images of leaves were classified into 4 categories: unlobed, lobed, dissected, and compound. These labels were applied to the tips of 5 phylogenetic trees of angiosperms (3 resolved at the genus level and 2 at the species level). By fitting a continuous-time Markov chain to the labelled trees, the authors claim that there is a significantly higher rate of transition to the unlobed leaf shape compared to transitions to more complex shapes.

      Simulation

      First, the authors validate a computational model (Runions et al) for leaf growth on an experimental dataset. By changing parameters in the model, they can recapitulate the morphological changes in the shapes of Arabidopsis leaves engendered by expression of two particular genes.

      Then the authors run an evolutionary model (without selection, just random mutations) on top of the computational leaf development model. As the random walk in parameter space reaches a stationary distribution, they look at both the proportions of the leaf categories in the steady state as well as the transition rates between different categories. The result is that transitions to unlobed leaves are more common than from unlobed leaves.

      We thank the reviewer for the helpful and clear summary of our work.

      General Comments

      The authors use angiosperm phylogenies from other works as the basis for the data analysis part of their work. Given the centrality of these phylogenies for their conclusions, more information is needed about how these phylogenies were constructed and what they mean. What is the timescale that they span? What method is used to infer them? What regions of DNA were sequenced in order to build the phylogenies? Also, maybe some more discussion of angiosperm evolution (e.g., when was the most recent common ancestor of all angiosperms?) would help put the study in context.

      We also need a more in-depth discussion of the computational model. What are all the $>100$ parameters doing, and what informs the seemingly strange mutational model that changes parameters by 3 orders of magnitude?

      I am confused about how the rates of transitions were inferred from the phylogeny. Here, one has a phylogeny inferred by some method (which needs to be described in more detail), and just the leaves are labelled. It is stated in the methods that BayesTraits was used to infer the transition rates. I realize this method is probably documented elsewhere, but a bit of a summary of how it works and how to interpret its results would (1) make the paper more selfcontained and (2) if the algorithm is credible, make the results firmer.

      We thank the referee for the suggestion to make the paper more accessible. The tool we use to infer transition rates from the phylogenies, BayesTraits, is standard in the field. However, the referee is right that for an interdisciplinary journal, it may be helpful to more fully flesh out how these methods work. To that end, we have added an additional section "Phylogenetic rate inference" in the supplementary information that includes a longer description of how BayesTraits works, and how we used it to infer transition rates from phylogenies.

      All trees are shown in the supplementary information section "Phylogenetic trees" with scale-bars showing the amount of time or genetic change that the trees span. For a broader discussion of angiosperm evolution, there is supplementary information section "The adaptive significance of leaf shape review".

      Regarding the more in-depth discussion of the computational model, we have added supplementary information section S1 "Leaf model details" to give a more detailed description of the leaf model.

      I am a bit skeptical of the authors' interpretation of the biological trend (of complex to simple leaf shapes) as being driven by neutral evolution. Why does one expect that the mutations generated by the random walk models described in the work are in fact neutral mutations?

      A random walk is a well-established way of modelling the dynamics of neutral evolution in the monomorphic regime, where the population has a narrow diversity of different genotypes. In the higher mutation rate polymorphic regime, where the diversity of genotypes in the population is larger, we also expect that a random walk should still recapitulate the correct average transition rates. The purpose of the simulations is not to model every aspect of population genetics, but to ask whether developmental bias alone is sufficient to generate the observed directional asymmetry. By assigning equal fitness to all viable leaves, we isolate the contribution of development from that of selection. The agreement with the phylogenetic transition rates therefore demonstrates sufficiency rather than exclusivity: selection may also contribute, but it is not required to explain the observed bias We discuss the evidence for the role adaptation in leaf shape further in supplementary information section "The adaptive significance of leaf shape review".

      If the entropy of simple leaf shapes is higher than that of complex leaf shapes, why did we have complex leaves at all? I suspect the authors might argue that this is due to selection. In that case, what allows these complex shapes to become simpler? Wouldn't they be losing the selective advantage that drove them to be more complex in the first place? Or maybe the idea is that the rates are inferred assuming some steady state that generates the phylogeny? I did not understand this point.

      The entropy language is a useful framing. Within that framework, one can view our study as showing that the entropy (defined here as the logarithm of the volume of parameter space mapping to a phenotype) of simple leaf shapes is higher than that of complex leaf shapes. If this entropy were to be ignored, then all states would be equally likely in our simulations, where we do not take fitness differences into account. What we show is that the differences in entropy -- related to differences in volumes of the parameter space that map to different phenotypes -- also affects the rates. The inferred transition rates for both simulation and phylogeny from unlobed to more complex shapes are lower than vice versa but not zero. Therefore, complex leaf shapes arise stochastically through mutation and in this model would eventually reach a steady state proportion, even in the absence of selection.

      Are the rates of transitions between leaf types inferred for the phylogeny assuming that the phylogeny is generated by the steady state of some Markov process? (I think the answer is no: in that case, how does one explain the initial condition?)

      The tool we use to infer transition rates from phylogenies—BayesTraits—allows the initial state at the root of the tree to vary during the numerical optimisation (Pagel, 1994). Therefore, it is not assumed that the initial state is generated by the steady state of the Markov process.

      If I take the mutation model (random walk) seriously, then shouldn't I expect that this steady state obeys detailed balance? In that case I should have $p_i r_{i\to j} = p_j r_{j\to i}$ for each of the occupancies $\{ p_i\}$ and transition rates $r_{i\to j}$ for the shape categories. How close are the rates inferred from the phylogenies to obeying detailed balance? Presumably, the Markov chain fitted to the simulation data obeys detailed balance because the mutation model itself does?

      BayesTraits allows off-diagonal transition rates of the rate matrix to vary freely during numerical optimisation (Pagel, 1994). Therefore, there is no requirement for the detailed balance to hold for the inferred rate matrix. For our simulations, the mutations are symmetric at the parameter level, therefore at this level, the process would be expected to obey the detailed balance.

      I find it hard to take the discussion of development seriously without some consideration of mechanics. Presumably, the mechanics are hidden in the computational leaf development model, but this model is not discussed in enough detail for the reader to know. It seems to me that the interesting question is: what are the {\it mechanical} constraints on development that drive the apparent trend in evolution towards simpler leaf shapes? Maybe it is something about the type of differential growth needed to make complex leaf shapes less robust to mutation. But in this case, I would assume that selection plays a role in the complexity of shape. In any case, a better understanding (or explanation) of the computational model is needed to make this interpretation.

      We thank the referee for the suggestion to make the paper more accessible. We have added a more detailed and pedagogical description of the model from (Runions, Tsiantis and Prusinkiewicz, 2017) in the supplementary information section S1 "Leaf model details". We also note that Fig. 5 in the methods that gives an overview of how the model works, including some mechanical aspects of development and growth.

      More generally, mechanics is one component of the developmental map that determines which parameter combinations produce viable leaf morphologies. Our analysis concerns the geometry of this complete developmental map, irrespective of whether its constraints arise from gene regulation, tissue mechanics, or their interaction.

      On the interesting question of what is causal, perhaps the example in figure 2 is helpful. We focus on two parameters, a morphogen repression strength, and a duration of growth. A key physical process here is called webbing, where cellular growth fills in the gaps between branching veins. This process flattens the leaf structure and creates a continuous, solid leaf blade (lamina). Strong webbing, characterized by a significant resistance to stretching and bending, results in a smoother margin (Runions, Tsiantis and Prusinkiewicz, 2017). The morphogen repression strength affects the physical parameters that determine how strong the webbing is. The duration of growth determines how long the leaf has to grow. Varying these two parameters varies the physical processes that determine leaf shape. The mechanics of growth operate downstream of these parameters that we vary in our evolutionary simulations according to the details of the leaf developmental model.

      Some discussion of timescales is needed, especially when invoking neutral evolutionary arguments. If a neutral mutation occurs, its time to fix in a population of size $N$ is $\sim N$ generations. What are the relevant angiosperm population sizes and the number of mutations that separate branches on the tree? Are timescales remotely consistent with e.g., the age of angiosperms on Earth?

      Neutral processes have a well-established role in key aspects of angiosperm evolution, for example genome complexity (Lynch and Conery, 2003). This would suggest that the relevant time scales and generation times are not completely prohibitive of neutral processes also playing a role in the evolution of angiosperm leaf shape. Effective population sizes in plants are highly variable but estimates span 10^3-10^6. Assuming diploidy (and therefore average fixation time of 4Ne) and generation times of 1-10 years, this gives fixation timescales of 10^3-10^7 years. This is within the timescales of the trees we analyse, which span >150 million years.

      Reviewer #2 (Public review):

      Strengths:

      The paper's underlying question is interesting, extending the authors' prior work on RNA along similar conceptual lines. The paper combines both image analysis of leaves and a computational analysis of a simple model of leaf development.

      Weaknesses:

      The entire paper is based on the Runion model. More intuition about the Runion model would be useful for a broader readership that cares about the evolutionary aspect of this, but may not know the developmental model in question. Obviously, this is prior well-established work, but 2 - 3 sentences highlighting the key structural aspects of such a model would be great. Currently, that intuition is found implicitly in a sentence on page 2 ("complex leaf shapes need more specificity in their GRNs than their simpler unlobed leaf shape"), but the reader is left wondering - is the Runion model a detailed mechanistic one with multiple interacting genes/proteins? If so, how many? Or is it just 2 - 3 genes but with complexity entirely in how long they are each expressed/when they are turned off, etc.

      We thank the referee for the suggestion to make the paper more useful for a broader readership. To that end, we have added a more detailed description of the (Runions, Tsiantis and Prusinkiewicz, 2017) model in supplementary information section S1 "Leaf model details".

      The Runions model has nearly 100 free parameters. Random walks in 100dimensional spaces have generic properties like a tendency to move toward regions of larger volume that have nothing to do with leaf biology. How do you disentangle the geometry of high-dimensional random walks from genuinely biological developmental bias? Would a toy model with 100 parameters and arbitrary phenotype categories also show "bias toward simplicity" if "simple" phenotypes occupy more volume?

      Our argument is largely independent of the number of parameters. While it is true that most of the volume is near the surface in a high-dimensional space, our argument is about the relative volumes of the sets of parameters that map to each of the four phenotypes, an entropic argument if you wish. The basic intuition is that a simple phenotype needs fewer parameters to be fine-tuned, and so a larger volume of parameter space will map to a simpler phenotype.

      The question about a toy-model with arbitrary phenotypes is helpful, because it allows us to clarify that what we are illustrating here with the biologically realistic example of leaf shapes is a much more generic principle. We can say with confidence that if the toy-model generates a many to one set of outputs (phenotypes) through an algorithmic process whose description length does not grow faster than logarithmically with the size of the genotype space, then it should produce a bias towards simplicity regardless of the number of dimensions, see for example Johnston et al. (2022) and Dingle, Camargo and Louis (2018) for a longer discussion of this more general point which is based on arguments from algorithmic information theory (AIT). We don’t use that framing in the current paper because the basic intuition for GRNs that more complex phenotypes need more parameters fine-tuned, and so have relatively smaller volumes, is more straightforward to understand that the more abstract AIT arguments. Our general prediction that this principle should hold more widely for GRNs can be made both by the more formal AIT route, or via the more heuristic fine-tuned parameter route.

      The discussion of Figure 4 (PCA of parameter space) uses "area" loosely when what's actually being measured is bin count in a 2D projection of a highdimensional space. I would think that, in general, PCA projections can be misleading about volume in the full parameter space, but I can't tell if that's an issue in this case. Some comments/thoughts here would be useful.

      The quantitative estimate of phenotype frequencies is computed directly in the full parameter space and does not depend on PCA. Ie. We estimate that the total volume of viable leaves maps to simple unlobed leaves about 80% of the time. However, the volume is extremely high-dimensional, and so hard to visualise. PCA is used solely to provide an interpretable visualization of this otherwise high-dimensional structure. The PCA plots in Fig 4 and Fig S16 are there to be illustrative, not quantitative. Because the volume differences are large, we do not think that the projections of the main PCA components would be misleading on at least the ordering of the sizes of the parameter space components that map to each leaf shape. We provided a similar analysis for other projections -- PC1-PC6 (supplementary information section "PCA occupancy for higher dimensions"), finding the same trend. To make this point clearer, we have now changed the sentence in the Fig. 4 caption slightly “This (reveals that --> illustrates how) unlobed leaves occupy a larger region of model parameter space than more complex shapes and that this larger space also contains the majority of more complex leaves.”

      The classifier validation section is in the Methods section, but it seems critical to the whole story. The < 80% agreement with manual classification could propagate to the rest of the estimates in the paper. Again, some comments/thoughts here would be useful.

      We have repeated the analysis of the agreement between by-eye and automatic morphometric classification. Generating a confusion matrix for the two classification methods shows that the agreement is high for unlobed, dissected and compound, with the main source of disagreement being leaves that were classified as lobed by-eye being classified as either unlobed or dissected by the automatic-morphometric method. The proportion of by-eye lobed leaves classified by the automatic morphometric method as either unlobed (27%) or dissected (23%) is relatively balanced, which we think will help cancel out some error as well. Moreover, we find that the agreement between the automatic-morphometric method and by-eye classification increases to 90.0% when using the categories unlobed and all other categories grouped into one. This is the most important classification for our finding that development and phylogeny are both biased towards unlobed.

      The authors should explain Mut2 and Mut5 in the main paper with a sentence or two, at least schematically, because how you mutate is obviously very relevant to interpreting a paper about biases in variation.

      In the results section we have added a sentence for more detail on the random walk.

      "[We mutated the initial sample using a random walk algorithm with two different mutational schemes, MUT2 (alg. 1) and MUT5 (alg. S2).] These algorithms work by iterating through model parameters one by one and perturbing the value by a small amount. We then [automatically classified the resulting shapes...]"

      Moreover, in methods section C there is already a more detailed description of both algorithms.

      “MUT2 (alg. 1) iterates through the parameters in a random order, and attempts to change the parameter by a value selected at random from an array of numbers randomly generated at 3 different orders of magnitude. MUT5 (alg. S2) is the same as MUT2 except the value each parameter is multiplied by 10% of the range of that parameter within the initial leaves (fig. S1). The aim here was to provide some way of accounting for the biologically relevant sampling range. "

      Moreover, the MUT2 algorithm is described in pseudocode in Algorithm 1 in the main text, and the pseudocode for MUT5 is in supplementary information section S1 C, as algorithm S2.

      The two mutational schemes use additive perturbations to individual parameters. Real mutations presumably affect regulatory networks in more structured ways (e.g., changing binding affinities that affect multiple parameters simultaneously). How sensitive are the results to the assumption of independent single-parameter mutations?

      The referee raises an interesting and well-known issue concerning this widely studied class of GRN models. Without a detailed understanding of how individual genetic mutations map onto model parameters, it is difficult to determine with confidence whether a mutation would produce correlated changes in certain sets of parameters. Our main argument, however, is that the primary source of the observed bias is geometric: the volume of parameter space (or equivalently, the entropy) corresponding to simple leaf morphologies is substantially larger than that corresponding to complex morphologies. As long as mutations explore parameter space approximately symmetrically, even if they involve correlated changes in multiple parameters, larger phenotype regions will tend to be encountered more frequently and retained for longer than smaller regions. We therefore expect the observed bias to be robust to many alternative mutation models, although quantifying this robustness is an interesting direction for future work.

      The connectedness argument is made using a 2D PCA projection. Is there a way to check this statement in the full parameter space or perhaps in higher dimensional projections to test the robustness of this result? Connected components can merge/split under different projections.

      Constructing the nearest neighbour graph for the full dimensional data results in the following no. connected components: unlobed-146, lobed-274, dissected-255, compound-315. This follows the same pattern identified for the PC1-PC2 projection, that unlobed splits into fewer connected components than other leaf shape categories.

      References:

      Dingle, K., Camargo, C.Q. and Louis, A.A. (2018) ‘Input–output maps are strongly biased towards simple outputs’, Nature Communications, 9(1), p. 761. Available at: https://doi.org/10.1038/s41467-018-03101-6.

      Johnston, I.G. et al. (2022) ‘Symmetry and simplicity spontaneously emerge from the algorithmic nature of evolution’, Proceedings of the National Academy of Sciences, 119(11), p. e2113883119. Available at: https://doi.org/10.1073/pnas.2113883119.

      Lynch, M. and Conery, J.S. (2003) ‘The Origins of Genome Complexity’, Science, 302(5649), pp. 1401–1404. Available at: https://doi.org/10.1126/science.1089370.

      Pagel, M. (1994) ‘Detecting correlated evolution on phylogenies: a general method for the comparative analysis of discrete characters’, Proceedings of the Royal Society of London. Series B: Biological Sciences, 255(1342), pp. 37–45. Available at: https://doi.org/10.1098/rspb.1994.0006.

      Runions, A., Tsiantis, M. and Prusinkiewicz, P. (2017) ‘A common developmental program can produce diverse leaf shapes’, New Phytologist, 216(2), pp. 401–418. Available at: https://doi.org/10.1111/nph.14449.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer 1 (Public review):

      Summary:

      This study presents a systematic investigation of parent-of-origin effects on gene expression using trio-based data from the Framingham Heart Study, which is notable for its relatively large number of trios. By combining whole-genome and RNA sequencing data, the authors examined the extent to which gene expression is influenced by whether genetic variants are inherited maternally or paternally.

      The authors report that parent-of-origin eQTLs are widespread, identifying 15,893 eQTLs from 14,733 variants and 1,824 genes that were significant in paternal, maternal, or joint tests but not detected by traditional eQTL approaches. They further classified these associations based on the relative strength and direction of paternal and maternal effects, highlighting a subset with opposing directions. The study also highlighted eGenes linked to known imprinted genes as well as those with opposing parent-specific effects, and observed that paternal eGenes are enriched for drug targets. Finally, the work revisits previous findings in which eQTL studies were used to interpret disease-associated loci, emphasizing that conventional eQTL analyses without testing the parent-of-origin may mislead gene prioritization efforts. The study recommends that future downstream analyses, such as Mendelian randomization, take into account the provided lists of SNPs and eGenes and exclude those with strong parent-of-origin effects when linking genetic regulation to disease risk.

      Strengths:

      The major strength of the study lies in the scale and quality of the dataset, the trio-based design, and the systematic application of statistical tests for parent-of-origin effects. The strengths thoughtfully employed Bayes factors rather than p-values to provide stronger evidence of association, which adds rigor to their analyses. These design choices provide compelling evidence that parent-of-origin effects are widespread and that conventional eQTL analyses miss a substantial fraction of regulatory variation. The results are clearly presented and supported by robust analyses, including the identification of opposing parental effects and the enrichment of paternal eGenes for drug targets. Notably, the two examples demonstrating how these findings can reshape disease gene prioritization highlight the broader impact of the study and encourage further work in the community to incorporate parent-of-origin effects.

      Weaknesses:

      The main limitations of the study are threefold.

      First, there is a lack of replication in independent cohorts, which is understandable given the difficulty of identifying datasets with a comparable number of trios, but replication would help establish the generalizability of the findings.

      We fully agree with the reviewer that replication in an independent cohort is a crucial step for establishing generalizability. As the reviewer notes, the Framingham Heart Study, with its 1,477 trios possessing both WGS and RNA-seq data, represents a uniquely powerful and, to our knowledge, currently unmatched resource for this specific type of parent-of-origin eQTL analysis.

      In the absence of an external cohort of comparable size and data richness, we have taken several steps to ensure the internal validity and robustness of our findings within the current study, which we will clarify and expand upon in the revised manuscript:

      Positive Control Validation: We explicitly used well-established, bona fide imprinted genes (e.g., MEG3, NDN, SNURF, as listed in Table 1 and Figure 1) as positive controls. The fact that our analysis correctly identifies their known parent-of-origin expression patterns (e.g., maternal eQTL for MEG3, paternal eQTL for NDN) serves as a powerful internal validation of our phasing methodology, statistical models, and significance thresholds. This demonstrates that our approach has the power to detect true POE signals.

      Conservative Calling Criteria: As the reviewer suggests, we prioritized specificity. Our definition of eQTL sets (Section 4.6) uses stringent thresholds (e.g., log<sub>10</sub> BF > 4 for primary signals and θ = log<sub>10</sub> 2 for exclusivity). We explored different θ parameters (Supplementary Table S2) and chose the one that minimized the inclusion of false positives, ensuring that our core gene sets (e.g., G<sub>1</sub>,G<sub>0</sub>,G<sub>2</sub>) are high-confidence discoveries.

      Rigorous Analytical Pipeline: As we note in the revised text, our conclusions are supported by a robust analytical pipeline. This includes trio-based phasing validated by simulation (Supplementary Table S1), the use of linear mixed models to control for relatedness and population structure, and the application of Bayes factors which inherently penalize variants with low minor allele frequencies, thereby reducing spurious associations.

      We believe these internal consistency checks and methodological rigor provide strong confidence in our findings. To further facilitate external replication, we will make the full list of POE eQTLs and eGenes available as a comprehensive resource (as noted in the Discussion and Supplementary Materials), enabling other researchers to validate these findings as appropriate datasets become available.

      Second, while Bayes factors are thoughtfully used to assess evidence of association, the paper does not fully explore how the chosen thresholds translate to the expected rate of false positives. For example, a minor allele frequency cutoff of 1% was applied, which seems somewhat arbitrary, and without reporting the allele frequency distribution of the identified eQTLs, it is unclear whether rare variants disproportionately contribute to the signals, potentially affecting the reliability of discoveries.

      We thank the reviewer for raising this important point regarding the calibration of our significance thresholds and the potential role of rare variants. We address this by clarifying the relationship between Bayes factors, prior odds, and false discovery rates, and by providing a more detailed characterization of the variants we identified.

      Bayes Factors and False Discovery: The reviewer is correct that the connection between a Bayes factor threshold and a false positive rate is not direct as it has to take into account of prior odds. As we briefly noted, for a given prior odds of association (e.g., 1 in 100 or 1 in 1000 for a cis-eQTL), a log<sub>10</sub> BF = 4 corresponds to a posterior probability of association (PPA) of 0.99 or 0.90 respectively. Consequently, 1 − PPA can be interpreted as the local false discovery rate (lfdr), as we have now explicitly stated in Section 2.2 (citing Soloff et al., 2024). Our choice of log<sub>10</sub> BF = 4 was therefore chosen to ensure a very low or modest lfdr (depending on the prior odds) for our primary findings.

      Minor Allele Frequency Threshold: The 1% MAF cutoff was indeed a pre-analysis filtering step. It was chosen based on the power afforded by our sample size of 1,477 trios. For variants rarer than 1%, our study is underpowered to detect associations, and any signals would be highly unstable. Importantly, the reviewer’s concern about rare variants disproportionately contributing to signals is further mitigated by our use of Bayes factors. As we note in Section 2.2, the prior used in our Bayes factor computation (with σ = 0.5 in the prior for effect sizes, as described in Section 4.4) inherently penalizes variants with small minor allele frequencies. This is because for a given effect size, the evidence for association is weaker for a rare variant than a common one. Thus, the combination of a pre-analysis MAF filter and the Bayesian analysis itself guards against spurious findings driven by very rare alleles.

      Allele Frequency Distribution: To directly address the reviewer’s request for transparency, in the revised manuscript we include a supplementary figure (e.g., Supplementary Figure S4) showing the distribution of minor allele frequencies (1000 genomes European descents) for the SNPs identified in paternal eQTL set S<sub>P</sub> and maternal eQTL set S<sub>M</sub>. This empirically demonstrate that our findings are not disproportionately driven by low-frequency variants and provide a more complete picture of the genetic architecture underlying these POE signals. We also add a sentence to the Results section (Section 2.5) summarizing this distribution.

      Third, the ancestry background of the study samples is not reported, which could be a confounding factor in the genetic analyses.

      We thank the reviewer for highlighting this omission. In the revised manuscript, we explicitly report the ancestry background of the Framingham Heart Study participants analyzed. Consistent with previous reports on this cohort, the vast majority of samples are of European descent.

      Crucially, as the reviewer suggests, population stratification can be a confounder in genetic studies. To mitigate this, our analysis employed a linear mixed model (Section 4.4) that includes a random effect with a covariance structure defined by the genetic relatedness matrix (GRM). This approach is specifically designed to control for spurious associations due to both subtle population structure and known relatedness among individuals, ensuring that our findings are robust to these potential confounders.

      Reviewer 2 (Public review):

      Summary:

      The authors have used 1477 sequenced trios with available gene expression data in the offspring to discover eQTLs that act in a parent-of-origin specific manner. The classified associated SNPs are tested for enrichment for GWAS hits, drug target genes, etc.

      Strengths:

      The manuscript presents an impressive analysis of a very rich data set of parent-of-origin eQTLs. To my knowledge, it is one of the largest studies of its kind, most analyses are sound, and the results are of interest to many in the field and potentially beyond. The different ideas of follow-up analyses are useful and make sense.

      Weaknesses:

      While in general the analyses are well-conducted, I noticed a major issue with the POE eQTL classification, which puts into question most of the downstream analysis. In light of this problem, most of the analysis would need to be rerun, which represents a major revision of the paper, but is straightforward to repair.

      We appreciate the reviewer’s concern and take it seriously. However, we believe the issue stems from a misunderstanding of our classification framework. We clarify our reasoning below, and we are confident that no re-analysis is necessary. In fact, our Bayesian approach was specifically chosen to avoid the very problem the reviewer raises.

      The major problem with the classification of POEs is that simply having significant maternal, but insignificant paternal effect is not an indicator of POE, this happens widely for SNPs with no POE whatsoever (it can happen by chance even when both maternal and paternal effects are the same and non-zero - the authors can see it via simulations under the null [maternal=paternal effect]).

      The reviewer raises a valid statistical concern: under the null hypothesis of equal maternal and paternal effects (β<sub>0</sub> = β<sub>1</sub>≠ 0), sampling variation could occasionally produce a scenario where one effect appears significant and the other does not. This is indeed a form of Type II error (failing to detect a true non-zero effect for one of the alleles).

      However, this is precisely why we chose Bayes factors over p-values. A key advantage of Bayes factors is that they are not blind to power. P-values are calculated solely under the null hypothesis and do not incorporate any information about the alternative hypothesis or the study’s power to detect it. Consequently, when power is low (e.g., due to minor allele frequency differences between paternal and maternal alleles), p-values can be misleading.

      In contrast, Bayes factors are computed under both the null and alternative hypotheses. They inherently incorporate power through the prior specification. As we note in Section 2.2, “Bayes factors penalize genetic variants with small allele frequencies to reduce false positives.” This means that a SNP where, by chance, one allele appears significant and the other does not—but where power is low due to allele frequency imbalance—will not receive a high Bayes factor, because the evidence is appropriately discounted.

      In order to be able to talk about POE, first, a significant difference between maternal and paternal effects needs to be claimed. Therefore, none of the 4 sets of POE eQTLs are justified. To me, the only relevant criterion to pick POE SNPs is the P-value when comparing the maternal and paternal effects.

      We respectfully disagree with the reviewer’s assertion that our approach to POE eQTL classification are not justified. There are multiple biologically meaningful patterns of parent-of-origin effects, and our classification scheme was designed to capture this diversity:

      (1) Paternal-specific eQTL (β<sub>0</sub> = 0, β<sub>1</sub> ≠ 0)

      (2) Maternal-specific eQTL (β<sub>0</sub> ≠ 0, β<sub>1</sub> = 0)

      (3) Opposing eQTL (β<sub>0</sub> ≠ 0, β<sub>1</sub> ≠ 0,β<sub>0</sub> × β<sub>1</sub> < 0)

      (4) Genotype eQTL (β<sub>0</sub>= β<sub>1</sub> ≠ 0)

      The reviewer’s proposed test (H<sub>0</sub>: β<sub>0</sub> = β<sub>1</sub>) collapses these distinct biological scenarios into a single binary outcome. For example: A purely paternal-specific eQTL (β<sub>0</sub> = 0, β<sub>1</sub> ≠ 0) would indeed show a significant difference, and would be captured by the reviewer’s test. However, a gene like ZNF890P in Table 1, where both effects are significant and in the same direction but of different magnitudes, would also show a significant difference. In the reviewer’s framework, this would be classified as a POE eQTL, yet biologically it behaves more like a genotype eQTL with an allelic imbalance. Our framework correctly separates these cases.

      Moreover, the reviewer’s proposed test is a nested special case of our broader approach. As we note in our response, our paternal-specific test (H<sup>0</sup>: β<sub>0</sub> = β<sub>1</sub> = 0 vs H<sub>1</sub>: β<sub>0</sub> = 0,β<sub>1</sub> ≠ 0) is a more constrained hypothesis that yields a subset of the SNPs that would be identified by the reviewer’s difference test, were it to have sufficient power. Our approach is therefore more conservative for classifying paternal- or maternal-specific eQTLs, not less.

      The definitions of the 4 groups are based on somewhat ad hoc priors, BF thresholds, etc. Also, in Section 4.6, the value of theta is arbitrarily chosen (along with the threshold of 4 to declare POE). In my opinion, the clean treatment of the 4 groups would start with a significant P-value (beta-maternal vs beta-paternal). Within this set, you can then use the original criteria presented in the paper, but only among these associations where there is solid evidence of different parental effects.

      We take strong issue with the characterization of our prior specifications and thresholds as “ad hoc” or “arbitrary.” In Bayesian analysis, prior specification is a principled and transparent modeling choice, not an arbitrary one.

      (1) Choice of log<sub>10</sub> BF = 4 threshold: As stated in Section 2.2, this threshold was chosen based on explicit considerations of prior odds and posterior probability of association. For a prior odds of 1:1000 (a reasonable guess for cis-eQTLs), this BF corresponds to a posterior probability of association of 0.91. If one prefers a more optimistic prior odds of 1:100, the PPA becomes 0.99. The threshold is therefore grounded in decision theory, not whim.

      (2) Choice of θ in Section 4.6: We explicitly state that we explored multiple values of θ(0, log<sub>10</sub> 2, log<sub>10</sub> 3) and chose θ = log<sub>10</sub> 2 because it “produced minimum G<sub>1</sub> and G<sub>0</sub> that contain known imprinted genes.” This is a principled, data-driven calibration step using positive controls, not an arbitrary selection. The transparency of this process is a strength, not a weakness.

      (3) Comparison to p-value thresholds: The reviewer suggests that p-value thresholds are somehow less arbitrary. However, the conventional p-value threshold of 0.05 is itself a historical convention with no universal justification. Moreover, as we note, p-values do not account for power differences across SNPs. A p-value of 5 × 10<sup>−8</sup> from a SNP with 40% MAF is not comparable to the same p-value from a SNP with 1% MAF, because the power to detect the association differs dramatically. Bayes factors automatically adjust for this through the prior, making them more comparable across variants, not less.

      In revision, we added a section in supplementary to review relationships between p-values, Bayes factors, and FDR.

      Recommendations for the authors:

      Reviewer 1 (Recommendations for the authors):

      Here are some suggestions to improve the study:

      (1) Provide information about the ancestry background of participants and consider including ancestry principal components in the eQTL models, as is commonly done, to account for population structure.

      We thank the reviewer for this suggestion. In the revised manuscript, we explicitly state that the participants in the Framingham Heart Study are predominantly of European descent, consistent with previous publications from this cohort. Regarding population structure, we respectfully note that our analysis already employs a linear mixed model (Section 4.4) that includes a random effect with a covariance structure defined by the genetic relatedness matrix (GRM). This approach is widely regarded as more robust than including a limited number of principal components, as it accounts for both fine-scale population stratification and known relatedness simultaneously.

      (2) Conduct sensitivity analyses using different Bayes factor cutoffs to assess the robustness of the findings.

      We appreciate the reviewer’s concern about threshold robustness. In fact, we already conducted a form of sensitivity analysis during the classification step. As described in Section 4.6 and shown in Supplementary Table S2, we explored multiple values of θ (0, log<sub>10</sub> 2, and log<sub>10</sub> 3) and observed how they affected the composition of our gene sets. The choice of log<sub>10</sub> BF = 4 for significance was similarly grounded in posterior probability calculations (Section 2.2). To further address the reviewer’s point, we add a Supplementary Table S3 for counts of eQTL and eGenes under different Bayes factor threshold. This demonstrates that our most significant claim, the abundance of POE eQTL, are not overly sensitive to the specific cutoff.

      (3) In the GWAS examples for KCNQ1 and CDKN1C, the assessment of whether the SNPs act as eQTLs for the two genes is based on a single BF threshold, which may be influenced by differences in gene expression levels. The authors could compare the corresponding effect sizes of these SNPs on both genes to provide a more nuanced investigation. While the limitation of missing data from other tissues is discussed in the paper, it remains possible that KCNQ1 plays a role in tissues more relevant to T2D.

      This is an excellent suggestion for a more nuanced investigation. We re-examined the effect sizes for the SNP rs2237892 in our published results. For gene CDKN1C, the paternal log<sub>10</sub> BF<sub>1</sub> = −0.477 and maternal log<sub>10</sub> BF<sub>0</sub> = 4.94, the normalized maternal effect in joint analysis is −4.86 vs −0.74 for paternal. Unfortunately, the published results has no eQTL for KCNQ1, which according to our selection creteria means maximum log<sub>10</sub> BF < 3 for all tests (genotype, paternal , maternal, joint). The concern for different gene expression level may affect BF is valid. We preempt this pitfall by quantile normalization of gene expression levels after controlling for GC content (as documented in Method Section). We agree with the reviewer that the lack of data from pancreatic tissues is a limitation. We add a sentence in revelant section to acknowledging that while whole blood is a valuable and accessible tissue, replication in T2D-relevant tissues (e.g., pancreas, adipose) would be an important future direction, and our findings provide a hypothesis for such targeted investigations.

      Reviewer 2 (Recommendations for the authors):

      Major comments:

      There are some literature elements missing:

      (1) Hofmeister has a newer and larger study [https://pubmed.ncbi.nlm.nih.gov/40770099/].Please cite that too; it also has POE pQTLs, which is relevant.

      (2) POE in pigs has been explored [https://www.nature.com/articles/s41467-02562243-6], please cite it.

      (3) An insightful review covering the mechanisms of POE for gene expression (https://www.sciencedirect.com/science/article/pii/S2352154618300482) should be cited.

      (4) Further studies on POE in gene expression in social insects (https://royalsocietypublishing.org and in mice (https://www.biorxiv.org/content/10.1101/2023.08.24.554674v1.full) are also relevant.

      We thank the reviewer for bringing these important references to our attention. We incorporated the suggested citations in the revision to provide a more comprehensive context for our work, including the newer POE pQTL study by Hofmeister et al., the findings in pigs, and the mechanistic review.

      While it’s OK to report and rank SNPs by BF, it is necessary to show association P-values as well. It is not explained in the text around the Table how the P-value is obtained in the Table. And it is important to show how their priors translate to FWER control. What is the FWER when picking SNPs at a certain BF value? 1-PPA and local FDR depend on the choice of the prior, but we need a prior-independent measure of FDR/FWER.

      We appreciate the opportunity to clarify. The p-value presented in Table 1 (column “P”) is indeed the frequentist p-value testing the null hypothesis of equal maternal and paternal effects (H<sub>0</sub> : β<sub>0</sub> = β<sub>1</sub>), as described in Section 4.5. We included this to provide a familiar metric for readers, but our discovery framework relies on Bayes factors for the reasons outlined in Section 2.2.

      Regarding error control, the reviewer is correct that 1-PPA is a local FDR that depends on the prior. We chose to control the local rate of false discoveries rather than the Family-Wise Error Rate (FWER) because FWER control (e.g., via Bonferroni) is often excessively conservative for exploratory analyses like eQTL mapping, especially given the correlation among tests due to LD.

      Our Bayesian approach provides a more nuanced measure of evidence at the level of each individual test, which is precisely what is needed for prioritizing SNPs with parent-of-origin effects.

      The demand for a prior-independent measure of FDR is conceptually problematic. Any probabilistic statement about a specific hypothesis being true or false necessarily requires a prior—this is a fundamental consequence of probability theory. Frequentist FDR, while prior-independent in one sense, does not provide a probability that a particular finding is false; it is a long-run error rate over many tests. Methods like q-values, often described as “prior-free,” still depend on implicit assumptions (e.g., the estimate of π<sub>0</sub>, independence of tests, and a mixture of effect sizes).

      In our specific context of cis-eQTL analysis, these assumptions are particularly questionable. LD induces correlation among nearby SNPs, violating the independence required for stable π<sub>0</sub> estimation. Moreover, effect sizes in a region are not randomly mixed—SNPs in high LD tend to have similar effect directions and magnitudes, which can bias the mixture model underlying q-value approaches. Our Bayesian approach, by modeling each SNP individually, avoids these cross-SNP assumptions.

      Importantly, while posterior probabilities depend on the choice of prior (π<sub>0</sub>), we have verified that our conclusions are robust across a wide range of plausible π<sub>0</sub> values (0.9,0.99,0.999). Given our extremely stringent Bayes factor threshold (BF<sub>j</sub> > 10<sup>4</sup>), the posterior probability for a maternal effect exceeds 0.90 for any π<sub>0</sub> < 0.999. Thus, the prior dependence is practically irrelevant for the SNPs we report.

      In revision, we added a section in Supplementary to describe the connections between p-value, Bayes factor, and FDR. We hope this will clarify that a (seemingly) prior independent FDR has a hidden assumption that cis-eQTL analysis is likely to violate.

      The major problem with the classification of POEs is that simply having significant maternal, but insignificant paternal effect is not an indicator of POE, this happens widely for SNPs with no POE whatsoever (it can happen by chance even when both maternal and paternal effects are the same and non-zero - the authors can see it via simulations under the null [maternal=paternal effect]). In order to be able to talk about POE, first, a significant difference between maternal and paternal effects needs to be claimed. Therefore, none of the 4 sets of POE eQTLs are justified. To me, the only relevant criterion to pick POE SNPs is the P-value when comparing the maternal and paternal effects. The definitions of the 4 groups are based on somewhat ad hoc priors, BF thresholds, etc. Also, in Section 4.6, the value of theta is arbitrarily chosen (along with the threshold of 4 to declare POE). In my opinion, the clean treatment of the 4 groups would start with a significant P-value (beta-maternal vs beta-paternal). Within this set, you can then use the original criteria presented in the paper, but only among these associations where there is solid evidence of different parental effects.

      We respectfully disagree with the reviewer’s assertion that a significant difference between maternal and paternal effects is the only valid criterion for defining POE, and we maintain that our classification is statistically sound and biologically meaningful.

      The Problem with the “Difference-Only” Approach: The reviewer’s proposed filter (a significant p-value for β<sub>0</sub> ≠ β<sub>1</sub>) is a single hypothesis test. Our goal was to classify eQTLs into multiple, distinct biological categories (paternal-specific, maternal-specific, opposing, etc.). The “difference-only” test collapses these categories. For example, a purely paternal-specific eQTL (β<sub>0</sub> = 0,β<sub>1</sub> ≠ 0) and a gene like ZNF890P (β<sub>0</sub> ≠ 0, β<sub>1</sub> ≠ 0, β<sub>0</sub> > β<sub>1</sub>) would both show a significant difference. In the reviewer’s framework, they would be lumped together, obscuring the fact that one is an imprinted gene and the other is a standard eQTL with allelic imbalance. Our framework correctly separates them.

      Bayes Factors are Not “Ad Hoc”: The choice of prior (σ = 0.5) follows established literature for linear model Bayes factors (Servin and Stephens, 2007). The threshold of log<sub>10</sub> BF = 4 was chosen based on its relationship to posterior probability (0.91-0.99 given reasonable prior odds), which is a transparent and principled decision rule. The selection of θ in Section 4.6 was calibrated using a positive control set of known imprinted genes, ensuring our definitions were conservative and accurate. This is the opposite of arbitrary.

      The Suggested Procedure Has Low Power: One can run the following simple R code to verify. We simulate maternal alleles xx and maternal alleles yy, then simulate phenotype with β<sub>xx</sub> > 0 and β<sub>yy</sub> = 0 (maternal effect only). We fit the joint model and compute p-values for the null β<sub>xx</sub> = β<sub>yy</sub> as suggested by reviewer. From the joint fit, we also extract p-values based on the null β<sub>xx</sub> = 0 and β<sub>yy</sub> = 0 respectively. The simulation was repeated 1000 times and p-values were stored in a matrix.

      We call positives based on suggested procedure, and compare number of positives called using marginal p-values at two threshold of 1×10<sup>−5</sup> and 1×10<sup>−6</sup> to declare significance. We used threshold of 0.01 to declare insignificance.

      The result demonstrates that the suggested procedure has a much lower power compared to the procedure based on marginal statistics.

      For the above reasons, the follow-up enrichment analysis is somewhat questionable. Most enrichments are non-significant, and it is likely because the SP and SM groups are diluted with SG SNPs. The P1-P9 groups have nothing to do with POE, and although the observation of increased enrichment for GWAS SNPs with increased pleiotropy is interesting, it is irrelevant for POE.

      We will address the dilution concern below. We agree that P1-P9 groups are not directly related to POE. But this is an interesting observation non-theless. As we found such an observation is missing in the literature, we ask to keep it in the paper.

      In the same way, section 2.7 is not supported; the claimed maternal and paternal POEs are heavily diluted by simple marginal associations. The same holds for sections 2.82.10. A striking example is Table 3: for clinical trial targets, paternal/maternal eQTLs behave just like simple marginal eQTLs (G<sub>G</sub>). A similar pattern emerges for combined target enrichment.

      The reviewer’s concern that our S<sub>P</sub> and S<sub>M</sub> sets are “diluted with S<sub>G</sub> SNPs” is precisely the issue our Bayes factor thresholds were designed to prevent. By requiring one effect to be significant and the other to be below a low threshold (θ), we explicitly excluded SNPs where both effects are significant and in the same direction (which defines S<sub>G</sub>).

      Regarding Table 3, the reviewer’s interpretation differs from ours. The fact that paternal eQTLs (G</sub>P</sub>) show significant enrichment for drug targets, while genotype eQTLs (G<sub>G</sub>) also show enrichment, does not imply dilution. Rather, it suggests there is an overlap in the biological importance of these gene sets, which is expected. The key message of the finding is the asymmetry: G<sub>P</sub> is significantly more enriched than G<sub>G</sub> (p=0.035 for combined targets), a pattern that would be washed out if G<sub>P</sub> were merely a diluted version of G<sub>G</sub>. This asymmetry supports the interesting biological hypothesis (Moore and Haig, 1991) we discuss. The non-significance for G<sub>M</sub> further highlights this asymmetry.

      I’m not sure how MR would be biased by POE: MR is conducted only if there is a marginal association, i.e., the average maternal and paternal effects are significant. If the expression is causal for a trait, the POE effect is propagated to the outcome; hence, the SNP effect on the exposure will be equally biased as the SNP effect on the outcome, and these cancel out, and the causal effect remains unbiased. Can the authors propose a concrete example of maternal/paternal effects that demonstrates their claimed bias?

      We thank the reviewer for this insightful question, which allows us to clarify our point with a concrete example from our data.

      Consider a scenario where one wishes to use Mendelian Randomization (MR) to test whether the expression of gene NECAB3 causally influences a particular trait (e.g., obesity). The reviewer is correct that if the causal effect is homogeneous, the average effect might still be captured. However, the bias we caution against arises in stratified analyses or in the interpretation of the genetic instrument itself.

      Take the SNP rs4911348 and its effect on NECAB3 (Figure 2). The genotype model shows no marginal association. Therefore, if a researcher were conducting a standard MR study using this SNP as an instrument for NECAB3 expression, they would discard it as an invalid instrument due to the lack of a marginal association. They would miss the true underlying biology entirely. The causal effect of NECAB3 on the trait would be masked in the full population.

      More subtly, even if a SNP has a marginal association, using it as an instrument while ignoring POE can lead to incorrect effect estimates in population subgroups defined by parent of origin. This is analogous to ignoring effect modification. For instance, if a treatment (exposure) has a different effect depending on which parent it came from (which is impossible, but the genetic propensity for the exposure does), failing to account for this can bias the instrumental variable estimate if the instrument’s strength varies by an unmeasured factor (parental origin).

      Our advice to “check the list of POE SNPs” is a practical caution: if the instrument for an exposure exhibits strong POE, the standard MR assumptions about the homogeneity of the instrument’s effect may be violated, potentially leading to biased estimates or incorrect conclusions about causality.

      Minor comments:

      (1) In Table 1, the last column header should be -log10(P), not ”P”.

      The column labelling is an editorial choice to prevent table overflow. This particularly labelling was explained in the caption.

      (2) While BFg/0/1/j are explained in the text, these notations should be explained in the Table caption as well.

      Added explanation in caption.

      (3) It should also be mentioned in the Table 1 caption how these top 10 SNPs were chosen.

      These are sentinel eQTL for each gene. We think the first paragraph of Section 2.3 explains clearly.

      (4) “may ”acquires” a cis-eQTL through” → ”may ”acquire” a cis-eQTL through”.

      Corrected. Thank you.

      (5) “which retained 16, 969 genes out of total 58103”, I assume the 58103 are transcripts, not genes.

      You are absolutely correct. We added transcripts after 58103.

      (6) In Equation (1), Z is not defined. In this concrete setting, isn’t it simply the identity matrix?

      Yes. Z is the identitity (loading) matrix for human study. We added a sentence to clarify in revision.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (on non-trivial pattern transformations):

      (3) All modelling is confined to one spatial dimension, and the very definition of a "non-trivial" transformation is framed in terms of peak positions along a line, which clearly must be reformulated for higher dimensions. It's well-known that diffusions in 1, 2, and 3 dimensions are also dramatically different, so the relevance of the three-class taxonomy to real multicellular tissues remains unclear, or at least should be explained in more detail.

      Reviewer #2 (on non-trivial pattern transformations):

      (5) The definition of non-trivial pattern formation is provided only in the Supplementary Information, despite its central importance for interpreting the main results. It would significantly improve clarity if this definition were included and explained in the main text. Additionally, it remains unclear how the definition is consistently applied across the different initial conditions. In particular, the authors should clarify how slopebased measures are determined for both the random noise and sharp peak/step function initial states. Furthermore, the authors do not specify how the sign function is evaluated at zero. If the standard mathematical definition sgn(0)=0 is used, then even a simple widening of a peak could fulfill the criterion for non-trivial pattern transformation.

      There was indeed a problem on how we defined non-trivial pattern transformations in the original version. This definition was not clear enough beyond 1D. We now provide a simple clear definition in the main text that applies to all dimensions (“P1” and “P2” in the second page of the introduction).

      As we now explain through the main text, even if the solution of the heat/diffusion equation depends on the dimension of the system, our classification of gene networks (and the mathematical analyses we use) does not depend on the dimensionality of the system. However, some aspects of the specific pattern transformations possible from these networks depend on the dimensionality of the system. In the current version of the article, every time we explain something about the resulting patterns in 1D, we also explain it for the resulting patterns in 2D and 3D. We also have added figures for the 2D cases (in current Fig.1 and Fig.9). We now explicitly explain how the possible resulting patterns in space can depend on the boundaries and shapes of the system (i.e. the distribution of cells in space) (see specially the 5th paragraph of the discussion).

      The criticisms about “slope-based measures” mentioned by reviewer 2, is now addressed in a paragraph at the end of the introduction (here we added it):

      “It is worth noting that these three basic initial patterns correspond to spatially discontinuous functions: in homogeneous with noise initial patterns, white noise is discontinuous by definition; in spike and combined spike-homogeneous initial patterns, there is a concentration discontinuity between cells on the edge of the spike and nearby cells outside the spike. However, once extracellular signal diffusion begins, these sharp boundaries are smoothed into differentiable gradients, where critical points can be properly defined (e.g., at the center of the initial spike).”

      The main concern among these relates to the validity of our linearization of the model equations and the extension of the results obtained for the linear system to the fully nonlinear system. In this regard, the reviewers’ comments are:

      Reviewer #1 (on linearization):

      (2) A central step in the model formulation is the linearisation of the reaction term around a homogeneous steady state; higher-order kinetics, including ubiquitous bimolecular sinks such as A + B → AB, are simply collapsed into the Jacobian without any stated amplitude bound on the perturbations. Because the manuscript never analyses how far this assumption can be relaxed, the robustness of the three-class taxonomy under realistic nonlinear reactions or large spike amplitudes remains uncertain.

      Reviewer #2 (on linearization):

      (2) Most of the proofs presented in the Supplementary Information rely on linearized versions of the governing equations, and it remains unclear how these results extend to the fully nonlinear system. We are concerned that the generality of the conclusions drawn from the linear analysis may be overstated in the main text. For example, in Section S3, the authors introduce the concept of dynamic equivalence of transitive chains (Proposition S3.1) and intracellular transitive M-branching (Proposition S3.2), which pertains to the system's steady-state behavior. However, the proof is based solely on the linearized equations, without additional justification for why the result should hold in the presence of nonlinearities. Moreover, the linearized system is used to analyze the response to a "spike initial pattern of arbitrary height C" (SI Chapter S5.1), yet it is not clear how conclusions derived from the linear regime can be valid for large perturbations, where nonlinear effects are expected to play a significant role. We encourage the authors to clarify the assumptions under which the linearized analysis remains valid and to discuss the potential limitations of applying these results to the nonlinear regime.

      We used three linearizations in the original version of the manuscript. One was to analyze hierarchic networks (in the Hierarchic networks section). In the new version of the article we do not use any linearization to study the hierarchic networks, so this problem is solved.

      The second linearization was in section S3 on transitive chains. We realized that this section is not really necessary at all for the article so we deleted it.

      We keep the third linearization but we now explain why such linearization is useful and valid in a section called “Linear stability analysis”. Thus, through this section we justify this choice (explicitly in its two first paragraphs).

      Regarding Reviewer 2 concerns about large perturbations, we acknowledge that the phrasing using “arbitrary height” may have been confusing. As we now explain in the linear stability analysis section, linear stability analysis assumes perturbations to be small.

      For the homogeneous-with-noise initial pattern, as we explain, these perturbations are assumed to be small because they are actually molecular noise.

      For the spike initial pattern and hierarchic networks the perturbation is not necessarily small. However, by the definition of the spike and combined homogeneous-spike initial patterns, all cells outside the spike start with the same concentration of the extracellular signals that are secreted from the spike (e.g. zero). Thus, even in the case in which extracellular signals concentrations in the spike would be unrealistically high, the amount of extracellular signal diffusing from it can be considered small by simply considering it at a small enough time interval. Thus, right outside the spike the diffusion of extracellular signals from the spike can be treated as a continuous small perturbation for which one can study the stability, as we do in the “Linear stability analysis section”. This we now explain at the end of the introduction and in the “Linear stability analysis” section when we talk about the initial patterns again.

      In the following, we respond to the remaining concerns raised by the reviewers:

      Reviewer #1 (Public review):

      (1) The Results section is difficult to follow. Key logical steps and network configurations are described shortly in prose, which constantly require the reader to address either SI or other parts of the text (see numerous links on the requirements R1-R5 listed at the beginning of the paper) to gain minimal understanding. As a result, a scientifically literate but non-specialist reader may struggle to grasp the argument with a reasonable time invested.

      We acknowledge that the original version of the main text may not be as clear as we intended. Initially, we believed that placing the more technical mathematical passages in the Supplementary Information would make the main text more accessible to readers. We were wrong. We have now moved crucial parts of the supplementary to the main text and adapted the rest of the text accordingly. The most important of those is the new “Linear stability analysis” section and the associated dispersion relation (e.g. Fig.6).

      Reviewer #2 (Public review):

      (1) We have serious concerns regarding the validity of the simulation results presented in the manuscript. Rather than simulating the full nonlinear system described by Equation (1), the authors base their results on a truncated expansion (Equation S.8.2) that captures only the time evolution of small deviations around a spatially homogeneous steady state. However, it remains unclear how this reduced system is derived from the full equations -specifically, which terms are retained or neglected and why- and how the expansion of the nonlinear function can be steady-state independent, as claimed. Additionally, in simulations involving the spike plus homogeneous initial condition, it is not evident -or, where equations are provided, it is not correct- that the assumed global homogeneous background actually corresponds to a steady state of the full dynamics. We elaborate on these concerns in the following:

      We are actually simulating the full nonlinear system described by Equation (1). In the current version we are more explicit about this. As we describe in the introduction and, now, through all the text several times (e.g. in the last paragraph of the model section and in the paragraph before the linear stability section), the aim of the article is to describe necessary requirements for non-trivial pattern transformations. We did not intent to describe all necessary requirements nor sufficient requirements. These requirements are at the level of gene network topology not at the level of f or its parameters. In other words, we just claim that gene networks having specific topological features can lead to some specific types of non-trivial pattern transformations but not to others. We do not say for which specific fs (or its parameters) these pattern transformations are possible, we just say that this can happen for some f, as long as these fulfill our requirements. We do show, however, that without some specific topological requirements there are non-trivial pattern transformations that are not possible, no matter the f (this explicitly stated in the last paragraph of the model section and in the paragraph before the linear stability section). Thus, all the simulations shown in the figures are just examples, with specific fs, of the types of non-trivial pattern transformations possible from each type of gene network topology.

      In all simulations we used the f of the Maini-Miura model. We could have chosen other ones but we happen to chose that f. The presentation of the Maini-Miura model has been revised to improve clarity (equation S6.1 in SI). This model we are simulating fully, we are not doing any linearization for the simulations. That may not have been explained clearly enough in the previous version of the article. We just happen to make a change of variable that may have been confused as a linearization. In the current version, the existence of a homogeneous steady state is parameterized by a tunable g<sup>*</sup>, that can be chosen as for spike initial patterns or g for noise-homogeneous and spike-homogeneous initial patterns. We have also included a proof that the model equations satisfy our conditions R1-5. Indeed, the model is non-linear as long as σ<sub>i</sub>≠0 for some gene product (as we explicitly assume).

      It is assumed that the homogeneous steady states are given by g_i=0 and g_i=c_i, where 1/c_i = \mu_i or \hat{\mu}_i, independently of the specific network structure. However, the basis for this assumption is unclear, especially since some of the functions do not satisfy this condition -for example, f5 as defined below Eq. S8.10.5. Moreover, if g_i=c_i does not correspond to a true steady state, then the time evolution of deviations from this state is not correctly described by Eq. S8.2, as the zeroth-order terms do not vanish in that case.

      In the revised manuscript, homogeneous steady states are parameterized by a tunable g<sup>*</sup>, which can be chosen as for spike initial patterns or g for noise-homogeneous and spike-homogeneous initial pattern. Function f(g) in (S6.1), as well as the specific non-linear entries used in certain simulations, are constructed such that g<sup>*</sup> is indeed a steady state of the system and that conditions R1-R5 are satisfied. We have also corrected some typos in section S6 (previously section S8) of the Supplementary Information, that we believe may have induced the confusion indicated by this reviewer.

      Additionally, the equations used contain only linear terms and a cubic degradation term for each species g_i, while neglecting all quadratic terms and cubic terms involving cross-species interactions (i≠j). An explanation for this selective truncation is not provided, and without knowledge of the full equation (f), it is impossible to assess whether this expansion is mathematically justified. If, as suggested in the Supplementary Information, the linear and cubic terms are derived from f, then at the very least, the Jacobian matrix should depend on the background steady-state concentration. However, the equations for the small deviation around a steady state (including the Jacobian matrix) used in the simulations appear to be independent of the particular steady state concentration.

      As described above we just chose an example f to exemplify the non-trivial pattern transformations possible from each class of gene network topologies. There is no special reason to include, or exclude for that matter, cubic cross-species interactions since the point is just to exemplify the types of possible pattern transformations from each type of gene network topology.

      In addition, we believe that part of the reviewer’s concern may have arisen from a notational ambiguity in the previous version of the manuscript, which has now been corrected: the matrix appearing in f(g) has been renamed from J to W<sup>T</sup>. As stated in the main text, the jacobian of the regulation function f(g) evaluated at the homogeneous steady state must coincide with the transpose of the network weight matrix. With the current equations (S6.1), we have , from which we easily get . Also, it is clear that the Jacobian of f(g) is not independent of g.

      This is why we believe that the differences observed between the spike-only initial condition and the spike superimposed on a homogeneous background are not due to the initial conditions themselves, but rather result from a modified reaction scheme introduced through a questionable cutoff.

      "In simulations with spike initial patterns, the reference value g≡0 represents an actual concentration of 0 and therefore, we must add to (S8.2) a Heaviside function Φ acting of f (i.e., Φ(f(g))=f(g) if f(g)>0 , Φ(f(g))=0 if f(g){less than or equal to}0) to prevent the existence of negative concentrations for any gene product (i.e., g_i<0 for some i)." (SI chapter S8).

      This cutoff alters the dynamics (no inhibition) and introduces a different reaction scheme between the two simulations. The need for this correction may itself reflect either a problem in the original equations (which should fulfill the necessary conditions and prevent negative concentrations (R4 in main text)) or the inappropriateness of using an expanded approximation which assumes independence on the steady state concentration. It is already questionable if the linearized equations with a cubic degradation term are valid for the spike initial conditions (with different background concentration values), as the amplitude of this perturbation seems rather large.

      The Heaviside function does not preclude inhibition, it precludes gene product concentration to be negative. In the current version of the article we do not use the Heaviside function but another similar, but continuous, function. Having this function can indeed affect the dynamics but: 1) does not violate our requirements on f 2) Does not affect which non-trivial pattern transformations are possible from which gene network topology. Without this function non-trivial pattern transformations are still possible from the spike initial pattern through hierarchical networks, in the way we describe in the article. The Heaviside function (and the one we now use) simply allows that to happen more easily, i.e. for a larger range of parameter values. With this function large inhibitions do not lead to negative gene products concentrations while without it, this can happen for some parameter combinations. None of the arguments nor proves in our article requires the Heaviside, or any similar function. Again this is simply because our aim is to identify topological requirements that are necessary, but not sufficient, for non-trivial pattern transformation. So an f that leads to negative gene products concentrations for some parameter combinations but to non-trivial pattern transformations for others, is still valid example of our points (although not the most interesting or realistic example f).

      We distinguish between the spike and combined spike-homogeneous initial patterns simply because they are biologically quite different, i.e. in the former the gene product in the spike is only expressed in the spike and nowhere else. As we describe in the current version the pattern transformations possible from these two different initial patterns are very similar. In the same way, which gene network topologies can lead to which types of non-trivial pattern transformations is not affected by using the Heaviside functions or not (although this can affect the range of parameter values in which this happens).

      Lastly, we note that under the current simulation scheme, it is not possible to meaningfully assess criteria RH2a and RH2b, as they rely on nonlinear interactions that are absent from the implemented dynamics.

      The implementation of nonlinear entries in f(g) whenever they are needed is now made explicit in the corresponding subsection in the main text and in section S6 in the Supplementary Information. This entries also satisfy conditions R1-R5 around the steady state given by g<sup>*</sup>. Again we should insist that the simulated fs are nonlinear (as now explicitly explained in the SI).

      (3) Several statements in the main text are presented without accompanying proof or sufficient explanation, which makes it difficult to assess their validity. In some cases, the lack of justification raises serious doubts about whether the claims are generally true. Examples are:

      "For the purpose of clarity we will explain our results as if these cells have a simple arrangement in space (e.g., a 1D line or a 2D square lattice) but, as we will discuss, our results shall apply with the same logic to any distribution of cells in space." (Main text l.145-l.148).

      The result of which gene network topologies can lead to pattern transformations are based on a linear stability analysis and some logical arguments. As we now explain through the text none of them depends on the number of dimensions nor on the shape of the arrangement of cells. The geometry of the domain can influence the specific form of the resulting patterns, but it does not alter the broader type of resulting patterns (e.g., periodic patterns, peaks emerging around a spike, etc.) that a given gene network topology can produce. We now explicitly discuss these dependencies in the 5th paragraph of the discussion.

      "For any non-trivial pattern transformation (as long as it is symmetric around the initial spike), there exists an H gene network capable of producing it from a spike initial pattern." (Main text l.366f).

      We now provide a more detailed justification of this statement and the limits of its applicability. This is now in section: “The ensemble of possible pattern transformations from spike initial patterns in H networks“. To make this section easier to understand, however, we have also done changes through all the hierarchic networks sections.

      "In 2D there are no peaks but concentric rings of high gene product concentration centered around the spike, while in 3D there are concentric spherical shells." (Main text l. 447ff).

      This result pertains specifically to pattern transformations arising from spike initial patterns. As defined in the text, spike initial patterns are radially symmetric (at least far away from the boundary). Since diffusion preserves radial symmetry, pattern transformations from spike initial patterns in two or three dimensions reduce to effectively one-dimensional transformations along each radial direction. In this framework, each pair of concentration peaks symmetric with respect to the spike in one dimension corresponds to a ridge surrounding the spike in two dimensions, and each ridge in two dimensions becomes a spherical ridge shell around the spike in three dimensions. In the current version we explain what happens in 1D but also, in the same places, what happens in 2D and 3D (and we have added figures to visualize this in 2D, e.g. Fig.1 and Fig.9)).

      (4) The study identifies one-signal networks and examines how combinations of these structures can give rise to minimal pattern-forming subnetworks. However, the analysis of the combinations of these minimal pattern-forming subnetworks remains relatively brief, and the manuscript does not explore how the results might change if the subnetworks were combined in upstream and downstream configurations. In our view, it is not evident that all possible gene regulatory networks can be fully characterized by these categories, nor that the resulting patterns can be reliably predicted. Rather, the approach appears more suited to identifying which known subnetworks are present within a larger network, without necessarily capturing the full dynamics of more complex configurations.

      We acknowledge that our explanation regarding the combination of sub-networks may have been too brief. We now provide a more detailed description in the section “Gene networks combining different classes of subnetworks” and in its sub-sections. There we explore the different ways in which signal subnetworks can be combined (upstream, downstream, in series, in parallel, etc.). However, this section cannot be understood (and that may have been the problem in the original version of the manuscript) without the linear stability analysis section that is now in the main text, and the associated discussion on the dispersion relation and results related to it. These are important because they apply to all gene networks and, thus, constrain the possible gene network topologies and the types of possible pattern transformations. In other words, whichever ways gene networks are combined, they will always be RD-stable (i.e. no pattern transformation) or RD-unstable of the first (periodic resulting patterns) or second kind (other patterns we discuss). In the current version, we combine this fact with other arguments to describe the types of pattern transformations possible by gene networks combining the different classes of subnetworks.

      (6) The manuscript lacks a clear and detailed explanation of the underlying model and its assumptions. In particular, it is not well-defined what constitutes a "cell" in the context of the model, nor is it justified why spatial features of cells -such as their size or boundaries- can be neglected. Furthermore, the concept of the extracellular space in the one-dimensional model remains ambiguous, making it unclear which gene products are assumed to diffuse.

      We now clarify all these points in the first three paragraphs of the “Methods: the Model” section. We have also included a figure for that clarification (Fig.3).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I suggest the following changes for each weakness I mentioned in the Public Review:

      (1) Presentation

      (R1.1) (a) Add a one-page "Key Requirements" table (e.g., immediately after the Model section) that lists every requirement code (R1-R5, I1-I2, RH1-RH2, etc.), its one-line statement, and the SI section where it is proved.

      In the new version of the article each requirement has its own paragraph starting with the requirement label, e.g. R1 (in bold): ….. We introduce each requirement there where they are justified or proven, otherwise the reader may not know where do they come from. We have also hyperlinked all requirements and most equations so that the reader can easily go back to the explanation of each requirement and equation.

      (R.1.2) Provide more figures illustrating the general structure of networks when you describe them; the network sketches could be folded into a single summary figure, so the reader sees all motifs at once. For example, in lines 304-311, it took me a while to understand if the requirement means just A -> k - ... ⊣ j, or it additionally requires A->...->j (through another pathway). It seems that the full requirement is A → k ⊣ j together with an independent positive route A → j. A figure describing the network structure, or at least a schematic "inline" plot in the spirit of what I just wrote, could help. This is just one example, but the text consists of a constant flow of such "diagrams encrypted in prose".

      We have followed the reviewer’s suggestions. Not all fit in a single figure so we have constructed new figures 4 and 5 for that purpose.

      (R.1.3) (b) Also consider supporting the main text with some key formulas and arguments from SI. My overall suggestion here is that it would be great to make the main text less prosaic and more self-consistent, if the journal requirements allow it.

      After the suggestions by both reviewers, and for the sake of clarity, we have actually moved (and clarified) several key parts of the SI into the main text. These include the whole “Linear stability analysis” and “Positive regulatory loops determine the kind of RD-instability” sections. These parts, although quite mathematical, facilitate the understanding of our results.

      (2) Linearisation

      (R.1.5) It's clear that keeping non-linearity is complicated and maybe redundant, but please, discuss the assumption of linearity explicitly, especially in the scope of relevance for the real systems, and explain why it's not important, if so. I guess that relaxing this assumption may affect the argumentation in many places, for example, equation (3) of the main text could break (i.e., if the signaling molecule can be consumed in some reaction of A+B->AB kind).

      We agree that the original version was not explicit enough about the reasons for the linear approximation. The first and last paragraphs of the section “Linear stability analysis” are explicitly devoted to justify this linearization. Moreover, the hierarchical network section is now written without using the linearization.

      We are not sure we understand which is the problem with the A+B→AB reaction. We are not assuming any specific f function, just the ensemble of functions that fulfill our requirements (R1 to R5). It is only for the simulations that we have to use a specific f. The reactions suggested by the reviewer could represent an f of the form d[AB]/dt=fAB([A]*[B])-m*[AB]**n for AB and d[A]/dt=-fAB([AB]) and d[B]/dt=-fAB([AB]), where fA and fB are functions that decrease with their arguments. We see no reason why there cannot be a fAB that fulfills our requirements. For example fAB=[A]*[B]/(K+[A]*[B])-m*[AB]. See also related comments in the public comments file.

      (R.1.6) Please, provide a separate section where you reformulate the definition of "non-trivial pattern transformation" for two- and three-dimensional domains, and summarize in this section why the analysis provided for 1D is relevant for higher-dimensional systems. By now, I'm not convinced.

      There was indeed a problem with the way we described non-triviality beyond 1D in the original version of the article. We have now refined the definition of pattern transformations so that it is understandable in 2D and 3D. This definition is presented in the introduction already (in P1 and P2). We have modified figure 1 accordingly.

      Reviewer #2 (Recommendations for the authors):

      Major Issues

      (1) Mathematical Proofs

      (R2.1) We strongly recommend that the authors revisit the mathematical derivations or provide a clear and rigorous justification for the assumptions made therein. These assumptions currently appear unjustified or overly simplistic, especially in light of the nonlinear dynamics the authors aim to describe. The authors should comment on why they expect their results to generalize to all complex network structures, as claimed, and not only apply to the simplified examples analyzed in the paper.

      The article has now been restructured to that end. Concerning the assumptions, they are now all explicitly described in the “Methods: the model” section. Concerning the derivations they are through all the results section. A major change in this line has been the moving of part of the supplementary into specific sections in the main text (and the consequent adaptation of the rest of the text). There are important points of the derivation that may have been buried into the old supplementary and that are crucial to understand the whole argument in the article. In fact, a large part of the results section is just a long argument to show that there are essentially only three classes of gene network topologies that can lead to non-trivial pattern transformations. These arguments are summed up in the last paragraph of the new section “Positive regulatory loops determine the kind of RD-instability” and in the first paragraph of the discussion. In brief:

      (1) Pattern transformation requires gene networks with extracellular signals

      (2) Applying previous mathematical results we show (given the broad requirements on f we have) that pattern transformation is only possible in gene networks that contain positive regulatory loops.

      (3) Applying previous mathematical results we show that in the gene networks in which these loops are extracellular, the only possible non-trivial pattern transformations lead to periodic resulting patterns.

      (4) Applying previous mathematical results we show that in the gene networks in which these loops are INTRAcellular, the only possible non-trivial pattern transformations do not necessarily lead to periodic resulting patterns.

      (5) Using simple logical arguments we also show that no non-trivial pattern transformations are possible in gene networks without negative interactions.

      (6) All the above points combined shows that there are only three classes of gene networks capable of nontrivial pattern transformations. 1) Those with intracellular positive loops, extracellular signals that do not affect themselves and some negative regulation by those (that we call hierarchic networks) 2) Those with intracellular positive loops and extracellular signals that affect themselves negatively (that we now call over-Turing networks) 3) Those with extracellular positive loops and an extracellular negative loops (that following previous work by others are called Turing networks).

      (7) Following previous research and different developmental arguments we explore the types of patterns transformations each of these three classes of gene networks can lead to. These types are characterized only in broad and potential terms. We say nothing about the parameters values for which any gene network leads to any specific pattern transformation. What we say is which types of pattern transformation may be possible (for some possible parameter combination) and which ones are not possible from gene network topology alone (based on the types of loops and so on).

      (R.2.3) Additional to the examples provided in the Public Review, claims such as "despite the large amount of theoretically possible gene network topologies, all gene network topologies necessary for pattern formation fall into just three fundamental classes and their combinations" (l. 34ff)

      This statement was originally intended as an introduction of the text following after it but it seems now clear that this was not apparent enough. This statement has been deleted but we convey a similar message letter in the text, now once its justification is provided. In fact, the justification for this statement is the summary we just described in the previous point (R.2.2) and it is discussed over the main text and summarized in the last paragraph of section “Positive regulatory loops determine the kind of RD instability”.

      (R.2.4) and "The same applies to the topologies we found not to be able to lead to non-trivial pattern transformation" (S7) are not or inadequately justified and should be either substantiated or significantly toned down.

      The same comments that above apply.

      (R.2.5) (a) We advise the authors to argue why it is enough to prove key results by considering linear dynamics (see S2-S7). While linearization is a common technique, the authors themselves emphasize the importance of nonlinearities in pattern formation throughout the paper.

      In the current version we provide an explicit justification for this in the section “Linear stability analysis”, especially in its first paragraph. Moreover, for the analysis of the hierarchical networks we do longer use any linearization.

      (R.2.6) (b) To make linear analysis meaningful, we suggest restricting the initial conditions to small fluctuations (e.g., small spikes or noise), which would justify using linearization to investigate the onset of non-trivial pattern formation. Alternatively, the authors should attempt to generalize the results to fully nonlinear dynamics, ideally for a broader class of functions f.

      As we now explain, the homogeneous-with-noise initial pattern already correspond to small perturbations around the homogeneous steady state (due to molecular noise). In addition, for the spike and spike–homogeneous initial pattern we now explicitly consider spikes of small amplitude. We acknowledge that the use of larger spikes in the previous version could lead to misunderstandings regarding the validity of the linear approximation, even though it does not contradict the assumptions underlying the analysis. In these initial patterns, pattern formation arises because the signal secreted from the spike diffuses into the surrounding domain, so that cells outside the spike experience only small deviations from the equilibrium concentration.

      Larger spikes may induce stronger deviations in cells located very close to the spike; however, because the spike occupies a region that is very small relative to the total domain size, these local effects do not influence pattern formation in the bulk of the domain. A similar situation occurs with boundary effects in cells located near the domain limits, which likewise do not affect the pattern formation process away from the boundaries. We have clarified this point in the revised manuscript, both in the final sentences of the Introduction and in the description of the initial conditions in the fourth paragraph of the “Linear stability analysis” section, where we explicitly state that each initial pattern can be interpreted as a perturbation of an otherwise homogeneous pattern.

      (R.2.7) (c) The assumptions required for the proofs should be explicitly stated and justified. At present, the logic behind the chosen constraints on f is unclear, and the flow of the argument suffers as a result.

      The actual justification for the requirements (i.e. constraints) on f are biological (and we now explain them more explicitly when we introduce these requirements). Most of the mathematical proofs do not require these requirements except when we explicitly say so.

      (R.2.8) (d) The illustrative functions provided in some of the proofs in the SI (e.g. S5.2.1 "To see this, let us consider, for example, that they are both quadratic monomials of the form f_k(g_A)=B_k g_A^2 and f_j(g_A)=B_j g_A^2") do not satisfy the authors' own stated conditions (e.g., this function violates requirement R4 (l.197 f)). More suitable examples should be selected to ensure consistency between assumptions and illustrations.

      We have changed the whole section (based on the comment R.2.9 from the same reviewer). We now provide arguments in the main text that generally do not rely on specific fs.

      (R.2.9) (e) Currently, all mathematical results are confined to the appendix. We recommend including key insights from the proofs in the main text to improve readability and to allow the main claims to stand on their own. For example, the section on the requirements RH2a and RH2b (l. 320 - l. 335)) would benefit strongly from the insights from S5.2.1

      We agree. We have moved the linear stability analysis and the dispersion relation section to the main text. We have also moved what used to be S5.2.1.

      (2) Simulations

      The simulations raise, as mentioned in the Public Review, several concerns regarding their generality and validity.

      (R.2.10) (a) We recommend validating the simulation results by comparing them with simulations of the full nonlinear equations. The authors should at least provide the equations for the full dynamics and explain how the expansion is performed and why it is valid. This also includes verifying the assumed steady states (g_i=0 and g_i=c_i, where 1/c_i = \mu_i or \hat{\mu}_i).

      We are simulating the whole non-linear equations. Here it is important to stress, as we do now in the main text, that our results apply to any f, as long as it fulfills our R1-R5 requirements. However, for the simulations in the figures we have to use a specific f (since there is an infinite amount of fs that fulfill our requirements). Again the figures are just examples to visualize the types of resulting patterns and gene networks we talk about.

      In the original version we may not have been clear enough about the equations used for the simulations. The presentation of the Maini-Miura model has been revised to improve clarity (equation S6.1 in SI). In particular, the existence of a homogeneous steady state is now parameterized by a tunable g<sup>*</sup>, that can be chosen as for spike initial patterns or for homogeneous-with-noise and spikehomogeneous initial patterns). We have also included a proof that the model equations satisfies our conditions R1-5. Indeed, the model is non-linear as long as σ<sup>i</sup>≠0 for some gene product (as we explicitly assume).

      The derivation of this cubic model from a separate expansion of general reaction-diffusion dynamics can be found in the original paper (Miura & Maini, 2004), with further applications to pattern formation that supporting its validity in subsequent works (Marcon et al., 2016; Diego et al., 2018). Importantly, this expansion is independent of the linearization performed in the main text of our article to derive the dispersion relation. The reference to this separate expansion in the previous version was included solely for contextual purposes; however, we have removed it in the revised manuscript to avoid potential confusion.

      (R.2.11) (b) The use of a Jacobian that is independent of the steady-state contradicts the assumption of nonlinearity (requirement R2 (l. 192f)) of f. We ask the authors to clarify this.

      We believe this concern arises from a notational ambiguity in the previous version of the manuscript, which has now been corrected: the matrix appearing in the regulatory term has been renamed from J to W<sup>T</sup>. As stated in the main text, the jacobian of the regulation function f(g) evaluated at the homogeneous steady state must coincide with the transpose of the network weight matrix. With the current equations (S6.1), we have , from which we easily get . Also, it is clear that the Jacobian of f(g) is not independent of g.

      (R.2.12) (c) In Figure S3 and similar simulations, the implementation of the nonlinear terms is ambiguous. The function f shown does not correspond to the Jacobian, and it remains unclear how these components are ultimately implemented in the simulation code. Additionally, as mentioned, it does not fulfill the necessary conditions for the global steady state.

      The implementation of nonlinear entries in f(g) whenever they are needed is now made explicit in the corresponding subsection of section S6 in the SI. With the new notation it becomes clearer that the fs used can fulfill the necessary conditions for the global steady state.

      (R.2.13) (d) The given function f_8 in S8.10.2 cannot correspond to the mentioned network since the number of gene products does not match the Jacobian and the network.

      This was a typo that has now been corrected.

      (R.2.14) (e) The given parameters for the figures in the SI do not match the figures. Please check and ensure that the correct figure is referenced (e.g., S8.2 Figure 3)

      This was a typo in the numeration of the subsections in the SI that has now been corrected.

      (R.2.15) (f) It is unclear which units are used, and the units used for the non-dimensionalization should be provided so one can relate them to biological systems.

      It is now explicitly stated in the revised version that the model equations are formulated in arbitrary units. This implies that the model dynamics are consistent with the characteristic units of any particular biological system under consideration. No non-dimensionalization of the model equations has been considered.

      (3) Conceptual and Structural Clarity

      The manuscript suffers from a lack of structural clarity, which affects both readability and scientific coherence.

      (R.2.16) (a) In one of the central figures (Figure 4) supporting their main claim, the naming of the network is not consistent with the main text. The network category referred to as "Over-Turing" is never mentioned in the main text. We suspect this should actually be labeled as the "noise-amplifying network."

      Indeed. This has now been corrected. We now use only the term “Over-Turing” in the article.

      (R.2.17) (b) The Supplementary Information includes an analysis of dispersion relations to classify patternforming networks, but this approach is not mentioned or referenced in the main text.

      This part of the SI has been moved to the main text and the dispersion relation has been fully and explicitly integrated in the overall argument of the article.

      (R.2.18) (c) In relation to Figure 6, we found that the concept of "diversity of possible final patterns" would benefit from a clearer definition and explanation. It is not immediately evident how this diversity is measured or what criteria are used to compare different networks. For instance, it is unclear why the Over-Turing network - which generates both periodic and noisy patterns - is considered to exhibit low diversity, whereas the Turing networks, which produce only periodic patterns, are described as having high diversity.

      This was just a large typo. The figure has been corrected. The reasons for this differences are now described in the last three paragraphs of the section “The ensemble of possible pattern transformations from H gene networks and spike initial conditions” for the hierarchical networks and in the last paragraph of the section “Pattern transformations in L- subnetworks from spike-homogeneous initial patterns ”, for the noise amplifying networks and in the seventh paragraph of the section “Pattern transformations in the combination of L+ and L- subnetworks” for the Turing networks.

      (R.2.19) (d) Additionally, the dependence of final patterns on initial conditions is not clearly described. It seems that this relationship is only analyzed for non-trivial pattern formations, but this is not explicitly stated. Clarifying these points in the caption of Figure 6 would greatly help readers understand the interpretation and significance of the results presented in this figure.

      Indeed, we have done nothing for the trivial pattern transformations. We are now more explicit about this already from the introduction. This article is only concerned with non-trivial pattern transformations. For each type of gene network we now provide a more detailed description of how the resulting pattern depends on the initial pattern (in the sections for each gene network).

      (R.2.20) (e) The significance statement is simply a verbatim repetition of parts of the abstract. This defeats its purpose, which is to articulate the broader implications of the work. We urge the authors to rewrite this section with a focus on significance rather than summary.

      We have now corrected this.

      (R.2.21) (f) We suggest including a dedicated figure to illustrate the biological model, depicting cells, intracellular and extracellular compartments, and the presence or absence of boundaries between adjacent cells. Such a figure would significantly enhance readers' understanding of the system being discussed.

      We have now done that. See new figure 3.

      (R.2.22) (g) We encourage the authors to strengthen the 2D and 3D results presented in the paper by adding supporting citations, sharing implementation details, or providing a more in-depth analysis of these systems. If such additions are not feasible, it may be best to remove references to the 2D and 3D systems to maintain clarity and focus.

      In the new version of the article we explain why our results on which gene networks can lead to pattern transformation do not depend on the dimensionality of the system. In fact, none of our proofs or arguments assumes or requires a specific number of dimensions. The networks are the same no matter the number of dimensions. The types of possible patterns can be seen as manifesting themselves differently depending on the number of dimensions. In the current version of the manuscript we explain now, every time we explain a resulting pattern, how the pattern is in 1, 2 and 3 dimensions and why. We have added Figures 1 and 9 for that purpose. As we explain in the text, the resulting patterns that are noisy would be noisy no matter the number of dimensions and the ones that are based on a spike in the initial pattern have necessarily radial symmetry (in any number of dimensions). Similarly the periodic patterns will be periodic no matter the number of dimensions (although some aspects of it will change). Similarly, in the 5th paragraph of the discussion we discuss the effects of the shape of the system and the boundary. There was a problem with the definition of pattern transformation we used, but this has now been corrected, in P1 and P2 in the introduction.

      (R.2.23) (h) The results section lacks a consistent structure. Section titles do not clearly indicate which phenomena or initial conditions are being analyzed, making it hard for readers to track the logical progression of the study.

      Now the results start with some introductory results with the subsections:

      “Basic requirements on gene networks capable of pattern transformation”

      The rest of the results are split into four clearly differentiated sections:

      “Gene network classification”

      “Linear stability Analysis”

      “Positive regulatory loops determine the kind of RD-instability”

      “Hierarchical Networks”

      “Emergent networks”.

      “Gene networks combining different classes of subnetworks”

      The last three sections have several sub-sections inside.

      We think that the titles of the sections are self-explanatory since hierarchical networks contain only H subnetworks while the emergent networks contain L+ or L- subnetworks and the last major sections is about how all these can be combined.

      Minor Issues

      (1) Notation and Terminology

      (R.2.24) (a) Variable naming is inconsistent throughout the paper. Terms like g_A(x) and A(x) (S5.2.1) are used for gene network concentrations without consistent usage. The naming of genes in networks also varies between the main text, SI, and figures. I.e., sometimes genes are labelled with small, sometimes with large letters, and sometimes with numbers.

      This has now been corrected.

      (R.2.25) (b) It would improve clarity to use distinct notations for intracellular vs. extracellular concentrations and gene expressions. Ensure networks and examples are consistent across all figures, captions, and supplementary materials. For example, RH2a and RH2b have different networks in the main text compared to the SI.

      As we now explain in the third paragraph of the “Methods: the model” section we consider, for simplicity, that gene products are either intracellular or extracellular. In that sense there is no possible ambiguity. As explained in that section, again for simplicity, we do not consider the receptor nor the signal transduction pathways of signals. This means that an extracellular gene product can “directly” regulate intracellular gene products. Because of that, we think that using different notations for extracellular and intracellular gene products would make things more confusing. We have corrected the misnaming between main text and figures.

      (R.2.26) (c) We suggest using distinct notation for the gene product itself and for its small deviation from a homogeneous steady state in the SI. This would help clarify whether specific statements apply only within the linearized regime or can be generalized to the full nonlinear dynamics.

      We do that in the new version of the article.

      (R.2.27) (d) Line 327 contains a mistake: g_k = g_j should be expressed as a proportional relationship. The division by g_A also seems unnecessary - please revise.

      This is now explained in a different way so this mistake does not apply.

      (2) Model Description

      (R.2.28) (a) Justify why boundary effects and spatial separation between cells can be neglected in the model.

      This is now discussed in the 5th paragraph of the model section. We do not claim that boundary effects are negligible. We claim, instead, that which are the gene networks that can lead to pattern transformations do not depend on the boundaries. The same occurs for the types of resulting patterns, in the coarse way we use, possible from each gene network and initial pattern.

      As stated in the first two paragraphs of the model section, the spatial separation between cells can be ignored because we assume there are many cells in the system and these are evenly spaced and sized (at least roughly). That is usually the case in animal development, although not always (there are exceptions in the very early stages of many marine invertebrates), and we do not claim to know exactly what happens in those cases: as we stated in the first paragraph of the introduction we assume systems made of many small cells.

      (R.2.29) (b) State explicitly that only extracellular gene products are assumed to diffuse - this is currently only mentioned in the SI.

      This is now explicitly stated early on in the first three paragraphs of the model section and also after the introduction of the model equations (1)-(3).

      (R.2.30) (c) In the Supplementary Information, the authors state that both extracellular and intracellular gene products can exhibit non-zero diffusion, which appears inconsistent with the conceptual framework and probably is a typographical error.

      This was indeed a typographical error. It is now corrected.

      (3) Assumptions and Requirements on f

      (R.2.31) (a) The equation for requirement R5 is incorrect as written in the main text and should be reformulated more rigorously. The condition should be stated for all constant values of g_i (and g_j) to avoid misinterpretation; otherwise, one might assume all matrix elements must have the same sign.

      This has now been corrected.

      (R.2.31) (b) Clarify what restrictions on f prevent pathological nonlinearities like 1/(g_k + \epsilon), which would contradict the assumed behavior at high concentrations.

      We do not understand this criticism. 1/(g_+\epsilon) fulfills our requirements on f and we do not see how is that pathological. We are unsure of what the reviewer means by the assumed behavior at high concentrations.

      (4) Figures and Captions

      (R.2.32) In Figure S3b, the diagram shows gene 5 being activated by gene 4, yet the caption states this is a negative regulation - please correct.

      This has now been corrected.

      (5) Readability and Formatting

      (R.2.33) (a) Improve navigation by hyperlinking references to equations, figures, and requirements throughout the document.

      In the new version we have inserted these hyperlinks.

      (R.2.34) (b) Adding hyperlinks to the requirements would additionally help the reader to keep track of them

      In the new version we have inserted these hyperlinks.

      (We.2.35) (c) Correct inconsistent or mismatched equation numbers and references. E.g. SI S5.1 is not referring to the correct equation (the equation it should be referring to would be Equation 3), and the reference to Figure 7 in part of the dispersion relation is wrong (as far as we see, this should be Figure 5).

      This has all been corrected now.

      (R.2.36) (d) Clarify ambiguous language in the introduction. For instance, the description of spike patterns (lines 136f) as a single cell spike contradicts the stated width (SI) and the visual representation involving 500 cells from the figures.

      This has now been corrected.

      (R.2.36) (e) The discussion of 2D and 3D simulations appears limited to the "noise amplifying" network. It's unclear whether a similar analysis was done for other network types.

      In Figures 1 and 9 and through the text we discuss all types of patterns in 2D and 3D.

      (6) Typos

      (R.2.37) Typos in the text (The following is just a small selection of the typos we came across. Since there are quite a few throughout the manuscript, we may not have caught all of them. We kindly recommend that the authors carefully proofread the full text to ensure consistency and clarity):

      We have corrected all the indicated typos and proofread the whole manuscript and SI.

      Reviewer #3 (Recommendations for the authors):

      Major concern:

      (R.3.1) Pattern formation can be induced by the positional information, and reaction-diffusion/Turing mechanisms is a foundational idea in the field. As in the references the manuscript cited, these paradigms were already clearly articulated and synthesized (e.g., Green & Sharpe's work (2015)). Moreover, the search for minimal network topologies that can generate Turing patterns has been extensively explored in Zheng et al. (2016). The novelty of the present work is unclear. It might offer a fresh perspective on an established problem, but it does not seem to present fundamentally new biological or mathematical advances.

      If the authors wish to strengthen the novelty and impact of the manuscript, they should consider explicitly acknowledging prior work and positioning their contribution as a formal extension or generalization, not discovery. To enhance the practical relevance of their work, the authors could demonstrate how their framework can be used to predict or classify gene network behaviors in pattern formation that are not easily identifiable through experimental approaches alone. For example, they could show how their classification helps distinguish between Turing, hierarchical, and noise-amplifying dynamics in complex or ambiguous biological systems, thereby offering a guiding tool for experimental design or interpretation.

      Indeed, the gene networks we identify have been identified before. We were and we are quite explicit about it, in the discussion, and we do cite the relevant work on that (including the one suggested by the reviewer). The novelty of the work is not identifying these gene networks, nor minimal ones, but showing that these are all the possible ones for pattern transformation (that there is no new type of network), this has not been done before (not even intended) and we are very explicit about that being our results (first paragraphs of the discussion).

      Minor concern:

      The writing style and language usage can be improved for clarity. Some explanations in the results and discussion can benefit from tight editing to eliminate redundancy and improve readability.

      We have corrected all the indicated typos and proofread the whole manuscript and SI.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Reviewer 3 (Public review):

      Comments on revised version:

      The current version of the manuscript is clear and complete. Kudos to the authors for their thorough revisions. My only remaining point concerns the definition of "report": "We define a report as any explicit behavioral response (whether verbal, manual, or otherwise) that communicates a participant's subjective state." It would be helpful to clarify whether this definition is intended to exclude purely internal, explicit self-reports that are not externally expressed. As currently formulated, the definition appears to require overt behavioral communication. However, this raises a conceptual issue in relation to the no-report paradigm literature, where the distinction between report, metacognitive access, and overt motor/verbal expression is precisely at stake.

      Could the authors specify whether "report" is meant to (i) be restricted to externally observable, behaviorally expressed reports, or (ii) extend to internally generated, explicit metacognitive judgments even when they are not communicated? Clarifying this point would help situate the manuscript more precisely within ongoing debates on the role of report in identifying neural correlates of consciousness.

      We thank the reviewer for prompting us to make this subtle but important distinction explicit. We agree that the two senses of "report", i.e., (i) externally observable, behaviorally expressed reports and (ii) internally generated, explicit metacognitive judgments that are not communicated, are conceptually distinct and that this distinction is precisely at stake in the no-report paradigm literature. We fully agree that sense (ii) (disentangling NCCs from covert metacognitive access) would be a valuable direction for future research. However, because the intracranial studies reviewed in the manuscript focus exclusively on distinguishing NCCs from overt behavioral reports, our definition is intentionally restricted to sense (i).

      To clarify this point in the manuscript, we added the following sentence at lines 111–114:

      "Note that the no-report intracranial studies described here attempt to distinguish NCCs from externally observable, behaviorally expressed reports, and not from internally generated metacognitive judgments that are not communicated."

    1. Author response:

      The following is the authors’ response to the original reviews.

      We have made several major changes in response to the comments and we feel that the manuscript is considerably stronger. In brief: 1. We have added substantial content about homeostasis and EI balance to the introduction. 2. We have addressed concerns about physiological relevance by performing calculations to show that the free calcium in our solutions is well within the physiological range, by citing previous studies showing that short-term plasticity is consistent across 33-38 ℃, and by doing simulations scaled to physiological temperatures to show that the key computational effects are retained. 3. We have addressed concerns about readability by extensive text rewrites, reformatting most of the figures, and by splitting figures into smaller, more focussed ones. 4. We have organized over 20 statistical evaluations and comparisons between our model and experiments into a table. 5. We have carried out additional calculations to examine how the optimal frequency for mismatch detection depends on parameters, and to show that mismatch detection remains even in the presence of stimulus jitter. 6. We have stated more clearly how our proposed mechanism for mismatch detection is based on transient plasticity-mediated skewing of EI-balance, and have added a schematic for the last figure to show this.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study uses optogenetics to activate CA3, while recording from CA1 neurons and characterizing the excitation/inhibition (E/I) balance. They observe use-dependent alterations in the E/I balance as a result of STP, and they develop a model to describe these observations. This is a very ambitious paper that deals with many issues using both experimental and modeling approaches.

      Strengths:

      This paper examines important principles regarding the manner in which synaptic circuitry and use-dependent synaptic plasticity can transform inputs and perform computations.

      Weaknesses:

      The use of selective ChR2 expression in CA3 cells is a good approach, but there are numerous issues that cause concern regarding the applicability of their slice recordings to physiological conditions and that make some aspects of their results difficult to interpret. Experiments are not performed under physiological conditions (high external calcium and low temperature), which makes the interpretation of their findings difficult.

      Calcium: We would like to reassure the reviewer that the free calcium levels in our solutions were at ~1.27 mM, well within the physiological range, since our aCSF solution used calcium buffers as well as CaCl2. We have added a section to the methods to show this calculation.

      Temperature: Klyachko and Stevens (J. Neurosci 2006) show that the facilitation, augmentation and filtering properties of the CA3-CA1 network were consistent between 33 and 38 degrees C, thus spanning our conditions of ~33 degrees C. Additionally, we have performed simulations to show that the mismatch detection computations remain pronounced (or are even strengthened) when simulation rates for kinetics and channels are scaled to physiological temperatures. Using a Q10 of 2, the scaling term for kinetics is ~37% faster. The outcomes are presented in Figures 7 and 9. We now state these points at the start of the results section:

      “Our bath solution had physiological levels of free ions including calcium (methods), and recordings were performed at 32-33 ℃ which has been shown in rats to yield similar short-term plasticity properties as at physiological temperatures (Klyachko and Stevens 2006b).”

      We have added a new section to the discussion “Relevance to in-vivo computation” in which we enumerate the caveats but also the points of convergence between our study and physiological conditions, to strengthen the interpretability of our results.

      In addition, the reliability of stimulating action potentials in CA3 pyramidal cells needs to be determined, particularly during high-frequency trains. If it is unreliable, there are alternative approaches that might prove to be superior, such as the use of somatically targeted ChR2.

      We acknowledge that somatically targeted ChR2 might have slightly improved the sparseness of stimuli, but even such localized expression could lead to unreliability if the position of the soma with respect to the illumination is such that the stimulus is near threshold. Instead, we have adopted a data-driven estimation of CA3 reliability. We reanalyzed our optically-triggered field potential readouts from CA3, to estimate their reliability individually and over trains (Figure 1).

      “Notably, the distribution of field amplitudes was very tight (Figure 1E), more so than the corresponding EPSPs (Figure 1H). Together with previous work using a similar optical stimulus system [6] we interpret this to say that the spiking responses from CA3 neurons to optical stimuli were consistent from trial to trial. The field response showed a slight decrease over the course of the pulse train of approximately 2% per pulse (regression fit slope=0.02, r2=0.05). We attribute this to ChR2 desensitization.”

      As a further bound to any functional outcomes of CA3 spiking (un)reliability, we point out that CA3-CA1 release probability is low (p~0.2). Any reduction in CA3 reliability is equivalent to reducing the probability of synaptic release, which is already treated as a stochastic process in our simulations. We were able to compare this to experiment as follows: We explicitly modeled the effect of different synaptic volumes as a surrogate for changing p_release in Figure 6-figure supplement 1, and mapped this to our data in Figure 6 D.

      “Then we compared the probability that each optical stimulus would elicit an EPSP (Figure 6 D). As expected, 15-square patterns (yellow dots) frequently gave an EPSP (77.5±11.7%), while 5-square patterns failed about half the time (51.4±16%). The simulated runs matched this (Table 1). The probability of failure reduced with increasing volume of the simulated presynaptic boutons, because larger volumes experienced smaller chemical noise (stochasticity) in synaptic release (Figure 6-figure supplement 1). We note that for the purposes of eliciting a postsynaptic response, any unreliability in optical stimulus-triggered firing of the CA3 neuron folds into the probability term for stochastic synaptic release. By matching this metric to experiment, we fine-tuned the volume scaling term for the presynaptic boutons to 0.2”

      In addition, a clearer, more detailed discussion of their model that distinguishes it from previous modeling studies would be helpful (and would make it seem less incremental).

      This is a good suggestion, as we regard our model as very substantially different from previous studies. We have incorporated this in the discussion as below:

      “Our current model is distinct in that it is truly multiscale, closely constrained by experiment, yet runs on modest hardware. It incorporates the network, a conductance based model of a CA1 pyramidal neuron, and chemical kinetic models of a population of stochastic synapses on its dendrite.

      Our network model is much reduced compared to models with exhaustive cellular and network-level detail44. Its simplicity enables extensive exploration of the network parameters and comparison with recorded activity under a series of well-controlled stimulus patterns (Figures 4-9).”

      We also point out that our proposed mechanism for mismatch detection is an advance over previous ones:

      “Leaving aside the obvious differences between auditory cortex and hippocampus, we frame our model as a transient differential tilt in EI balance (Figure 3, Figure 8A,B, Figure 10B), in distinction to the fresh-afferent model. This makes our model robust over a wide range of stimulus and network conditions (Figure 9), and has the functional implication that transient responses remain at about the same amplitude over a prolonged stimulus sequence (Figure 8B, Figure 10B), rather than declining.”

      Reviewer #2 (Public review):

      Summary:

      The authors investigate EI balance in the CA3-CA1 projections, emphasizing synaptic depletion and the implied rebalancing of excitatory and inhibitory projections onto a single CA1 Pyramidal cell. They present physiological results with optical stimulation in CA3 and measuring various response features in CA1, showing signatures consistent with the adjustment of EI balance. In particular, the authors emphasize a transient effect where the neuron escapes from EI balance, which can be used for mismatch detection. They partially replicate these results in a computational model that looks at detailed properties of synaptic plasticity in CA1.

      Strengths:

      The authors provide compelling evidence that non-specific modulation of synaptic plasticity, combined with their differential effects on excitatory and inhibitory neurons, can be used by CA1 excitatory neurons to detect changes in the population activity of CA3 neurons. Indeed, they provide insight into the potential computational role of transient EI imbalance.

      Weaknesses:

      The authors observe that "little is known about how EI balance itself evolves dynamically due to activity-driven plasticity in sparsely active networks." This is an overstatement, or better an understatement, given the extensive literature on EI balance (e.g. Wen W, Turrigiano GG. Keeping Your Brain in Balance: Homeostatic Regulation of Network Function. Ann Rev Neurosci. 2024. https://doi.org/10.1146/annurev-neuro-092523-110001 PMID:38382543). This way of framing the question does a disservice to the field and fails to contextualize the current research properly.

      We agree that we could have presented this better. Our focus was on short-term (<1 second) EI balance changes, but our statement did not set this context clearly. We rewritten and expanded the introduction to place our work in context of the substantial previous work on plasticity and homeostasis in EI balance.

      The evidence is incomplete because the authors do not show a specific relationship between synaptic change in CA1 and EI balance adjustment, i.e., the alternative could be that this is an unspecific effect unrelated to the specific regulation of EI balance and its functional role in the hippocampus and the cortex.

      We don’t quite follow this point. We have devoted Figures 2 and 3 to showing a specific relationship between short-term plasticity on CA3->CA1 synapses, and EI balance. In Figure 2 we show how E and I responses evolve over a pulse train. In Figure 3 we explicitly show the plasticity in E and I synapses, and then map it onto EI balance. In Panel 3E to G all these points come together and we show how gamma (the measure of nonlinearity of summation) evolves over a series of pulses in parallel with plasticity in E and I. We have added some new data in Figure 7A, B to show how E and I contribute to mismatch detection.

      Indeed, the paper drifts from addressing EI balance to elucidating the mismatch detection.

      We acknowledge that we did not sufficiently articulate the role of EI balance terms in our subsequent analysis of mismatch detection. We have added several figure panels (Figure 7A, B), added a summary schematic (Figure 10) and redone the text and discussion. With these changes we make the point that mismatch detection can be better framed as a transient shift in EI balance.

      “we frame our model as a transient differential tilt in EI balance (Figure 3, Figure 8A,B, Figure 10B), in distinction to the fresh-afferent model. This makes our model robust over a wide range of stimulus and network conditions (Figure 9), and has the functional implication that transient responses remain at about the same amplitude over a prolonged stimulus sequence (Figure 8B, Figure 10B), rather than declining.”

      The second shortcoming is that they do not show that the stimulation of the CA3 neurons occurs in a physiologically realistic regime.

      We have responded to the concerns about calcium concentration and temperature above in the response to the first reviewer. From the text:

      “Our bath solution had physiological levels of free ions including calcium (methods), and recordings were performed at 32-33 °C which has been shown in rats to yield similar shortterm plasticity properties as at physiological temperatures (Klyachko and Stevens 2006b).”

      In addition, there is a concern about the mapping between physiological activity and our stimuli. It is true that the patterned stimuli we delivered were artificial. We make the point that they are nevertheless a much closer map to sparse physiological patterns than conventionally obtained through Schaffer collateral volleys:

      “We use optical patterned stimuli to stimulate a cross-section of CA3 neurons with a variety of distributed patterns, theta, and other frequency rhythms. These stimuli are sparser and more dispersed than Schaffer collateral electrical stimuli which tend to stimulate adjacent fibres and in most cases are very strong.”

      We have added a section to the discussion “Relevance to in-vivo computation” to more completely address these points.

      Nor do they analyze what the impact will be of the excitatory transient in "mismatch detection", and CA1,

      We are unsure what the reviewer means by the excitatory transient. At the level of CA3, we observe a narrow optically triggered field response for each light pulse. At the level of CA1, we monitor the responses due to activation of E and I synapses, and are able to observe peaks for each of the light pulses. We have analyzed all these features in figures 1 through 3, and they are also explicitly included in the model. Based on the reviewer’s comment we have further characterized the field responses in CA3:

      “We observed a small amount of ‘ringing’ of the field response which we interpret as either CA3 spiking in a burst, or recurrent activation of the CA3 neurons (Figure 1 supplement 2). The ringing was down to ~5% within 8 ms, supporting our treatment of the optical input as a tightly time-delimited event, and setting a low bound to any contribution to patterns by recurrence.”

      When this would occur at the level of the whole population, i.e., the physiological impossibility of triggering uncontrolled chaotic excitatory responses.

      Again, we are unsure what population or chaotic responses the reviewer has in mind. As mentioned above we have further characterized the field readouts of population responses in CA3 and have established tight limits on recurrent activity (Figure 1-figure supplement 2). In case the reviewer is looking for the outcome at the entire CA1 network as a whole, our experiment figures 1GH,J,K,L show sharp, single peak CA1 neuronal responses.

      In particular, when we consider CA3 as an attractor memory system, the range of deviations (mismatches) that a CA1 neuron can be exposed to and detect, given the model presented in this paper, might be below those generated due to CA3 pattern-completion dynamics.

      While this is an interesting question for further work, our study focuses on a tighter question, that of mismatch detection downstream of the CA3. As indicated above and in Figure 1figure supplement 2, our field and patch recordings show that under our stimulus conditions, the internal dynamics of the CA3 produce minimal delayed or recurrent signals. Thus, by design, the CA3 layer in our system acts as an almost pure input layer with minimal internal dynamics. In the discussion we address some of the possibilities that may arise from pattern computations in CA3 and other upstream areas:

      “We speculate that upstream areas may encode higher order stimulus features such as gaps, duration, intensity, localization, and frequency steps into distinct input patterns. Our proposed EI-balance shift mechanism could be a common end-point for all of these. This would transform quite complex mismatch detection tasks into a uniform computation of pattern change, generalizing the mechanism to stimuli which were previously considered to require a more complex network-level implementation”

      In addition, the match between the model and the physiological results is not fully quantified, leaving it to the reader to make a leap of faith.

      While the original version had numerous points of comparison between physiology and model, we agree that the values were scattered. In this revision we have tabulated them and performed additional statistical comparisons between model and data for a total of over 20 comparisons for the cell electrophysiology and network readouts (Table 1). We have also organized the preceding chemical kinetic comparisons in the supplements to Figure 4. We regard our study as one of very few to undertake quantitative experimental comparisons over such a range of readouts, experiments, and scales.

      In addition, the manuscript suffers from poor analysis and presentation. The work could be improved by putting more effort into translating results into insightful metrics.

      We acknowledge that the presentation needed improvement. We have performed a major rewrite and reorganized many of the figures. As mentioned above, we have tabulated numerous metrics (Table 1) and have characterized EI balance and its evolution due to plasticity in a pulse train (Figures 2 and 3). For higher-level metrics, the new figures now extensively explore how mismatch sensitivity depends on parameters, stimulus patterns, and repeat frequency (Figures 7, 8, 9). We have added a discussion section “Relevance to invivo computation”

      Overall, the authors have not achieved their original aim to show that the observed phenomenon is relevant to computation in CA1 or the brain outside of a highly controlled in vitro setup and reductionist single cell model.

      We feel that with this revision we have more clearly shown that our measurements are relevant to in-vivo computation, both through improved clarity and additional analysis. We have added a section “Relevance to in-vivo computation” in the discussion which enumerates the steps we have taken to support the relevance of our study. In the revision we have also performed several modelling extrapolations which encompass in-vivo conditions, such as testing jitter and frequency range. In a broader sense, in vitro work by design, is meant to be highly controlled so as to be able to get at mechanisms, and in our study we have delivered a range of physiologically relevant stimulus combinations to bridge the gap.

      The authors combine several techniques for in vitro whole-cell patch-clamp recordings with patterned optical stimulation of the CA3 network in the mouse hippocampus, which is consistent with the state-of-the-art.

      They introduce a metric of similarity between expected and observed response patterns, called gamma. The name is confusing given the wide use of the label gamma for oscillation frequencies above 20 Hz. Gamma is calculated as (E*O)/(E-O). This means that gamma approximates infinity as the difference goes to 0, to mention one of the problems. This metric is not interpretable, and it is not clear why the authors did not follow a standard approach, e.g., likelihood, correlation, or percent error.

      We acknowledge the potential for confusion, however we felt it would be more confusing to change nomenclature. The metric gamma is derived from previous published work (Bhatia et al, eLife 2019) describing nonlinearities in summation, which is cited. In that study and the current one, there was no instance in which gamma became unreasonably large. It is true that the term gamma is used for many concepts, but we feel that the contexts are so different between summation nonlinearity and oscillation frequencies that confusion is unlikely. We have taken care with the wording in the text to further disambiguate the usage.

      The authors aim to replicate the physiological results with an "abstract model of the hippocampal FFEI network. In practice, this is a conductance-based model of a single CA1 neuron, including chemical kinetics-based multi-step neurotransmitter vesicle release. This is an abstraction from the FFEI network that the paper starts with.

      We stress that the full model was used for all simulations except synaptic chemistry parameter fitting. We have clarified this point in the text and discussion section. From the text following Figure 4:

      “We used this full model, with optical stimulus, CA3, Interneurons, CA1 neuron, probabilistic connectivity, and presynaptic signaling chemistry, for all subsequent calculations in this study.”

      The model has 256 integrate-and-fire CA3 neurons, 256 interneurons, plus 200 inhibitory and 100 excitatory synapses onto the CA1 neuron, in each of which we have distinct multistep transmitter release kinetics.

      It raises the question whether this is the right level at which to model the computational impacts of EI imbalance on CA1 neurons. Given the highly reduced model they have elaborated, the generalization to the complete CA3-CA1 network that the authors suggest can be achieved in the discussion is overoptimistic. Network models of CA3 and C1 must be considered, together with afferents from the entorhinal cortex to accomplish this generalization.

      We hope we have clarified that we do indeed base all our calculations on the full FFEI model converging onto the CA1 neuron whose connectivity influences circuit function, and we feel that this is necessary and sufficient for our goals in this study.

      While the role of the recurrent CA3 network and EC would be interesting topics for future work, the scope of our study is to model the computational impact of EI imbalance in the FFEI network of CA3-> CA1 on CA1 neurons.

      The authors reveal a potentially interesting physiological feature of CA1 excitatory neurons under very specific stimulus conditions.

      We thank the reviewer for considering the work as interesting. We would like to clarify, however, that our stimulus conditions are actually multidimensional. Specifically, we have varied frequency, pattern, and number of inputs for burst stimuli, and we have also examined Poisson train inputs. In the model we have examined spiking responses, and theta modulated stimuli. In the revision we have also included jittered synaptic input, and obtained frequency dependence of the mismatch detection. To our knowledge this is among the more multidimensional stimulus-response and modeling studies on this system.

      It could warrant follow-up studies to place EI imbalance in a physiologically realistic context.

      Reviewer #3 (Public review):

      Summary:

      This work shows experimentally and computationally that single CA1 neurons can perform mismatch detection on patterned CA3 inputs and that STP and EI balance underlie this detection.

      Strengths:

      It has been known that STP can enhance the EPSP when the corresponding presynaptic input exhibits abrupt changes in firing rate. This work provides experimental evidence and further computational support for the hypothesis that the basic computation through STP is useful for detecting abrupt changes in the spatial pattern of synaptic inputs at the Schaffer collaterals. Further, their results indicate the novel view that mismatch detection is most efficient when gamma-frequency bursting inputs exhibit mismatches between theta cycles.

      Weaknesses:

      Their model assumes that patterned activities in CA3 do not have overlaps. However, overlaps between memory engrams have been shown. Therefore, this assumption may not hold, and whether the proposed mechanism is valid for overlapping CA3 inputs needs further clarification.

      We see that our account of the methods needs clarification, since we explicitly incorporate overlap in our model. First, from the experiments themselves, we say that we expect overlap:

      “This was also consistent with the observation of a wide field of excitability around individual CA3 neurons [6] (Figure 1-figure supplement 1). From this we expect that there is some overlap in the sets of CA3 neurons activated by different patterns, and this overlap increases with more stimulus squares.”

      In the model, we systematically examine the effect of overlap and have added several figures to make the point (Figure 9 Bi, Figure 9Ci, Figure 4-figure supplement 6, Figure 7figure supplement 1, Figure 9-figure supplement 1).

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      The use of selective ChR2 expression in CA3 cells is a good approach, but there are numerous issues that cause concern regarding the applicability of the slice recordings to physiological conditions and that make some aspects of the results difficult to interpret.

      Weaknesses:

      (1) Some aspects of this study seem somewhat incremental. There is a rich literature on the study of excitation and inhibitory synapses and the issue of EI balance. There are a great many related studies that are not cited (off the top of my head: Pouille and Scanziani 2001, Mittmann, Chadderton and Hausser 2004, Atallah and Scanziani 2009, but there are many, many more). A great many of the ideas presented in this study have already been published previously (Klyachko and Stevens, 2006, and numerous other related manuscripts).

      We agree that the topic of EI balance has a very substantial literature. We have incorporated many of the mentioned articles and others in our introduction and discussion. Our study explicitly links several strands of work on EI balance with short-term plasticity and spatial patterning:

      “The current study integrates several research themes of EI balance, short-term plasticity, and network computation to systematically characterize and model the properties of a network with feedforward inhibition. We complete the experiment-model-prediction-testing loop and show that differential changes on E and I synapses may provide a mechanism for single neurons to extract interesting features of spatiotemporal inputs through STP (Asopa and Bhalla 2023), while keeping mean activity steady.”

      We find that our sparse optical stimulation protocol gives qualitatively distinct results, and is amenable to investigation of more complex spatial pattern dependent effects. We have explicitly discussed the mentioned paper by Klyachko and Stevens, and numerous others, to point out where our study differs. From the discussion:

      “For example, studies using field electrode stimulation of the Shaffer collaterals report a sustained shift to excitation during burst input (Klyachko and Stevens 2006a). In contrast, our sparse optical patterned stimuli results in a small window of escape from EI balance around pulse 2 or 3 in a burst (Figure 3), following which both E and I undergo depression to restore balance (Figure 3, 8). Thus, spatial patterning intersects with short-term plasticity to add another layer of timing control through gating of E-I balance.”

      (2) There are multiple technical issues that call into question the relevance of this study for physiological conditions and the study of STP.

      (a) Their experiments were performed in elevated external calcium (2 mM) compared to physiological calcium (1.1-1.5 mM). This will have a major influence on the probability of release and short-term plasticity.

      This concern does not take into account the composition of our solution, which incorporated calcium buffers to give free calcium levels of ~1.27 mM. We have provided detailed calculations in the methods section.

      “Our bath solution had physiological levels of free ions including calcium (methods), and recordings were performed at 32-33 °C which has been shown in rats to yield similar short-term plasticity properties as at physiological temperatures (Klyachko and Stevens 2006b).”

      (b) Their experiments were performed at reduced temperatures (32-33 {degree sign}C). This is alright for many studies, but this is an important deficiency for the particular issue of EI balance and STP, and the relevance of conclusions based on these conditions.

      Klyachko and Stevens (J. Neurosci 2006) show that the facilitation, augmentation and filtering properties of the CA3-CA1 network were consistent between 33 and 38 degrees C, thus spanning our conditions of ~33 degrees C. Additionally, we have performed simulations to show that the mismatch detection computations remain pronounced (or are even strengthened) when simulation rates for kinetics and channels are scaled to physiological temperatures. Using a Q10 of 2, the scaling term for kinetics is ~37% faster. The outcomes are presented in Figures 7 and 9.

      (c) I like the selective expression of ChR2 in CA3 pyramidal cells, but they have not provided any information on the effect of stimulation on the firing of CA3 cells (Extended Data Figure 1 is not enough). Is it reliable for single stimuli or stochastic?

      We have used field recordings in the CA3 to put tight bounds on the properties of CA3 firing (Figure 1, Figure 1-figure supplementar 2.) The field recordings show that on average the firing is highly reliable. We explicitly characterize the probability of eliciting EPSPs through Poisson patterned stimuli in Figure 6 D. As discussed in the text (excerpted below) any stochasticity in firing folds into the parameters for p_release, and synaptic firing is itself stochastic.

      We note that for the purposes of eliciting a postsynaptic response, any unreliability in optical stimulus-triggered firing of the CA3 neuron folds into the probability term for stochastic synaptic release.

      Do CA3 cells fire once or multiple times?

      This was a useful point, and we examined our field potential data more closely based on this.

      “We observed a small amount of ‘ringing’ of the field response which we interpret as either CA3 spiking in a burst, or recurrent activation of the CA3 neurons (Figure 1-figure supplement 2). The ringing was down to ~5% by the third peak which occurred within 8 ms, supporting our treatment of the optical input as a single brief event, and setting a low bound to any contribution to patterns by recurrence.”

      Are the spikes precisely timed, or do they vary?

      Based on the field potentials, the spikes are precisely timed (Figure 1-figure supplement 2D, E).

      “fEPSP Peak Width distribution centred around 1.2 ms, but no peak was wider than 1.6 ms, suggesting tight synchrony in case multiple CA3 neurons were spiking.”

      Are there use-dependent changes in the ability of optogenetic stimulation to evoke spiking?

      Yes, and this is characterized in figure 1 panel I. The decrement is about 2% per pulse.

      The CA3 regions are highly interconnected with recurrent collaterals. Does stimulation during trains alter the activity in the CA3 region as a result of these collaterals?

      Based on the CA3 field recordings, almost all CA3 activity is optically triggered (Figure 1figure supplement 2). Figure 1C shows narrow fEPSPs in a burst.

      This would be a particularly important issue during trains. Would they have gotten more readily interpretable results if they had used a somatically targeted ChR2 variant?

      We feel it is unlikely that a somatically targeted ChR2 would change outcomes. All our analysis assumes overlap of excitation of CA3 pyramidal neurons, that is, a given spot illuminates multiple cells to different degrees, and that there will be neurons which are activated by more than one spot. Somatic targeting does not eliminate activation due to scattering and out-of-focal-plane illumination.

      In extended Figure 5, they show stimulus patterns used to stimulate. I need some more explanation. Are they stimulating in the cell body region only, or are they stimulating in the vicinity of dendrites?

      Extended Figure 5 (now Figure 4-figure supplement 6) indicates the stimulus patterns in the model. The experimental illumination pattern was 336µm x 187.2µm oriented so that the long axis of the pattern lay along the CA3 cell body layer (methods). Sample stimulus patterns are illustrated in Figure 1 panels D, G and J. Given scatter and out-of-plane illumination we expect that dendrites will also be stimulated. This is corroborated in Figure 1figure supplement 1 where we find that in addition to a strong ‘receptive field’ at the soma, there is a dispersed region of weaker activation. We cannot say definitively whether this dispersed region is due to light scatter, out-of-plane illumination, or dendritic activation. However, even somatically targeted ChR2 would elicit multi-neuron activity due to scatter and out-of-plane soma activation.

      If that is the case, there are a great many complications that arise, and it seems to be an approach that could unreliably activate a great many CA3 cells.

      We have now put in a paragraph to discuss this, and to set bounds to the unreliability.

      “To monitor the strength and consistency of the total resultant optogenetic activation of the CA3 layer, we used an extracellular field electrode in the CA3 stratum radiatum (Figure 1A, methods). The field response correlated well with optically-driven CA1 PC depolarization (Figure 1E-G), and scaled with the size of the pattern (Figure 1F). This was also consistent with the observation of a wide field of excitability around individual CA3 neurons (Bhatia et al. 2019) (Figure 1-figure supplement 1). From this we expect that there is some overlap in the sets of CA3 neurons activated by different patterns, and this overlap increases with more stimulus squares. Notably, the distribution of field amplitudes was very tight (Figure 1E), more so than the corresponding EPSPs (Figure 1H). Together with previous work using a similar optical stimulus system (Bhatia et al. 2019) we interpret this to say that the spiking responses from CA3 neurons to optical stimuli were consistent from trial to trial.”

      We also note that any CA3 firing unreliability folds into the stochastic release terms, as discussed in an earlier point.

      (d) As far as I can tell, they did not examine the effects of blocking NMDA receptors in their slice experiments. This seems like a very important experiment to perform if they really want to understand EI balance.

      The reviewer is correct that we did not block NMDA receptors. While this would have teased apart contributions of NMDAR and AMPAR to the overall response, our analysis of EI balance required the intact synapse and hence this decomposition (which has been done in previous studies) was not needed for our analysis.

      Based on a-d it is not clear that their conclusions regarding EI balance and STP are relevant under physiological conditions, and their findings are difficult to interpret.

      We have addressed the concerns about physiological conditions when it comes to the Ca2+ levels and temperature. We do not feel that points c and d alter the interpretation of our findings.

      Minor:

      (3) Their model has only 1 type of interneuron, whereas there are many. CA3-interneuron synapse has very different plasticity for different types of interneurons, and different types of interneuron synapses onto different parts of the CA3 cell. They need to justify lumping all of these types of interneurons.

      We agree that our model had a coarse-grained representation of interneurons as a single class. We feel this is an appropriate level of detail because it fits well for our experiments, and keeps the model tractable.

      “We have, of course, simplified the network, most notably in the use of only one inhibitory interneuron class which maps to parvalbumin-positive fast-spiking interneurons with perisomatic connectivity. This level of detail was chosen as it was able to quantitatively fit a large number of observations with minimal circuit complexity.”

      (4) How many parameters can they adjust in their model? It seems that with so many parameters, their model is not very good at times (extended Figure 3B and E, for example).

      Our model has 6 free parameters for the network (Table 1), and another 7 parameters each for the E and I presynaptic plasticity models (Figure 4A and Supplementary Data). The presynapse plasticity parameters are directly assigned from the burst response recordings using the parameter fitting as described in the Methods. Normalized RMS differences between model and experiment for presynapse parameters are presented in Figure 4-figure supplements 1-3, panel F. Most traces lie below 0.3, which is a good fit. We have now tabulated numerous comparisons between model and experiment (Table 1). In all but 1 of 20 tests, the model value lies within the experimental range.

      “Overall, we were able to quantitatively replicate almost all features of the experimental dataset in our multiscale model incorporating presynaptic signalling, postsynaptic electrophysiology, and abstracted network connectivity and responses. Between the datasets in Figure 4-figure supplements 1 to 3, Figure 5, and Figure 6, we were able to substantially constrain the parameters in our model, from chemical to cellular physiology to network.”

      Additionally, we have included a new Figure 9 to systematically do parameter sweeps. From this we conclude:

      “...mismatch detection in our model is robustly present and can be tuned over a wide range of network parameters and model assumptions, with the notable exception that it is absolutely dependent on the presence of STP.”

      (5) They use the term short-term potentiation (STP), but plasticity is not just enhancement; there is also depression. That is why many others opt for the more inclusive "short-term plasticity".

      We agree that this was unclear. We meant to use “Short Term Plasticity” and have now clarified this in the text.

      Reviewer #2 (Recommendations for the authors):

      The paper is poorly written and would benefit from a more careful preparation of the manuscript. In the opinion of this reviewer, it does not meet the expected quality for a paper of this type. Reviewing the paper was somewhat frustrating, requiring puzzling through details that were not well described. Also, failing to put clear labels on figures and their low quality did not help.

      We have worked substantially on the readability in the revision. We have made numerous changes to the text and figure legends, and have reworked several figures, with the goal of addressing concerns about readability.

      The introduction lacks proper context for EI balance and the hippocampus.

      We have substantially rewritten the introduction to more clearly place our work in the context of the relevant literature. We touch upon short-term plasticity and computation, on homeostasis, on EI balance and on network correlates of plasticity such as mismatch detection.

      The data analysis is superficial, and insufficient effort is put into compressing complex data into insightful metrics.

      We have done substantial rewrites to address this concern. There are two kinds of metrics we have developed for this study: those that measure the goodness of fit between simulations and data (consolidated into Figure 4-figure supplements 1 to 3 and in Table 1), and those which capture high-level features such as sublinearity of summation due to EI balance (Figure 3), selectivity for mismatch detection (Figures 7 to 9), and peak frequency for mismatch selectivity (Figure 9). We have also performed additional simulations as per reviewer suggestions, which give metrics for dependence of transition detection on network parameters, and for sensitivity of mismatch detection to input spike jitter.

      The only attempt to do this was the gamma measure, which left one wanting (see above).

      We have responded to the points about the gamma measure above.

      Figures are low-quality, labels are missing,

      We have substantially reworked figures, their labels, and legends. The automated mapping from our high-resolution figures to PDF seems to have blurred many of the figures, however, links to the originals should be there in the revision.

      And the analysis stays too close to the data without presenting a clear quantitative synthesis and insight.

      Please see response above. We have tried to balance the process of characterizing numerous readouts and making a model that closely matches experiment, with the high-level insights by way of computational outcomes such as mismatch detection in a variety of more physiological contexts (pulse trains and theta patterned inputs, Figures 7 to 9).

      Key results and mapping between physiology and the model are kept subjective and not quantified.

      Please see response above. We have consolidated our comparisons between physiology and experiments into Figure 4-figure supplements 1 to 3 and Table 1.

      In addition, the similarity measure gamma, which is introduced to express the relationship or the modulation of the response, is mathematically naïve and not well-motivated. It will approach infinity when expected and actual values become more and more similar. While this might be the range where sensitivity is required.

      Please see response above. The metric gamma is derived from previous published work (Bhatia et al, eLife 2019) describing nonlinearities in summation, which is cited. In that study and the current one, there was no instance in which gamma became unreasonably large. It is true that the term gamma is used for many concepts, but we feel that the contexts are so different between summation nonlinearity and oscillation frequencies that confusion is unlikely. We have taken care with the wording in the text to further disambiguate the usage.

      Some detailed observations:

      P2: What is an "interesting" feature?

      We have replaced the word “interesting” with “salient”:

      “We complete the experiment-model-prediction-testing loop and show that differential changes on E and I synapses may provide a mechanism for single neurons to extract salient features of spatiotemporal inputs through STP (Asopa and Bhalla 2023), while keeping mean activity steady.”

      P6 L110: However, over the pulse train, E and I underwent distinct STP profiles (Figure 1 M).

      What makes them distinct?

      This panel is now removed. A clearer account is presented in Figure 2D,E and F:

      “The EPSC showed a trend of early potentiation followed by depression (Figure 2D, 2E), while the inhibition underwent depression from the start (Figure 2 D, Fi)”

      P6 L115: Why can recurrent excitation in the CA3 segment be excluded?

      We thank the reviewer for pointing us to a more detailed analysis, which is now presented in Figure 1-figure supplement 2. We have added the following text:

      “We observed a small amount of ‘ringing’ of the field response which we interpret as either CA3 spiking in a burst, or recurrent activation of the CA3 neurons (Figure 1-figure supplement 2). The ringing was down to ~5% by the third peak which occurred within 8 ms, supporting our treatment of the optical input as a single brief event, and setting a low bound to any contribution to patterns by recurrence.”

      P8 F2A: How are the responses normalized?

      In the text we state:

      “All the PSPs of an 8-pulse train were normalised to the probe pulse.”

      We have added this line into the legend.

      “Traces were normalised to a reference pulse 0, delivered 300ms before the burst.”

      Explain why, given this normalization, the 15 square stimulation is less effective than the 5 square one.

      We acknowledge this was unclear. In the revised text we explain:

      “For the EPSCs, the 15-square trials had a higher reference pulse and higher stimulus overlap (discussed below), hence their normalised peak values were smaller (Figure 2E).”

      F2D: Where do you show that the biphasic response is a statistically significant deviation?

      Thank you for pointing out this missing analysis. We have added it in Figure 3A.

      P9 149: E should be E&F.

      Corrected.

      P9 L150: Explain the "ii" indexing.

      Corrected.

      P10: It is a bit clumsy to call the measure gamma. For general observation on the equation, see the general remark above.

      Please see discussion on this. We are reusing a published term.

      P10 L175: How do your results and F3G show divisive inhibition?

      In the current study we’re not setting out to show divisive inhibition, as that work has been published (Bhatia et al, eLife, 2019). We’ve corrected the text accordingly.

      “Using responses from the reference pulse, we replicated earlier observations (Bhatia et al. 2019; Wehr and Zador 2003) showing divisive normalisation, and obtained a median gamma of 7.16 (95% CI = 4.76 - 10.2)(Figure 3G).”

      Becomes

      “By comparing observed vs. expected responses, we replicated earlier observations (Bhatia et al. 2019; Wehr and Zador 2003) showing sublinear summation, and obtained a median gamma of 7.16 (95% CI = 4.76 - 10.2) (Figure 3G).”

      P16: How does F5 demonstrate a good match between model and physiology?

      We acknowledge we left this out. In F5E we show the model and experiment distributions over different frequencies. Our previous analysis only reported frequency dependence, and now we have added the comparison of response amplitudes. We have inserted the analysis and consolidated the results into Table 1.

      P18 l281: 15-square patterns (yellow dots) almost always gave an EPSP, while 5-square patterns frequently failed.

      Where can I see this? It is mentioned in the caption, but legends are absent.

      In the original source file and in the original confirmation pdf from eLife, the figure legend is present, and has an entry for panel C and D.

      “C,D:probability of trigger to generate a peak in the EPSP trace”

      In the revised version we have quantified these values and put the comparisons into Table 1:

      “Then we compared the probability that each optical stimulus would elicit an EPSP (Figure 6 D). As expected, 15-square patterns (yellow dots) frequently gave an EPSP (77.5±11.7%), while 5-square patterns failed about half the time (51.4±16%). The simulated runs matched this (Table 1).”

      P19 l307: Overall, we were able to replicate numerous features...

      Please be specific. What exactly did you replicate? How is it statistically demonstrated?

      This is a good point, we have updated the text to more systematically work through comparisons and metrics. We have also added some further metrics for features of the responses in Figures 5 and 6. As a way to organize all our comparisons we have added Table 1.

      P22 l341: The transient responses must be proportional to the overlap. Please quantify this effect more precisely.

      In Fig 7 panels L and O we had previously quantified the amplitude of transient responses with respect to two parameters closely related to overlap: pattern sparseness and probability of connections from CA3 to CA1. In Figure 9Bi we show that there is a complex and frequency-dependent relationship between overlap and mismatch responses. In the revision in figures 8 and 9 we have recast the “pattern sparseness” term as the more intuitive “overlap”. These are related almost linearly with a negative slope (Figure 9-figure supplement 1).

      P22 l342: What does "in E" mean?

      Should be Figure 7E for the original version. In the revised paper we have removed this panel.

      l347: I cannot follow. How do these single traces (7C-E) show these effects?

      We acknowledge that the figure and legend did not clearly indicate the timings of the transitions. We have completely redone and reduced figure 7 to simplify the presentation. The timing of transitions between patterns is now indicated using red triangles.

      What does denser connectivity refer to?

      Denser connectivity refers to a higher value for probability of connection between CA3 and CA1. In the revised version we have changed the figure to refer to stimulus overlap:

      “None of the transitions in Figure 8D (dense stimuli, 34% overlap) were significant, but two transitions in Figure 8E were significant (sparse stimuli with 2.5% overlap, p = 1.53e-5 and 6.1e-5).”

      P26: It is unreasonable to expect a reader to put this puzzle together.

      We acknowledge that this is a large and complex figure. In response to the reviewer’s input we have split the figure between Figures 7 and 9, and removed some panels, so as to make it easier to navigate.

      Reviewer #3 (Recommendations for the authors):

      (1) Which parameters are crucial for determining the preferred frequency (i.e., gamma frequency) for mismatch detection? This point should be addressed further.

      This is an interesting suggestion and we have performed additional simulations to address it. It turns out that the frequency tuning is very broad, over almost the entire gamma range from 40 to 200 Hz, and is indeed tuned by simulation parameters. We have placed these findings in Figure 9 in the new version of the paper.

      (2) The meanings of horizontal and vertical color bars should be explained in the legend of Figure 2A. Do they show the average values over columns and rows? A similar question applies to Figure 3G.

      We have removed the marginal heatmaps from Figures 2 and 3 as they were not contributing to the interpretation.

      (3) I wonder whether the proposed mismatch detection is tolerant against timing jitters in repeated presynaptic spike patterns. This information allows us to infer the accuracy required for neural code using population spike patterns.

      This is a good suggestion. We have run additional simulations to quantify this. It turns out that jitter has a clear effect on mismatch detection, and affects 5-square (low-overlap) patterns differently from high overlap (15 square) patterns. The latter see a boost in selectivity with 6 ms jitter. This comparison is now in Figure 7E ii and 7 Eiii

    1. Author response:

      The following is the authors’ response to the previous reviews

      eLife Assessment

      This study makes a valuable contribution to understanding how negative affect shapes food-choice decision making in bulimia nervosa by leveraging a mechanistic drift diffusion model to quantify the weighting of tastiness and healthiness attributes. The evidence is solid, supported by a randomized crossover design and generally appropriate statistical analyses. However, the interpretability of the findings is limited by ambiguities in the affect manipulation, particularly regarding whether neutral and negative inductions yielded reliably distinct affective states at the time of task performance in the bulimia nervosa group. Consequently, session-related differences in model parameters cannot be unequivocally attributed to negative affect rather than to uncontrolled state or contextual factors, and clearer separation of affective conditions alongside analyses aligned with the paired data structure would strengthen the conclusions.

      We thank the Editor and Reviewers for their careful summary of the study's strengths and for their constructive feedback.

      The eLife Assessment identified two specific limitations that qualified the strength of evidence:

      (1) ambiguity regarding whether the two affect inductions yielded reliably distinct affective states in the BN group at the time of task performance, and (2) analyses that were not fully aligned with the paired data structure. We have directly addressed both concerns in this revision. We provide explicit statistical evidence confirming that neutral and negative inductions yielded distinct affective states in the bulimia nervosa group; and we have re-analyzed all DDM parameters using updated mixed-effects regressions with an unstructured covariance matrix that appropriately accounts for the paired data structure. For completeness, we have also added the requested difference-in-difference analysis. Both approaches yielded conclusions consistent with those originally reported.

      In light of these revisions, we would be grateful if the Editorial Team would consider whether the strength of evidence rating might be updated from "solid" to "convincing." All changes in the revised manuscript are marked in blue.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      Using a computational modeling approach based on the Drift and Diffusion Model (DDM) introduced by Ratcliff and McKoon in 2008, the article by Shevlin and colleagues investigates whether there are differences between neutral and negative emotional states in:

      (1) The timings of the integration in food choices of the perceived healthiness and tastiness of food options in individuals with bulimia nervosa (BN) and healthy participants (2) The weighting of the perceived healthiness and tastiness of these options.

      Strengths:

      By looking at the mechanistic part of the decision process, the approach has potential to improve the understanding of pathological food choices.

      Weaknesses:

      I thank the authors for revising their manuscript.

      I still notice that the authors did not go through their manuscript to look for wordings refering to a prediction interpretation of their results while I already highlighted the inappropriateness of this wording in my two first rounds of reviews: e.g. there is still "we used zero-inflated negative binomial models to predict the three-month frequency" and I can find other statements like this. The design of their study does not allow such claims.

      We thank the Reviewer for identifying cases where the term “predicted” may mislead readers about the causal nature of our claims. We have made the following edits (changes are italicized):

      Methods (lines 516-518): “For these exploratory analyses, we used negative binomials to test the association between parameter estimates and the three-month frequency of retrospectively reported Objective Binge Episodes (OBE) and Subjective Binge Episodes (SBE).”

      Figure 5 (lines 881-882): “Affect-induced changes in information onset were associated with more frequent subjective binge episodes.

      The authors answered my major concern regarding the experimental induction towards a negative or a neutral state before running the food decision task. My concern is: BN patients already seemed to be already in a high negative state before undergoing the neutral induction, while these patients are in a lower negative state before undergoing the negative induction. It is therefore not surprising that patients seem to report a similar level of negative state after the two inductions (according to the figure of the authors' previous article). Of note is that the additional analysis the authors ran within the BN group only provides a significant result: this result shows that there has been an induction but does not rule out that patients were in the exact same magnitude of negative state to perform the task as the figure in their previously published article suggests it. The major issue is to show that:

      (1) As compared to the neutral induction, there has been a higher variation in negative state after as compared to before the negative induction.

      (2) The magnitude of the negative state after the negative induction is higher than the magnitude of the negative state after the neutral induction.

      The first point shows that the induction worked. The second point shows that the participants are in two distinct states. Without showing the second point, it may be possible that one induction increases the negative state of participants to the same level as the one of the second induction that has not increased anything.

      Within this context, how is it possible to associate, in patients, a difference in the DDM between the two sessions to a negative state (which is one of the main focus of the article) rather than to another parameter that has not been captured? A similar situation would be in an experiment studying the consequence of stress, a stressfull induction over relaxed participants attending the lab has high chances to raise the level of stress of those participants to the same level as the one that the same participants would experience after a neutral induction when these participants attend the lab with an already high level of stress. In that case, would it be approrpiate to claim that a difference at a task performed after the induction would be related to stress while the participants would be at the same level of stress when performing the task despite the fact that the induction worked ?

      In the experiment performed by the authors, the additional analysis to perform would be a paired sample t-test (or the appropriate non-parametric test) to check whether the magnitude of negative state of BN patients was different between the negative and neutral conditions after the induction only. If not, associating the difference at the DDM with negative states in BN is highly misleading.

      We thank the Reviewer for pressing on this point, and we apologize that our previous response did not make this sufficiently explicit. We agree with the Reviewer that two things must be demonstrated: (1) that the negative induction produced a greater change in negative affect than the neutral induction, and (2) that the magnitude of post-induction negative affect was higher following the negative induction than the neutral induction. We had included the results of analyses addressing both points in the Supplementary Materials of our previous submission, but we appreciate that we had not made this clear in our response.

      Regarding point (1), the mixed-effects model in Supplementary Table S1 yielded a significant Affect Condition × Timing interaction (β = 20.43, SE = 6.35, t = 3.22, p = 0.002), confirming that negative affect increased significantly more from pre- to post-induction in the negative condition than in the neutral condition. This is further supported by within-BN-group analyses in the Supplementary Materials: the negative affect induction produced a large, significant increase in negative affect (mean difference = 20.36, SE = 4.21, t = 4.84, p < 0.0001, Cohen's d = 0.97), whereas the neutral induction was not associated with a significant change in negative affect (mean difference = 7.16, SE = 4.21, t = 1.70, p = 0.327, Cohen's d = 0.34).

      Regarding point (2), we directly compared post-induction negative affect between conditions within the BN group, as requested by the Reviewer. The magnitude of negative affect was significantly higher following the negative mood induction than after the neutral mood induction (mean difference = 17.40, SE = 4.21, t = 4.13, p = 0.0003, Cohen's d = 0.83). This large effect size confirms that participants with BN were in meaningfully distinct affective states when performing the food decision task under the two conditions.

      Together, these analyses establish (1) that the induction worked as intended, and (2) that the two post-induction states were both statistically and practically distinct. We have added explicit language to the manuscript to make both of these points clear (lines: 181-185):

      Critically, post-induction negative affect within the BN group was significantly higher following the negative affect induction than after the neutral affect induction (mean difference = 17.40, SE = 4.21, t = 4.13, p < 0.001, Cohen's d = 0.83; see Supplementary Materials for full details), confirming that BN participants completed the food decision task under meaningfully distinct affective states across the two sessions.

      I read carefully the authors' answer related to mixed models: they claim that mixed models take into account correlations within their repeated data. The specification of the structure of the covariance matrix allows to control only partly for that. I notice that the authors did not specify the structure of that matrix: the article they refer to justify the appropriateness of their analyses is not adapted. The specification of the structure of the covariance matrix needs to address, in a mixed model, the difference in handling 4 repeated data per participants that cannot be paired as compared to 4 repeated data that can be paired (two per session with one before and one after the neutral or negative priming sessions, if I count right). Of note is that a covariance structure that is left free of constraint for the fit of the model does not capture appropriately the pairing of the data: it has all chances to capture the covariance in a different way. And a covariance structure that has constraints has more chances to lead to a model that cannot be estimated because of an absence of convergence of the algorithms.

      By the way, a single two-sample t-test (or a Mann-Whitney test if appropriate), and not a set of multiple paired-sample t-test as the authors suggest, would answer the goal of the authors to test for what they call the three-way interaction in their comment. This test would be performed between the two groups of participants (BN/controls) with the computation for each participant separately: (assessment after neutral induction-assessment before neutral induction)-(assessment after negative induction-assessment before negative induction). This analysis answers points 1, 2 and 4 they raise together with my point of controlling for the paired data. I would have agreed with their choice of a mixed model if they had an unbalanced dataset within each participant.

      We thank the Reviewer for this clarification, and we apologize that our previous response did not adequately distinguish between two different sets of analyses: (1) analyses of DDM parameter estimates, which involved four observations per participant (2 affect conditions × 2 food types); (2) trial-level analyses of choice and response time behavior, where each participant contributed many trials per condition and the dataset is genuinely unbalanced across participants due to trial exclusions – precisely the situation where mixed-effects models with participant-level random slopes are appropriate. The concern about covariance structure applies specifically to the DDM parameter analyses, but does not apply to our trial-level analyses.

      We also want to clarify a point about the task design that may have caused confusion. The Food Choice Task was administered only once per session, after the mood induction (i.e., once after negative mood induction, and once after neutral mood induction). As detailed in Figure 1, the task was not completed pre-induction. The four observations per participant in the DDM parameter analyses therefore reflect 2 affect conditions × 2 food types assessed within each condition, not a pre/post structure. This does not change how we address the concern about covariance structure, as there is still a nested feature of interest (food type within condition), but we wanted to correct this misunderstanding explicitly.

      For the DDM parameter analyses, we agree with the Reviewer that the original random effects structure did not adequately account for the paired nature of the four within-person observations.

      We have addressed this in two ways.

      First, we re-estimated the mixed model specifying an unstructured covariance matrix using the nlme package, which places no constraints on the correlation pattern among the four withinperson observations. We acknowledge the Reviewer's point that an unconstrained covariance matrix is not guaranteed to recover the within-session pairing structure. We explored whether a more constrained specification would be preferable. Specifically, we tested a nested random effect of affect condition within subject, which would directly encode the pairing of Low-Fat and High-Fat observations within each session. However, this model failed to converge. This is not a numerical issue but a fundamental identification problem: with only two observations per session per subject, the session-level and residual variance components cannot be separately estimated. We therefore selected the unstructured model as a more conservative option. Importantly, even if the unstructured model does not explicitly encode the pairing, it is a more general mathematical formula which would not impose incorrect constraints on the correlation structure.

      Consistent with our original findings, the mixed model with an unstructured covariance matrix yielded a significant three-way interaction (Group × Condition × Food Type: β = 0.28, SE = 0.12, t = 2.36, p = 0.020). All simple effects analyses have been updated to reflect the models with this covariance structure, and these are reported in the updated Supplementary Tables.

      Second, following the Reviewer's suggestion (adapted to the actual design structure, in which the Food Choice Task was administered once per session after the mood induction rather than before and after), we computed a difference-in-difference score for each participant's relative attribute onset parameter (τ<sub>s</sub>) following the affect inductions: (negative condition, high-fat − negative condition, low-fat) − (neutral condition, high-fat − neutral condition, low-fat). This score directly encodes the paired structure by construction, bypassing the covariance specification problem entirely. Consistent with the Reviewer's recommendation to use a non-parametric test where appropriate, we used a Wilcoxon rank-sum test (equivalent to Mann-Whitney U) to compare these difference scores between groups. The results confirmed that BN participants showed significantly larger food-type-specific changes in τs following negative affect induction relative to HC (W = 156, p = 0.018). We then applied this approach to all other DDM parameters (i.e., ω<sub>taste</sub>, ω<sub>health</sub>, α, τ<sub>ND</sub>, and z), and report these results alongside updated mixed-effects model results in the Supplementary Materials. The conclusions drawn from the difference-in-difference analyses were consistent with those from the mixed-effects models across all parameters.

      Both approaches converge on the same conclusion and we report both sets of complementary results in the manuscript: the updated mixed-effects models address the full factorial design in a single framework, while the added difference-in-difference analyses explicitly resolve the covariance specification problem by encoding the paired structure directly into each participant’s score, as the Reviewer recommended.

      Reviewer #2 (Public review):

      Summary:

      Binge eating is often preceded by heightened negative affect, but the specific processes underlying this link are not well-understood. The purpose of this manuscript was to examine whether affect state (neutral or negative mood) impacts food choice decision-making processes that may increase likelihood of binge eating in individuals with bulimia nervosa (BN). The researchers used a randomized crossover design in women with BN (n=25) and controls (n=21), in which participants underwent a negative or neutral mood induction prior to completing a food-choice task. The researchers found that despite no differences in food choices in the negative and neutral conditions, women with BN demonstrated a stronger bias toward considering the 'tastiness' before the 'healthiness' of the food after the negative mood induction.

      Strengths:

      The topic is important and clinically relevant and methods are sound. The use of computational modeling to understand nuances in decision-making processes and how that might relate to eating disorder symptom severity is a strength of the study.

      Weaknesses:

      Sample size was relatively small, and participants were all women with BN, which limits generalizability of findings to the larger population of individuals who engage in binge eating. It is likely that the negative affect manipulation was weak and may not have been potent enough to change behavior. These limitations are adequately noted in the discussion.

      We thank the reviewer for their thorough description of the strengths and weaknesses of this study.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The research investigates the frequency-dependent effects of transcutaneous tibial nerve stimulation (TTNS) on bladder function in healthy humans and via a computational model. The authors report that low-frequency (1 Hz) TTNS accelerates the urge to void, while highfrequency (20 Hz) TTNS delays it, corroborated by a computational model suggesting brainstem-mediated mechanisms. The work bridges experimental and theoretical approaches to propose a novel framework for TTNS applications in urinary retention.

      Strengths:

      (1) The integration of human experiments and computational modeling is a major strength. The model successfully replicates bladder dynamics and provides mechanistic insights into frequency-dependent effects.

      (2) Identifies potential therapeutic applications for urinary retention, a condition with limited non-invasive treatments.

      (3) Figures are clear and illustrative, and supplementary materials provide essential methodological depth.

      (4) Controlled experimental design (eg., single-blinded, fluid/caffeine restrictions, etc), detailed computational model parameters and validation against animal data, transparency in data exclusion criteria and statistical adjustments.

      Weaknesses:

      (1) The study uses healthy participants; extrapolation to clinical populations (e.g., urinary retention patients) requires validation.

      The authors have included a statement noting this and explaining that future work will explore this.

      (2) The simulated bladder capacity (100-150 mL) is lower than physiological ranges (300400 mL). While the authors note this, the impact on model validity should be further addressed.

      The authors acknowledge that the simulated bladder capacity and voiding efficiency of the model are lower than human physiological ranges. They have added an additional explanatory paragraph detailing this limitation and proposing the animal training data as a possible cause. Despite these limitations we do not believe this prevents the model from being used to explore proof-of-concept hypotheses (e.g., presence of frequency dependence, potential mechanistic bases) as in the present paper.

      (3) The model omits nociceptive afferents, limiting its applicability to pathological conditions like overactive bladder.

      The authors acknowledge that this is a limitation of the model, and have included a paragraph in the paper’s discussion detailing the limited scope of our in silico approach and clarifying the extent to which the results may be interpreted.

      (4) The lack of significant differences in urge intensity between groups (despite timing differences) warrants deeper discussion. Is the primary effect on efferent activity (as suggested) rather than sensory perception?

      The authors acknowledge that this is a surprising result and as such have deepened the discussion of the pilot study results, including hypothesizing as to potential explanations and suggesting further research in the area.

      (5) One of the highlights of this study is the identification of the effect of low-frequency (1 Hz) tibial nerve stimulation (TNS) on facilitating bladder contraction. Although the authors have clarified this effect in healthy participants, it would strengthen the conclusion if a UAB animal model (e.g., PMCID: PMC7927909, PMC8163611, PMC7847056, PMC8799394) were used to evaluate the same effect.

      The use of animal models is out with the scope of this study which aimed to act as a proof of concept work using a primarily computational approach backed by preliminary human data. The authors acknowledge that this does limit the strength of the conclusions. However, several animal models have been utilized in previous work (as cited in the publication) that demonstrate an excitatory effect of low-frequency tibial nerve stimulation. This work builds upon these previous studies to strengthen the case for a frequency dependent effect of the intervention.

      Reviewer #2 (Public review):

      Summary:

      Tibial nerve (electrical) stimulation (TNS) has emerged over the past 15 years as a non-invasive method to treat bladder overactivity, but interestingly, new animal work has suggested that TNS could actually be used to excite the bladder when appropriately tuning the stimulation frequency, effectively inverting its effect, perhaps opening the door to treat different conditions (e.g., UAB). The present study tests how healthy people respond to low and high frequency TNS, with the authors showing that they can substantially delay people's first sensation of bladder fullness with high frequencies (20Hz, shown many times before) but also that they can slightly hasten people's first sensation with low frequencies (1Hz, new result in humans). Moreover, the authors develop a computational model of interconnected conductance-based simulated neurons arranged in a physiologically plausible circuit that reproduces some aspects of the frequency-dependent effects of TNS. Their simulations suggest that we might expect low-frequency TNS to also increase the duration of bladder contractions in humans. The study highlights a potential new research direction, optimizing TNS stimulation parameters to increase basal bladder excitability.

      Strengths:

      The main strength of the work is to call attention to a new possibility of inverting the effect of TNS in humans by manipulating stimulation frequency, opening new indications for the therapy. This is highly relevant because of the recent popularity of TNS and its non-invasiveness, which lends itself to rapid testing and evaluation for new conditions and a high willingness to adopt. The authors convincingly demonstrate a modest excitatory effect on bladder sensation with low-frequency TNS, which clearly warrants further investigation.

      The high-level design of the hypotheses, concepts, and experiments is clearly articulated in both the methods and in particularly clear diagrams, letting the reader focus their attention on the most important findings.

      It is rare to develop a new computational model of the lower urinary tract at a systems level, and even more so for it to incorporate circuits in the spinal cord and brainstem centers, and this work undoubtedly advances the field's ability to engineer such systems. Further, because the model is comprised of linked conductance-based point-neurons, it is an excellent tool to investigate how an arguably plausible wiring diagram for neural control of the LUT could result in stimulation frequency-dependent effects on pelvic efferents. It is a proof of concept demonstrating how their mechanistic hypothesis of TNS could be implemented neurophysiologically by the nervous system.

      Weaknesses:

      The main drawback of the work is the frequent over-interpretation of the results. The human study and computational model are both proof-of-principle studies because the experimental effect size and sample size are modest, and the computational model is poorly validated and does not generate physiologically typical cystometric responses in simulations that are designed to recapitulate nominal LUT behavior.

      Despite the stated caveats about the small effect in the human study, it should be emphasized throughout that this result is most reasonably interpreted as showing the possibility that TNS can have a low-frequency excitatory effect that merits follow-up, rather than a conclusive demonstration. The effect size is small (as the authors note) and should be placed in context with some minimally clinically important difference, if possible. The result is statistically significant, but even this may be subject to revision due to the small sample and the effect of post-hoc outlier removal and data analysis choices.

      Acknowledged, the authors have included caveats in the discussion making clear that the present results should be interpreted as a proof of concept rather than a definitive demonstration. We note that in combination with existing animal findings these results strengthen the case for the existence of an unexplored excitatory effect of TTNS in human beings that may have valuable clinical implications if generalised.

      Given the apparent mismatch between the model and the cystometric behavior at the systems level in the "normal" case (e.g., low capacity, low voiding efficiency, omitted pressure profiles, frequency, etc.) and the absence of quantitative model validation (e.g., it was not compared directly with any experimental data from human urodynamics or rodent cystometry, beyond the initial fit to the neural data, no sensitivity analyses were performed, no goodness of fit computed, etc.) the discussion should be much more circumspect about interpreting the results at a systems level and should probably contain a paragraph explicitly detailing the limitations of the model. The subsequent interpretation should focus narrowly on the neural circuitry, rather than things like contraction duration, where the model is at its strongest. As written, the authors over-interpret what the in silico study can reasonably be used to infer about LUT function.

      The authors have reworded the discussion section, including a limitations paragraph containing caveats about the interpretation of the results. We make clear that a systemslevel perspective should be maintained and that futher research is required to validate and generalise these results.

      More justification is needed for why the contraction duration of the model is the central focus of analysis, when it connects only tentatively to the human study results, which focus on urgency. While not necessarily incorrect, a clearer link or motivation should be offered for how this informs our understanding of frequency-dependent TNS afferent or efferent inhibition during filling (which was the focus of the human studies and the abstract). In other words, why doesn't the model reproduce the 1Hz excitation effect of expediting void onset (or urgency in the human study), and why is it justified to look at contraction duration as a surrogate measure?

      The authors acknowledge this issue, and have included an additional section to the discussion considering the disparity between afferent and efferent effects observed across the pilot study and computational experimentation. The need for further research within this area to disentangle the complex nature of the frequency dependence has been stressed.

      The authors claim that "voiding behavior occurred earlier [at 1Hz stim in the model]", pointing to Figure 6A as evidence, but this panel appears to show a single example model run where 1Hz voiding occurs only ~1s earlier (display makes this very hard to estimate). This is insufficient evidence to support the claim. Later, it is stated that "TNS did not ... void much earlier". The claims should be made compatible, and all such claims should have reasonable supporting evidence.

      The authors have included additional information in the supplementary materials to support the claim.

      This information includes the bladder volume profile of a number of simulations under 0Hz and 1Hz conditions as well as the average void-onset time (i.e., simulated time before first void).

      There are a number of reporting concerns that can be easily addressed:

      (1) Human Study:

      (a) To interpret the human study analysis, a fuller description of the "optional 10m inute extension" is necessary. How were participants presented with this option, how was blinding preserved, what fraction of participants accepted, and did phase 1 results influence their decisions to continue?

      The authors have included additional clarification detailing how blinding was maintained during the washout period. Additionally, we have included a section in the results which details participation rates for the washout period. Given that only one participant declined participation in the washout period we do not believe it is necessary to conduct an analysis on what factors influenced participation.

      (b) For reproducibility, details about the TNS parameters should be articulated, such as the method of determining "motor thresholds" (unless this is synonymous with "urge to urinate"), the shape of the stimulation pulses (e.g., biphasic, charge balanced), typical applied current, etc.

      The authors have included the requested information and added two figures to the supplementary materials detailing the parameters of the equipment and the exact electrode placement used during the pilot study.

      (2) The Computational Model

      (a) The code availability statement for this type of work is inadequate. The model used for simulations in this work, as well as the code used to initialize (and randomize synaptic connections), needs to be hosted publicly because i) a model this intricate is extremely hard to reproduce/verify without code, ii) simulations are an essential piece of the argument, iii) hosting code requires very little overhead. Although there is an appropriate level of detail in the model description, it would not be possible to reproduce the model in any reasonable amount of time (or at all) because of the implementation-level details that are, understandably, omitted from the methods (e.g., what is a "unit", what 'exactly' do the connections in the PMC and PAG diagrams relate to, what were the final parameters used for all conductances, which parameters were "matched" to the original papers and which were not, etc.).

      The authors have included a link to a public GitHub repository where any interested individuals may download and use the code on their own machines for their own purposes. The repository, which includes a readme file detailing the operation of the model, as well as the thoroughly documented code provide the necessary transparency as suggested by the reviewers. We hope that by making the code open-source in this manner further research efforts by any interested researchers will be stimulated.

      (b) Critical cystometric/urodynamic values that are typically analyzed to assess healthy LUT function are detrusor pressure (timeseries) and/or post-void residual or voiding efficiency (scalars). These should be included to verify that the model is representative of the "normal" case. This is especially important because the model's "normal" behavior appears to have extremely low voiding efficiency (Figure 6A).

      The authors acknowledge this limitation and as such have modified the simulation files to calculate and return: detrusor pressure, post-void residual, bladder capacity, and voiding efficiency (calculated post-hoc from these values). It should be noted however, that implementing this change required that the computational results be re-run using the new code. As such, the exact details of Figure 5 now differ slightly (though the high-level results and implications remain unchanged).

      While the high-level results surrounding the frequency-dependence of TTNS and the likely brainstem specific cause of this effect remain unchanged, there were minor changes in the results of the computational projection experiments that necessitated a re-write of a portion of the results section.

      Additionally, the authors have added a section exploring the low-voiding efficiency of the model at baseline and potential explanatory factors.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      (1) In Figure 6Cii, the high frequency is labeled as 10 Hz, but it should be 20 Hz. The authors should correct this in the figure legend.

      Acknowledged, the typo has been corrected.

      Reviewer #2 (Recommendations for the authors):

      (1) Data and Analysis:

      (a) Greater detail on analysis exclusion is warranted. What does it mean to have "greater than normal water intake"? Why was a large "urge duration" grounds for exclusion? Was its threshold set post-hoc, which group was that participant from, and does its inclusion (or not) affect the results of the analysis substantially?

      The authors acknowledge the issue of data removal. As such, to address this limitation an alternative analysis was conducted. Rather than frequentist methods, a Bayesian modelling approach and post-hoc ROPE analysis was conducted which included a greater proportion of the dataset (excluding only those who did not undergo neuromodulation, or who directly met the exclusion criteria for the study). This approach was taken as bayesian methods are better suited for smaller sample sizes such as the one utilised in the present work. The ROPE analysis provides additional evidence for a real-world relevance of the effect on bladder function. Though the authors acknowledge that these results are preliminary they hope they will provide initial evidence for the translation of a novel effect of TTNS into human participants.

      (b) It is my understanding that Figure 4C is a plot of G1Hz and G20Hz on the horizontal from 4A and G1Hz and G20Hz on the vertical from 4B-"before". Hopefully, this is correct, and perhaps there is some way to state more simply what data are being reported, as it took me some time to understand.

      The authors confirm that figure 4C is a representation of data from figure panels A, and B. Thee horizontal axis represents the temporal “"urge onset” and the vertical axis the subjective intensity experienced at this point. To clarify this, the authors adjusted the axis labels to make clear the data being reported. Additional clarification was also added to the figure legend.

      (c) The choice of units in Figure 6 makes interpretation harder than it needs to be. Although not SI units, the field commonly reports volume in ml and duration in seconds or minutes (certainly not ms). The horizontal on Figure 6A is especially confusing, since sim cycles are not clearly defined, nor is the reason for the 20ms of them, or if the 1000s of total simulation time means compute-time or simulated time. Is Figure 6A (20ms/cyc)(50000cyc)(1s/1000ms)*(1min/60s) = 16.67 min of simulated time? If so, does the model show >6 voiding events in that time under normal conditions (which probably requires some explanation, since that is unusual)? Later (L216), other terminology of "simulation run" is introduced and further complicates the interpretation of how much simulated time is passing.

      Acknowledged, the authors have updated the units used in figures througout the publication to match standard SI notation (Fig 4: M<sup>3</sup> -> ml, Fig. 5A:M<sup>3</sup> -> ml, 20ms cycles -> seconds, ms->seconds). Authors have also updated the language used in the figure and the paper to make clear that the figure is referring to 500 seconds (16.67 mins) of simulated time.

      (d) It appears that in Figure 6B that a contraction duration of 0ms means no contraction at all - unclear if that is also true for everything below the horizontal dashed line.

      (e) Using p-values for analyzing differences between average model outputs (Figure 6C) is not appropriate, since one can run the model as many times as needed, making any negligible effect size statistically significant.

      The authors acknowledge that the computational nature of the second analysis limits the statistical tests that may be reasonably applied. As such, they have rewritten the results and discussion section to instead compare mean differences/effect sizes without reliance on p-values specifically.

      (2) Clarity and Presentation:

      (a) Figure 3 should be removed since it describes an experiment not conducted in this study and whose data was used only for model fitting, not an integral component of the model concept, analysis, or results. A short description and a paper reference are sufficient.

      The authors acknowledge this feedback and have removed Figure 3 from the publication. We have instead provided a reference and brief description of the data used to fit the parameters of the model.

      (b) L46, based on my understanding, should read something like "...may be a frequency dependent of TTNS, where low frequencies up-regulate bladder activity while higher frequencies downregulate it."

      Acknowledged, this section has been reworded to improve clarity.

      (c) Generally speaking, there is nothing "paradoxical" about a frequency-dependent response to e-stim, which happens throughout the nervous system and even in the LUT with pudendal sensory stimulation. "Surprising", "useful", "underexplored", etc., are all closer to the authors' meaning.

      Acknowledged, the authors have avoided the use of the term paradoxical to better represent the original intent of the research findings.

      (d) I am used to "washout" rather than "runoff", but this is a journal style decision, and either is fine.

      Acknowledged, the authors have replaced the use of the term runoff with washout and adjusted figure 1 to reflect this change.

      (e) L51 "analytically" is a mathematical keyword reserved for closed-form solutions, which is not what the authors actually refer to. Something like "computationally" or "in silico" is closer to their meaning.

      Acknowledged

      (f) L172 "abnormality" should be "non-normality".

      Acknowledged

      (g) L148 "Like the original model", presumably referring to Gorski?

      Correct, wording has been changed to make this clear.

      (h) L208-220 Unclear precisely what is meant by "intensity of the voiding events" or "temporal nature of the cycle".

      Acknowledged, the authors have provided additional clarification to avoid confusion.

      (i) Figure 6C Is "baseline" the nominal model without stimulation, while the "all connected" is the nominal model with stimulation? And all the rest of the conditions indicate what was cut in silico?

      Acknowledged, authors have reworded the figure legend to improve clarity.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      In this study, the authors set out to determine how two classes of kinase inhibitors, which stabilise a disease-relevant enzyme in either an active (Type I) or inactive state (Type II), influence its organisation and interactions with microtubule filaments in cells. Using the state-ofthe-art in-cell structural imaging approaches, they examine how these compounds affect the formation of protein filaments and their association with microtubules, and succeed in defining the underlying structural basis for these differences.

      A major strength of the work is the application of in-cell cryo-electron tomography combined with correlative imaging, which enables direct visualisation of protein organisation in a near-native cellular context. The data convincingly demonstrate that the Type I inhibitor compound stabilising the active state promotes extensive LRRK2 filament formation and microtubule bundling, whereas compounds stabilising the inactive state markedly reduce these interactions. The structural analysis further provides insight into how conformational states relate to filament organisation, including modelling of previously unresolved regions of the protein.

      These findings are internally consistent and align well with prior biochemical and structural studies, many of which were performed by the same team.

      There are, however, some limitations that should be noted. The experiments rely on overexpression of the I2020T mutant form of the LRRK2 protein, which is a rare variant, in a single cell type (293T cells), which may not fully reflect endogenous behaviour or wild-type LRRK2 in a physiological context. In addition, while the imaging data are compelling, the functional consequences of the observed filament formation and microtubule association remain unclear.

      The study therefore provides strong descriptive and structural insight, but more limited evidence linking these observations to cellular or disease-relevant outcomes.

      Overall, the authors largely achieve their aims, and the results support their central conclusion that different classes of kinase inhibitors have distinct effects on protein organisation in cells. The work represents an important advance in understanding how small molecules can reshape protein architecture in a cellular environment, with potential implications for therapeutic strategies. The methodological approach will also be of broad interest to the field, as it highlights the power of in-cell structural biology to study dynamic protein assemblies that are difficult to capture using traditional approaches.

      We thank the reviewer for their thoughtful and positive assessment of our work. We appreciate their recognition that in-cell cryo-electron tomography and correlative imaging provide a powerful approach for directly visualizing how small-molecule inhibitors reshape LRRK2 organization in a cellular environment.

      We agree that the use of overexpressed LRRK2I2020T in HEK293T cells represents an important limitation of the present study. This experimental system was selected because it enabled visualization and structural analysis of inhibitor-dependent LRRK2 assemblies in cells. However, the extent to which these observations apply to endogenous LRRK2, wild-type protein, other disease-associated variants, or physiologically relevant cell types remains to be established.

      We also agree that the functional consequences of inhibitor-dependent LRRK2 filament formation and microtubule association remain unresolved. The goal of the present study was to define how type I and type II kinase inhibitors alter the cellular organization and structural state of LRRK2. Our data demonstrate that these inhibitor classes have markedly different effects on LRRK2 filament formation and microtubule association in cells, and provide a structural framework for understanding these differences. Future studies will be required to determine how these assemblies influence LRRK2 signaling, microtubule-based processes, and diseaserelevant cellular phenotypes.

      We thank the reviewer for highlighting both the methodological significance of this work and its potential implications for understanding how therapeutic molecules remodel protein architecture in cells.

      Reviewer #2 (Public review):

      Summary:

      Mutations in Leucine-Rich Repeat Kinase 2 (LRRK2) are a major cause of Parkinson's disease. LRRK2 PD-related mutations all result in increased kinase activity. Therefore, LRRK2 has been the focus of the development of kinase inhibitors. So far, two classes of kinase inhibitors have been identified: type 1 LRRK2-specific inhibitors that stabilize LRRK2 in a closed active-like conformation and broad-range type 2 inhibitors that stabilize LRRK2 in an open inactive-like conformation. Basiashvili et al. used here in cell structural biology to study the effect of both type 1 and type 2 inhibitors on the localization and structural conformation of LRRK2-I2020T.

      Strengths:

      They showed that Type 1 and not Type 2 inhibitors induce LRRK2 filament/ on microtubules.

      Furthermore, they were able to build a structural map of full-length LRRK2 I2020T bound to a Type 1 inhibitor in a closed kinase confirmation. Together, this work thus confirms the data of previous studies that showed that LRRK2 Type 1 and 2 inhibitors differently affect filament formation.

      Weaknesses:

      All conclusions are fully supported by the provided data. However, as the authors indicated themselves, the physiological relevance of LRRK2 microtubule binding is questionable. Furthermore, although the authors used a full-length LRRK2 protein, like in previously published structures, the resolution of the N-terminal domains is rather poor. Therefore, it also remains unclear what we learn from this structure compared to the previously published structures.

      We thank the reviewer for their positive evaluation of our study and for recognizing that our conclusions are supported by the data.

      We agree that the physiological relevance of LRRK2 filament formation and microtubule association remains an important open question. Our study was designed to determine how type I and type II inhibitors affect the cellular organization and structural conformation of LRRK2. We explicitly acknowledge that future studies using endogenous LRRK2, disease-relevant cellular systems, and functional assays will be necessary to determine the biological significance of inhibitor-induced microtubule association.

      We also appreciate the reviewer’s comment regarding the resolution of the N-terminal domains. Although the N-terminal density does not support detailed atomic interpretation, its visualization provides information about the global organization of full-length LRRK2 within an inhibitorinduced, microtubule-associated assembly in cells. Importantly, our study does not claim highresolution structural determination of the N-terminal regions. Rather, the advance is the in-cell structural observation of full-length LRRK2<sup>I2020T</sup> in a type I inhibitor-stabilized, closed-kinase conformation, together with density indicating that the N-terminal repeat regions adopt an organization within the microtubule-associated lattice.

      We have revised the manuscript to clarify this point and to more carefully distinguish the structural information supported by the density from interpretations that would require higherresolution data.

      Reviewer #3 (Public review):

      Summary:

      This paper describes new insights into the effects of type-I and type-II LRRK2 inhibitors on HEK293T cells that over-express GFP-labeled LRRK2-I2020T. Using correlative light microscopy and cryo-electron tomography, a type-I inhibitor leads to the extensive decoration of microtubules with LRRK2, which is not seen for a type-II inhibitor. Subtomogram averaging reveals that LRRK2 binds to the microtubules in a closed-kinase conformation, with density for the N-terminal arms.

      Strengths:

      The paper is well written; the CLEM and cryo-ET appear to be done to a high standard. Consequently, I have only minor comments.

      Weaknesses:

      The resolution of the subtomogram averages is somewhat limited, but the authors have adequately limited the number of degrees of freedom in the fitting of their atomic models by only allowing rigid-body transformations of separate parts of LRRK2.

      The authors should include FSC curves between the rigid-body fitted atomic models and the various sub-tomogram average maps.

      We thank the reviewer for their positive assessment of the manuscript and for recognizing the quality of the correlative imaging and in-cell cryo-electron tomography analyses.

      We also appreciate the reviewer’s recognition that our interpretation of the maps was appropriately constrained by fitting domains as rigid bodies, rather than attempting unsupported high-resolution model refinement.

      We thank the reviewer for highlighting this and apologize for the oversight. We have added all the missing FSC curve plots of subtomogram maps presented in this study in Extended Data Figure 8.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      I think the current study is OK as it is, and the authors have taken this as far as they can.

      In future work, for either the authors or others in the field, it will be important to determine whether endogenous LRRK2 can be recruited to microtubules in response to compounds that stabilise the active state, particularly in cell types that are more relevant to Parkinson's disease. Does this cause a roadblock that impacts microtubule-driven transport? Establishing whether such recruitment occurs under physiological expression levels will be critical for assessing the broader relevance of the findings.

      In addition, it would be valuable to evaluate whether these Type 1 compounds have detrimental cellular effects linked to altered endogenous LRRK2-driven microtubule association, and whether inhibitors that stabilise the inactive state offer a potential advantage by avoiding this phenotype.

      We thank the reviewer for insightful recommendations for future studies.

      Reviewer #2 (Recommendations for the authors):

      (1) Figure 5: What is map C, and how is it different from the other maps? The authors indicate that the resolution of the N-terminal domains is moderate. How certain are the authors of the fit of these domains? Since map C is not provided in the supplemental, it is not possible to check this.

      We apologize for this oversight. We have updated the text to reflect how the map C was calculated. Now the text reads:

      “Additionally, we performed subtomogram analysis in Dynamo on a larger LRRK2<sup>IT</sup>decorated lattice that contained three layers of LRRK2<sup>IT</sup> density around the microtubule; we refer to this average as map C. Refinement was focused on the central four LRRK2<sup>IT</sup> subunits to better resolve additional protein densities within this larger lattice. In map C (Fig. 5A; Ext. Fig. 7).”

      In addition, we updated the figure 5D-F to demonstrate clear fit of the N-terminal domains into the presented map. We also added an Extended Data Figure 7 to the supplemental materials to highlight the fit of the model in the map and highlight the areas that would correspond to the Nterminal domains of LRRK2. We hope these updates demonstrate a good fit and justify observations highlighted in the paper.

      (2) The authors convincingly confirm that LRRK2 Type 1 and 2 inhibitors differently affect filament formation and that type 1 LRRK2-specific inhibitors stabilize LRRK2 in a closed activelike conformation. However, from the way the paper is written, it is unclear what we learn from this new structural data. How similar is the current structure compared to the previous structures? What is the novelty?

      We thank the reviewer for noting that this is unclear and giving us the opportunity to highlight it in the manuscript. We have added the following sentence in the discussion:

      “However, how the N-terminal repeats of LRRK2 are organized when the protein is in its closedkinase conformation remained unresolved. Stabilization of LRRK2 in a closed-kinase conformation by MLi-2 treatment and microtubule association reduces conformational heterogeneity to permit structure determination of full-length LRRK2<sup>IT</sup> with the N-terminal repeats undocked from the catalytic core. Therefore, the key novelty of this structure is that it captures full-length LRRK2<sup>IT</sup> in a cellular, microtubule-associated closed-kinase state and shows that kinase closure is compatible with an undocked N-terminal architecture. This distinguishes the in situ closed-kinase state from previously described in vitro intermediate active states.”

      Minor comments:

      (1) "Its C-terminal catalytic region is composed of WD40, Roc GTPase, Kinase and COR (RCKW) domains."

      Suggest changing this to Roc GTPase, Cor, Kinase and WD40 (RCKW) domains for clarity/following of abbreviation.

      We have made this change.

      (2) "In the MLi-2 treated cells, LRRK2IT strands were organized around microtubules with a regularly spaced lattice, similar to the LRRK2IT strands in cells not treated without the inhibitor (Fig. 3A-E)"

      Phrasing, correct the underlined portion.

      We have made this change.

      (3) "While average pitch. rise, and handedness of the filaments of the rate GZD-824 treated LRRK2 filaments were similar..."

      Punctuation.

      We have made this change.

      (4) "Our results clarify the relationship between kinase conformation, repeat undocking, and microtubule association. Increased microtubule association observed for I2020T mutant favors repeat undocking, a prerequisite for kinase closure and filament assembly"

      Do the authors mean undocking by the N-terminal repeats or repeatedly undocking of these domains?

      We meant undocking of the domains, and have corrected the sentence to clarify this.

      (5) "Together, these findings provide a structural view of full-length LRRK2 in a closed kinaseconformation and capture a resolved snapshot along its conformational continuum"

      Needs a space.

      We have made this change, and thank the reviewer for pointing it out.

      (6) "Microtubule decoration by LRRK2IT has not been studied in cell types that endogenously express high levels of LRRK2, such as lung epithelial cells and brain-resident immune cells including microglia and macrophages44. Thus, it remains possible that aberrant LRRK2microtubule interactions occur under physiological expression conditions, potentially disrupting homeostatic intracellular transport and being further exacerbated by type I LRRK2 inhibitors, as suggested by in vitro studies23,45."

      Many studies have studied the localization of endogenous LRRK2, however were not able to detect filament localization on microtubules. Moreover, to my knowledge, there is also no clear evidence that type 1 inhibitors disrupt microtubule transport in cells expressing endogenous levels of LRRK2.

      Therefore, I suggest to rephrase or remove this paragraph.

      We agree that the current evidence does not establish that this occurs broadly in cells. However, to our knowledge, cells or tissues with high endogenous LRRK2 expression have not yet been systematically examined in this context. We therefore present sparse decoration of hyperactive LRRK2 on microtubules as a possibility rather than a strong conclusion. We have also previously shown that type I inhibitors disrupt microtubule transport in vitro, but determining whether a similar effect occurs in cells is ongoing work and beyond the scope of the present manuscript.

      Reviewer #3 (Recommendations for the authors):

      (1) P4: The first section of the Results refers to LRRK2 localising to microtubules in the presence of the type-I compounds, and to the cytosol with the type-II inhibitor. Aren't microtubules in the cytosol also?

      We meant cytosolic LRRK2, we have revised the text to reflect this. It now reads:

      In cells treated with MLi-2, we observed LRRK2<sup>IT</sup> in extended filaments, puncta, and diffuse in the cytosol (Fig. 1D-E; Ext. Fig 1A-D). In contrast, when cells were treated with GZD-824, LRRK2<sup>IT</sup> was mostly localized to puncta and distributed throughout the cytosol, with reduced filament formation (Fig. 1F-G; Ext. Fig 1E-H), in agreement with our previous work [23,24,40].

      (2) P4: second column, halfway down. I don't understand how the 16 and 8 neighbours are derived from Figure 3J-K. Perhaps indicate this in the figure?

      Thank you for bringing this to our attention. We have added an Extended Data Figure 5 to clarify this point. The Extended data figure 5 highlights and annotates the immediate neighboring LRRK2 densities in the MLi-2- and GZD-824-treated lattices, making clear how the 16 and 8 nearest-neighbor values were assigned from the observed lattice organization.

      (3) P6: first column, halfway down: perhaps make it explicit that only rigid-body fitting was performed because of the limited resolution?

      We have incorporated this useful suggestion. The text now reads:

      “We split this model in three parts: the WD40 and C-lobe of the kinase, the N-lobe of the kinase with ROC and COR domains, and the LRR and ANK domains, aligned and fitted each of these three to our map A (Fig. 4D-F). Given the limited resolution of the map A, we fit the model as three rigid bodies without atomic refinement.”

      (4) P6: same column near the bottom: what is map C? and how was it calculated? Also, it is not clear to me from Figures 5D-F whether the statement "clearly correspond to the LRR-ANK-ARM domains" is justified by the map. From Figure 5D-F, I see a rather poor fit in a low-resolution map. This needs to be toned down or better illustrated.

      We apologize for the oversight. We have updated the text to clarify how the map C was calculated. Now the text reads:

      “Additionally, we performed subtomogram analysis in Dynamo on a larger LRRK2<sup>IT</sup>decorated lattice that contained three layers of LRRK2<sup>IT</sup> density around the microtubule; we refer to this average as map C. Refinement was focused on the central four LRRK2<sup>IT</sup> subunits to better resolve additional protein densities within this larger lattice. In map C (Fig. 5A; Ext. Fig. 7).”

      In addition, we updated the figure 5D-F to better demonstrate the fit of the N-terminal domains into the presented map. We also added an Extended Data Figure 7 to the supplemental materials to further highlight the fit within the map and indicate the areas that correspond to the N-terminal domains of LRRK2. We hope these updates clarify how map C was calculated and better illustrate our interpretation of the additional densities.

    1. Author response:

      We would like to thank the reviewers for their careful analysis of our manuscript. We appreciate their insightful suggestions for improvement. We intend to address each of their comments in our revision, with the major points outlined below.

      (1) Reviewers 1 and 2 both highlighted the importance of the specificity of our genetic and optogenetic manipulations in the interpretation of our results. We agree that this point is essential. We will expand our discussion to incorporate more references demonstrating the specificity of our genetic approach, the networks engaged, and potential caveats.

      (2) We acknowledge the importance of validating the dystonic nature of our model as noted by Reviewer 2 and the value of more objective quantification of dystonic crisis as requested by Reviewer 1. We will discuss the potential as well as the difficulty of developing this kind of classification due to the non-stereotypic nature of dystonic movements and the lack of objective, measurable definitions even in clinical settings.

      (3) Reviewers 1 and 2 also requested additional discussion of the role of the iCNN to CL thalamus projection in driving dystonic crisis. We will clarify our claims on this point to more accurately reflect what we can confidently interpret from our current experiments and discuss the value of further experiments in the future.

      (4) We agree with Reviewers 1 and 2 that the effects of repeated stimulation are intriguing and deserve further investigation in the future. We will expand our discussion of this point to provide additional context and describe potential mechanisms that could explain our observed results, which may be tested in further studies.

      (5) Reviewer 1 noted that the clinical dataset could be discussed in more detail to support the translational relevance of our findings. We will provide additional information on the characteristics of our patient sample and potential confounding variables.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The manuscript by Yang et al. investigates the relationship between multi-unit activity in the locus coeruleus, putatively noradrenergic locus coeruleus, hippocampus (HP) sharp-wave ripples (SWR) and spindles using multi-site electrophysiology in freely behaving male rats. The study focuses on SWR during quiet wake and non-REM sleep, and their relation to cortical states (identified using EEG recordings in frontal areas) and LC units.

      The manuscript highlights differential modulation of LC units as a function of HP-cortical communication during wake and sleep. They establish that ripples and LC units are inversely correlated to levels of arousal: wake, i.e. higher arousal correlates with higher LC unit activity and lower ripple rates. The authors show that LC neuron activity is strongly inhibited just before SWR detected during wake. During non-REM sleep, they distinguish "isolated" ripples from SWR coupled to spindles and show that inhibition of LC neuron activity is absent before spindle-coupled ripples but not before isolated ripples, suggesting a mechanism where noradrenaline (NA) tone is modulated by HP-cortical coupling. This result has interesting implications for the roles of noradrenaline in the modulation of sleep-dependent memory consolidation, as ripple-spindle coupling is a mechanism favoring consolidation. The authors further show that NA neuronal activity is downregulated before spindles.

      Strengths:

      In continuity with previous work from the laboratory, this work expands our understanding of the activity of neuromodulatory systems in relation to vigilance states and brain oscillations, an area of research that is timely and impactful. The manuscript presents strong results suggesting that NA tone varies differentially depending on coupling of HP SWR with cortical spindles. The authors place their findings back in the context of identified roles of HP ripples and coupling to cortical oscillations for memory formation in a very interesting discussion. The distinction of LC neuron activity between awake, ripple-spindle coupled events and isolated ripples is an exciting result and its relation to arousal and memory opens fascinating lines of research.

      Weaknesses:

      I regretted that the paper fell short of trying to push this line of idea a bit further, for example by contrasting in the same rats the LC unit-HP ripple coupling during exploration of a highly familiar context (as seemingly was the case in their study) versus a novel context, which would increase arousal and trigger memory-related mechanisms. Any kind of manipulation of arousal levels and investigation of the impact on awake vs non-REM sleep LC-HP ripple coordination would considerably strengthen the scope of the study.

      Comments on revised version:

      The authors have added methodological details to the results section after the first round of reviews, improving the manuscript readability. Some points might still be improved, for example, the authors use a delta/gamma ratio to track cortical states for example, but there is no methods section corresponding to this metric. Authors write that higher SI corresponds to a lower arousal state that is associated with "more synchronized cortical population activity, higher ripple rate and reduced LC neurons firing" but there are no references or analysis to support this statement, only examples showing changes in SI over a few minutes.

      We thank Reviewer #1 for the positive evaluation of our study and for highlighting its strengths and potential avenues for future investigation.

      We have specified in the Methods the calculation of SI as a delta/gamma ratio and provided the frequency ranges used for each band: “Artefact-free EEG signals were band-pass filtered using a Butterworth filter implemented in Matlab 2024a (MathWorks, Natick, MA). Subsequently, deltaband power (δ, 1–4 Hz), theta-band power (θ, 6–10 Hz), and the θ/δ power ratio were computed within contiguous 4-second epochs.”

      We agree with the reviewer and have acknowledged in the Discussion that incorporating behavioral assays will be essential for achieving a mechanistic understanding of the observed network dynamics and their functional role in memory consolidation. Such experiments are beyond the scope of the present study but represent an important direction for future research. We have also revised the Discussion to avoid overstated claims and to ensure that our interpretation remains appropriately supported by the current data. Discussion (last paragraph): “Conducting behavioral assays before electrophysiological recordings, along with spatially and temporally precise modulation of LC activity during recording sessions, will be essential for achieving a mechanistic understanding of network dynamics and its functional role for memory consolidation in future investigations.”

      Reviewer #2 (Public review):

      Summary:

      In this study, authors studied the synchrony between ripple events in Hippocampus, cortical spindles and Locus Coeruleus spiking. The results in this study together with the established literature on the relationship of hippocampal ripples with widespread thalamic and cortical waves, guided authors to propose a role for Locus Coeruleus spiking patterns in memory consolidation. The findings provided here, i.e. correlations between LC spiking activity and Hippocampal ripples, could provide basis for future studies probing the directional flow or the necessity of these correlations in the memory consolidation process. Hence, the paper provides enough scientific advance to highlight the elusive yet important role of Norepinephrine circuitry in the memory processes.

      Strengths:

      Authors were able to demonstrate correlations of Locus Coeruleus spikes with hippocampal ripples as well as with cortical spindles. Specific strength of the paper is in the demonstration that the spindles that activate with the ripples are comparatively different in their correlations with Locus Coeruleus than those which do not.

      Weaknesses:

      The claims regarding the roles of these specific interactions were mostly derived from the literature that these processes individually contribute to the memory process, without any evidence of these specific interactions being necessary for memory processes. There are also issues with the description of methods, validation of shuffling procedures and unclear presentation and the interpretation of the findings, which are described in points that follow. I believe addressing these weaknesses might improve and add to the strength of the findings.

      Comments on revised version:

      The authors addressed all of my major concerns during the revision. As a result, the study now provides convincing evidence as well as improved presentation of results, that makes this manuscript important to the broader field of neuroscience, beyond the specific sub-field.

      We thank Reviewer #2 for the positive assessment of our work and for recognizing both its strengths and its potential to stimulate future research in this area. We agree that assessing memory function is essential for understanding how noradrenergic signalling influences the network mechanisms underlying memory consolidation. While such experiments are beyond the scope of the present study, we acknowledge this important limitation in the Discussion and identify it as a key direction for future research. Discussion (last paragraph): “Conducting behavioral assays before electrophysiological recordings, along with spatially and temporally precise modulation of LC activity during recording sessions, will be essential for achieving a mechanistic understanding of network dynamics and its functional role for memory consolidation in future investigations.”

      We added more details in the Methods and expanded the Figure 4 legend to improve the results presentation.

      Reviewer #3 (Public review):

      This manuscript examines how locus coeruleus (LC) activity relates to hippocampal ripple events across behavioral states in freely moving rats. Using multi-site electrophysiological recordings, the authors report that LC activity is suppressed prior to ripple events, with the magnitude of suppression depending on ripple subtype. Suppression is stronger during wakefulness than during NREM sleep and least pronounced for ripples coupled to spindles.

      The study is technically sound and addresses a timely and important question regarding how LC activity interacts with hippocampal and thalamocortical network events across vigilance states. While the findings are interesting, they remain observational in nature. Following revision, the manuscript has substantially improved in both presentation and interpretation of the results, and most concerns have been addressed satisfactorily. I therefore only have a few minor considerations that the authors may wish to explore further in the current study or in future work, as these directions could provide additional mechanistic insight and would likely be of considerable interest to the field.

      The authors demonstrate clearly that tonic LC firing rates preceding ripples differ significantly between wake-associated ripples (highest LC firing), isolated ripples during NREM sleep (lower LC firing), and spindle-coupled ripples (lowest LC firing). They also appropriately note that baseline firing differences will naturally influence the magnitude of LC suppression, which they also observe (highest LC reduction for wake ripples, then isolated ripples and last spindle-coupled ripples).

      However, this aspect could be explored further, as it may provide additional insight into the regulation of spindle-associated ripple events. Since LC activity appears to decline gradually prior to ripple occurrence (Suppl. Figure 2), it would be interesting to test whether this gradual reduction helps organize the emergence of isolated versus spindle-coupled ripples. For example, isolated ripples may occur during the initial phase of LC decline, whereas spindle-coupled ripples may preferentially emerge when LC activity reaches its lowest levels. Such a relationship could also be consistent with the stronger synchronization observed for spindle-ripple coupling.

      Related to this point, it would also be informative to examine whether isolated spindles occur more randomly in time, whereas spindle-associated ripple events appear more temporally clustered. If a single isolated spindle occurs, the associated LC suppression might be more pronounced. In contrast, when multiple spindle-associated ripple events occur in succession, LC activity may already be reduced following the first event, resulting in smaller additional suppression preceding subsequent events. Exploring this possibility could help clarify how LC dynamics shape the temporal emergence of ripple-subtypes

      We are grateful to Reviewer #3 for the positive evaluation of our manuscript and for the constructive comments highlighting the significance of our findings and their implications for future studies. We agree that a more comprehensive investigation of cross-regional coupling and its modulation by the LC–NE system represents an important and still insufficiently explored area of research. Further elucidating the complexity of these interactions will be essential for understanding how noradrenergic signalling shapes large-scale brain network dynamics across behavioral states. We acknowledge it in the Discussion: “A more comprehensive investigation of cross-regional coupling and its modulation by the LC–NE system represents an important and still insufficiently explored area of research. Further elucidating the complexity of these interactions will be essential for understanding how noradrenergic signaling shapes large-scale brain network dynamics across behavioral states.”

      Recommendations for the authors:

      Reviewer #3 (Recommendations for the authors):

      Figure 4: It would be helpful to show the unshuffled data at the front (it is hidden partly behind the unshuffled data). Also, the unshuffled data are not introduced in the text for this figure. Would be helpful. Please also add color bars to improve interpretability.

      To improve readability and facilitate interpretation, we revised Figure 4. Specifically, we 1) reordered the plots to present the unshuffled (ripple) data at the front; 2) expanded the figure legend to provide a more detailed description of the shuffling procedure; and 3) removed the unnecessary color fill from the box plots in panels B and C, while retaining the labels.

      Figure 7: The color coding appears wrong in panel F (mean curves in F do not correspond to time traces in G). This should be checked and corrected if necessary.

      We have corrected the colour coding in Figure 7.

    1. Author response:

      We thank the reviewers for their thoughtful and constructive comments, and we plan to implement many of their suggestions to improve the paper. We agree that the manuscript would benefit from a clearer and more evidence-based presentation of how feedback responses relate to subsequent learning responses. To address this point, we will perform additional analyses and modeling, including model-free analyses of the phasic and tonic components. These analyses will allow us to test whether the tonic component remains the dominant predictor of the learning response without relying on the specific assumptions of the tonic/phasic decomposition model.

      We also agree that the manuscript would benefit from a more detailed discussion of the mechanisms that may shape the temporal evolution of feedback responses and their relationship to subsequent learning. We will therefore expand the discussion of this issue and relate our findings to adaptive feedback control and continuous-time models of motor adaptation, which may provide useful frameworks for interpreting the relationship between feedback responses and learning responses.

      Finally, we agree that the scope and limitations of the current experimental paradigm should be discussed more explicitly when considering the generality of our findings. We will therefore discuss whether and how the present results may generalize to broader forms of sensorimotor learning and adaptation. We will also

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This valuable study analyses correlations between traits of Chinese frog species and their Red List status, finding differences between adults and larvae and thus pointing to the importance of considering different life-cycle stages in this and possibly other animal groups when assessing species extinction risks. The current study is, however, incomplete because of unclear threat categories for tadpoles, the omission of other key species traits, and insufficient statistical analysis.

      Thank you very much. We have revised the manuscript according to the reviewers' comments. The parts highlighted in red in the manuscript are the revised portions.

      Public Reviews:

      Reviewer #1 (Public review):

      The manuscript shows that different traits of adults and larvae correlate with Red List status. The authors argue that this shows a big gap in the conservation of amphibians and that the traits of all life stages should be taken into account in amphibian conservation. Specifically, amphibian conservation should do more for the habitats where the larvae live.

      The manuscript is well written and easy to understand. The methods are sound.

      While the study will make an interesting contribution to conservation science, there are many things that I disagree with.

      (1) I don't think that amphibian larvae and their requirements are a "blind spot" as the title suggests. When reading the manuscript, I didn't learn how conservation practice should change in response to the results.

      Thank you very much for your suggestions. The description of the 'blind spot' was inappropriate, and we have revised it. Investigating the relationship between life history traits and threat status can help us understand which species are more vulnerable to extinction. Furthermore, we can predict the potential threat severity of species that have not yet been assessed. Because we still lack knowledge about the biodiversity of many taxonomic groups. For example, as of early 2024, over 34% of Chinese anuran species have been described in the last ten years, and 100 - 200 new species are still being discovered globally each year. Under these circumstances, given the current investment in biodiversity conservation, it is nearly impossible to assess the threat status of every species and develop conservation strategies. Therefore, predicting the threat status of species is very important for biodiversity conservation, as it will provide support for the subsequent formulation of specific conservation policies. Among the already described animals species, most have complex life history cycles. Moreover, species face threats not only at the adult stage; those with certain traits at other life stages may also be vulnerable to threats. For example, our study takes amphibians as an example and shows that groups with larger body sizes at the tadpole stage may face more serious threats.

      (2) I wonder whether the relationship between species traits and extinction risk is of great importance for conservation. If a species is Data Deficient on the IUCN Red List, then species traits could be used to predict its Red List category. However, for other conservation projects, I don't see how this would work. How would traits be linked to captive breeding, conservation translocation, pond construction or habitat management in general? In some cases, I can envision a link between species traits and pond hydroperiod.

      Thank you very much for your suggestions. Understanding the relationship between traits and threat status is of great importance for the conservation policies and the allocation of conservation resources, especially when conservation resources are insufficient. As mentioned earlier, the current conservation resources are insufficient to support us in surveying and assessing every Data Deficient (DD) species, not to mention the large number of new species being discovered each year. By predicting threat status, we can identify which groups or species should be prioritized for research, such as population size and distribution range surveys, so that specific conservation strategies can subsequently be developed.

      (3) Species traits are body size and morphological traits. That makes sense. However, one of the species traits was microhabitat. I find it far-fetched to call habitat a species trait. This is standard habitat ecology. It is well known that habitats matter and that different habitat types face different threats, and consequently, the species that live in those habitats. Furthermore, habitat and morphology may be confounded. For example, tadpoles in lentic and lotic habitats have very different morphologies. So is it habitat or morphology?

      Thank you very much for your suggestions. The type of habitat in which a species lives affects the threats it faces. In many studies on the relationship between extinction risk and traits, microhabitat or habitat type is widely used as a predictive variable. For example, in studies on Squamata, whether a species is distributed on islands or peninsulas has also been included as a trait. Following your suggestion, we have revised the sentences to refer to 'morphological traits and microhabitat information'. Many morphological traits of species are related to habitat selection, but not all traits associated with habitat selection have been measured or have sufficient data. Therefore, it is necessary to include microhabitat type as an independent variable. Additionally, we calculated the Variance Inflation Factor (VIF) prior to the regression analysis to ensure that the analysis was not affected by multicollinearity.

      (4) I don't know how the threat status of Chinese amphibians is determined. IUCN has multiple reasons why a species can be Red Listed. One reason is range size, and another reason is population decline. Personally, I don't think they should be pooled in an analysis because they are fundamentally different reasons why a species has a high extinction risk. A reduction in population size of greater than 30% in 10 years or 3 generations is not the same thing as a small distribution range. Another issue is that IUCN developed the Green Status of species. The Green Status shows that even a species which is LC on the Red List may be significantly depleted.

      Thank you very much for your valuable suggestions. The assessment method of the China Biodiversity Red List is the same as that of the IUCN Red List, both of which are based on population size and area of distribution. We fully agree with your point that analyses should be conducted according to specific threat types. Unfortunately, the full report of the latest version of the China Biodiversity Red List, released in 2023, has still not been published. Therefore, we were unable to perform the relevant analyses.

      (5) The species traits in Table 1 are mostly functional/morphological and body size related (and microhabitat). While there may be correlations between traits and Red List status, it is unknown whether this is correlation or causation. In addition, it is difficult to know the conservation interventions that may be necessary now that we know that relative head with and Red List status are correlated.

      Thank you for pointing out the important distinction between correlation and causation. Your comment is very insightful, and we have revised our manuscript to further clarify the scope and limitations of our study. The aim of our study is to identify which traits show statistical associations with extinction risk, thereby providing testable hypotheses for future research. We acknowledge that the mechanisms underlying the associations between certain morphological traits (e.g., head length, tympanum diameter) and extinction risk remain unclear, and these findings cannot yet be directly translated into well-established management measures. Nevertheless, the value of our study lies precisely in generating hypotheses about traits that warrant prioritized investigation of their causal mechanisms, as well as offering clues for the initial allocation of conservation resources. Following your suggestion, we have discussed the limitations of the study in the Discussion section of the manuscript.

      (6) In the discussion, the authors explain why body size and other traits may affect extinction risk and whether there is a causal relationship. I agree that body size may have a direct effect because larger species are harvested more frequently (it was interesting to learn that tadpoles are harvested as well). However, as macroecological studies show, smaller species often have larger populations than larger species. Abundance may matter.

      Thank you very much for your suggestion. Following your advice, we have revised the discussion section regarding body size.

      (7) I found it much harder to understand why relative head length and tympanum size correlated with Red List status. I wasn't convinced by the arguments in the discussion. Typanum size may be related to hearing and anthropogenic noise. Several studies are cited which show that frogs alter their calling behaviour in response to noise. Crucially, however, they describe changes in behaviour or properties of the advertisement call, yet none show that noise has effects on population viability. If some anthropogenic stressor affects individuals, then this does not mean that it will cause a population decline. When IUCN published the second global amphibian assessment, did they list noise as a major threat to amphibians?

      We appreciate your insightful comments and fully agree with your assessment. Indeed, the hypothesis that noise threatened anuran amphibians lacks direct evidence. While relevant studies indicate that anthropogenic noise causes auditory masking in anurans and reduces individual reproductive success, the IUCN has not listed noise as a primary threat to amphibians. Although acoustic communication is vital for amphibian reproduction and is susceptible to noise interference, there is currently no definitive evidence proving that noise extensively impacts amphibian survival. Therefore, in the revised manuscript, we retained it as a hypothesis to be tested and explicitly clarified that current evidence is limited to behavioral changes. Regarding the correlation with relative head length, we acknowledge that the underlying mechanism remains unclear; it may stem from phylogenetic signal residuals or unidentified ecological factors (such as diet or locomotor ability). In the Discussion, we revised this part as a correlation requiring further investigation.

      (8) There are statements that the tadpole stage is the most important stage: "a critical period for amphibian survival" (line 78-79). While there is high mortality in the tadpole stage, tadpole survival is rather unlikely to affect population survival. Many population models show this. See, for example, Biek et al. 2002 in Conservation Biology. Other papers have argued that the postmetamorphic juvenile stage is most important (Petrovan and Schmidt 2009 Biological Conservation).

      We greatly appreciate your comment. We agree that the original statement was overly absolute. The most critical life stage for population persistence can differ across species, and many studies have shown that other stages may be more important. Accordingly, we have revised this sentence as you suggested.

      (9) The authors repeatedly make the statement that amphibian conservation should focus more on the tadpole stage. I don't understand why this statement is made. For example, a major activity in amphibian conservation is the restoration and de novo construction of ponds (see Calhoun et al. 2014 PNAS, Moor et al. 2022 PNAS). Ponds are habitats for tadpoles. Others removed fish from amphibian breeding sites because fish prey on tadpoles (and adults; see Vredenburg 2004 PNAS). Semlitsch (2002 in Conservation Biology) argued that the management of pond hydroperiod is a critical element of amphibian recovery plans. Ponds should be temporary because this effectively removes predators that consume tadpoles. Clearly, the tadpole stage is not a neglected stage in amphibian conservation.

      Thank you for pointing this out. The literature you cited (Calhoun et al., 2014; Moor et al., 2022; Vredenburg, 2004; Semlitsch, 2002) convincingly demonstrates that the tadpole stage has received a certain degree of attention in amphibian conservation practice. Our original statement was indeed problematic. What we intended to convey is that information on the tadpole stage needs to be integrated into conservation assessment frameworks and conservation planning. For example, many studies on the relationship between functional traits and threat extent have not included tadpole-related information. Compared with our knowledge of adult amphibians, we know far less about tadpoles, and for many species, information on the tadpole stage is entirely lacking. Therefore, we call for tadpoles to receive greater attention in future research relative to the current situation.

      Recommendations for the authors:

      Reviewing Editor Comments:

      Conceptual problems:

      (1) Many conservation measures for amphibians target larvae; thus, globally, this is not a blind spot. If this is different in China, it would be important to point this out.

      We thank the reviewer for the thoughtful comment. We recognize that the tadpole stage has indeed received attention in amphibian conservation practice, and our original statement was therefore imprecise. Our intended argument was that tadpole-stage information should be integrated into conservation assessment frameworks and conservation planning. For instance, many studies examining the relationships between functional traits and threat extent have failed to include data on tadpoles. Our understanding of tadpoles remains far more limited than that of adult amphibians, and for a large number of species, no information on the tadpole stage is available. Consequently, we advocate for substantially greater research attention to tadpoles than they currently receive. We have revised the text accordingly.

      (2) While traits may be used to predict Red-List status, it is not clear how they could inform conservation measures. This should be discussed.

      Thank you for your comment. The aim of our study is to identify which traits show statistical associations with extinction risk, thereby providing testable hypotheses for future research. We acknowledge that the mechanisms underlying the associations between certain morphological traits (e.g., head length, tympanum diameter) and extinction risk remain unclear, and these findings cannot yet be directly translated into well-established management measures. Nevertheless, the value of our study lies precisely in generating hypotheses about traits that warrant prioritized investigation of their causal mechanisms, as well as offering clues for the initial allocation of conservation resources. Following your suggestion, we have discussed the limitations of the study in the conclusion section of the manuscript.

      (3) The Red-List categories may not be appropriate to link traits to extinction risk. It would be important to explain how these are defined for China and how this may affect the analysis (e.g. linking larval traits to larval extinction risks would be difficult if Red-List criteria do not consider larvae).

      Thank you very much for your suggestions. The assessment method of the China Biodiversity Red List is the same as that of the IUCN Red List, both of which are based on population size and area of distribution. The assessment process is independent of species' morphological traits. Consequently, analyzing correlations between traits and Red List categories does not constitute circular reasoning or contain any inherent logical contradiction. On the contrary, it is precisely because the two are independent that statistically significant associations between traits and extinction risk can have predictive value and inform conservation actions. In the revised manuscript, we clarified the independence of Red List assessments and rephrase any potentially misleading wording (e.g., changing "threat category of tadpoles" to "threat category of the species (assessed based on adults)").

      Methodological problems:

      (4) Choice of traits. Are morphological traits sufficient (add e.g. fecundity)? Justify the use of habitat traits (also, if additional ones would be included: geographic and altitudinal ranges, habitat specificity).

      Thank you for your suggestion. We fully agree that traits such as geographic range, elevational range, fecundity, and habitat specificity have important effects on extinction risk. The core objective of this study is to compare the stage-specific differences in the associations between extinction risk and morphological and microhabitat traits of adults versus tadpoles. Moreover, spatial traits such as geographic range are inherently highly correlated with the threat status of species, and including them might mask life-stage-specific signals. We will acknowledge this limitation in the discussion and identify the above-mentioned traits as important directions for future research.

      (5) Model choice: models have high uncertainty, thus better use model averaging and AICc instead of AIC. Overall, the statistical analysis and model selection procedure are poorly described; only summary results are presented.

      We greatly appreciate the reviewer's suggestion. Accordingly, we re-analyzed the data following your advice. In addition, the description of the methods has been supplemented.

      (6) Caveats: the data only allow for correlational analysis; causation cannot be derived from observational data. Furthermore, with a limited number of species, the number of predictors should not be too large.

      Thank you for your suggestion. Studying the relationship between traits and species threat status is important in conservation biology. Although such studies can only reveal statistical associations between traits and extinction risk rather than infer causality, they can generate hypotheses to facilitate future research. Additionally, this type of study can help predict the threat severity of unevaluated species, which is highly valuable for developing biodiversity conservation plans. In this study, 299 species were included in the analysis, and nine predictor variables (eight morphological traits plus one microhabitat type) were used. The ratio of sample size to number of variables was approximately 33:1, and variance inflation factor (VIF) tests indicated that multicollinearity was within an acceptable range (VIF < 5). Therefore, the risk of model overfitting is low. We will add this clarification in the revised manuscript.

      Reviewer #2 (Recommendations for the authors):

      (1) My first major concern is the species threat categories for tadpoles. The authors obtained the extinction risk data from the China Biodiversity Red List or IUCN. However, the assessment of threat categories, whether by the China Biodiversity Red List or IUCN, is based solely on adults. That means that the threat categories for both adults and tadpoles are the same, which can be seen in Figure 1. Since there is no specific assessment of threat categories for tadpoles, I have concerns about whether it is reasonable to relate species traits of tadpoles to the extinction risk for adults. I think it is one of the reasons why there is no study examining the association between functional traits and extinction risk in tadpole stages.

      We thank the reviewer for raising this important point, as it addresses a key prerequisite issue. The Red List assessment evaluates species, not individual life stages. The threat categories of both the IUCN and China Biodiversity Red Lists are determined based on criteria such as population size and geographic range of the species. The assessment process is independent of species' morphological traits. Consequently, analyzing correlations between traits and Red List categories does not constitute circular reasoning or contain any inherent logical contradiction. On the contrary, statistically significant associations between traits and extinction risk can have predictive value and inform conservation actions. In the revised manuscript, we will explicitly clarify the independence of Red List assessments and rephrase any potentially misleading wording (e.g., changing "threat category of tadpoles" to "threat category of the species (assessed based on adults)").

      (2) My second major concern is about the Data Analysis. The authors built and compared three types of models, i.e., PGLS_BM, PGLS_OU, and GLS_no_phylogeny. They claim that the OU-based PGLS model provided the best fit for both adult and tadpole datasets. Although the result seems reasonable, it is not clear how the OU-based PGLS model was obtained and what it exactly means. It seems to be a full model including all the predictor variables. However, since eight morphological traits and one microhabitat data of both adults and tadpoles were collected, there should be 29-1=511 candidate models. Unless the best model has an Akaike weight (wi) > 0.90 in all the OU-based PGLS models, it has substantial model selection uncertainty. If this is the case, the model average should be used, and weighted estimates of regression coefficients and unconditional standard errors that incorporate model selection uncertainty are better statistical methods (Burnham & Anderson, 2002).

      Thank you very much for your suggestion. Species' traits are related to evolutionary relationships, with more closely related species tending to be more similar. In the original manuscript, the three models we compared (PGLS_BM, PGLS_OU, GLS_no_phylogeny) were intended to select the optimal evolutionary covariance structure. Since we were more interested in the differences between adults and tadpoles, after selecting the OU structure, we actually used a single full model that included all traits to estimate the regression coefficients for each factor. Following your advice, we have added a model averaging analysis and revised the manuscript accordingly.

      (3) In addition, the Second-Order Information Criterion AICc, but not AIC, should be used for model selection. You have at least 9 variables (eight morphological traits and one microhabitat data) or 11/13 variables for the parameter estimates (Table 1). However, you have only 299 species included in the analysis (n = 299), which is relatively small compared to the number of variables (n/k << 40). Therefore, the AIC corrected for small sample size (AICc) should be used.

      We greatly appreciate the reviewer's suggestion. Accordingly, we re-analyzed the data following your advice.

      (4) Previous studies found that amphibian species with large body size, restricted geographic and elevational ranges, low fecundity or high habitat specificity are frequently predicted to have higher extinction risk (Cooper et al., 2008; Sodhi et al., 2008; Botts et al., 2013; Lips et al., 2003; Murray & Hose, 2005). The authors only included morphological traits and one microhabitat data point in the analyses. I wonder whether they can collect more trait data associated with extinction risk, such as geographic and elevational ranges, fecundity traits, or diet/habitat specificity, so as to gain more insight into the study.

      Thank you for your suggestion. We fully agree that traits such as geographic range, elevational range, fecundity, and habitat specificity have important effects on extinction risk. The object of this study is to compare the stage-specific differences in the associations between extinction risk and morphological and microhabitat traits of adults versus tadpoles. Moreover, spatial traits such as geographic range are inherently highly correlated with the threat status of species, and including them might mask life-stage-specific signals. In the Methods, we acknowledge this limitation and identify the above-mentioned traits as important directions for future research.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This study provides an important assessment of how body size influences the occurrence of macro-organisms in urban areas across the globe. Size in most plants, but only some animal families, was positively associated with urban tolerance. The data set is impressive, but the evidence for broad-scale conclusions is incomplete due to methodological issues that need to be resolved.

      We have substantially revised the manuscript to resolve the methodological issues raised, including clarifying the definition, calculation, and interpretation of urban affinity (formerly named urban tolerance), and tightening the scope of our conclusions to align directly with the evidence presented.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      The authors integrate multiple large databases to test whether body sizes were positively associated with which species tolerate urban areas. In general, many plant families showed a positive association between body size and urban tolerance, whereas a smaller, though still non-trivial, percentage of animal families showed the same pattern. Notably, the authors are careful in the interpretation of their findings and provide helpful context for the ways that this analysis can be generative in shaping new hypotheses and theory around how urbanization influences biodiversity at large. They are careful to discuss how body size is an important trait, but the absence of a relationship between body size and urban tolerance in many families suggests a variety of other traits undergird urban success.

      We appreciate this thoughtful and balanced assessment of our work and fully agree with the reviewer’s interpretation. In particular, we share the view that the heterogeneous and often weak association between body size and urban affinity across many families is an important result in its own right, underscoring that no single trait is likely to explain urban success across the tree of life. As the reviewer notes, our intention was not to present body size as a universal predictor, but rather as a widely available, integrative trait that can help reveal where general patterns do and do not emerge. We view the lack of a consistent relationship in many families as strong motivation for future work that explicitly integrates additional functional traits and ecological contexts, and we have clarified this perspective in the revised manuscript.

      Strengths:

      The authors aggregated a large dataset, but they also applied robust filters to ensure they had an adequate and representative number of detections for a given species, family, geography, etc. The authors also applied their analysis at multiple taxonomic scales (family and order), which allowed for a better interpretation of the patterns in the data and at what taxonomic scale body size might be important.

      We thank the reviewer for highlighting these strengths of the study. Considerable effort went into assembling, harmonizing, and filtering these data across taxa, regions, and taxonomic resolutions, and we were deliberate in applying conservative thresholds to ensure that species-level urban affinity estimates were based on adequate and comparable sampling. We hope that, beyond the specific results presented here, the compiled dataset and analytical framework will serve as a valuable resource for future studies aiming to explore additional traits, taxa, or mechanisms underlying species’ responses to urbanization.

      Weaknesses:

      My main concern is that it is not fully clear how the measure of body size might influence the result. The authors were unable to obtain consistent measures of body size (mean, median, maximum, or sex variation). This, of course, could be very consequential as means and medians can differ quite a bit, and they certainly will differ substantially from a maximum. And of course, sex differences can be marked in multiple directions or absent altogether. The authors do note that they selected the measure that was most common in a family, but it was not clear whether species in that family that did not have that measure were removed or not. This could potentially shape the variability in the dataset and obscure true patterns. This may require additional clarity from the authors and is also a real constraint in compiling large data from disparate sources.

      We appreciate this important point and agree that heterogeneity in how body size is measured (e.g., mean vs. maximum values, sex-specific measures) is a real but unavoidable challenge when compiling organismal trait data across such a broad taxonomic scope. We would like to clarify that our analytical approach was explicitly designed to minimize the influence of this heterogeneity rather than ignore it. Specifically, for each family we retained all species for which at least one body size estimate was available, rather than removing species that lacked a particular measurement type. When multiple body size measures existed for a species, we selected the measurement type that was most commonly available within that family in order to maximize comparability among species while retaining sample size. Importantly, differences among body size measurement types (including units, measurement detail, and whether values reflected means, maxima, or sex-specific estimates) were further accounted for by (i) log-transforming all body size values and (ii) centering and scaling body size values within each measurement type, which was included as a random effect in the hierarchical models. This approach reduces the influence of systematic differences among measurement types on estimated relationships with urban affinity. We have added a sentence to the methods clarifying that species with a single measurement type were not removed from analyses:

      “Importantly, this procedure did not result in the exclusion of species lacking a particular body size measurement type; rather, all species with at least one available body size estimate were retained, with measurement heterogeneity explicitly accounted for through hierarchical modeling.”

      We agree that variation in body size definitions may still contribute residual noise and potentially obscure weak relationships, and we now emphasize this more clearly as a limitation of large-scale trait syntheses. However, because our primary inference focuses on the presence, absence, and direction of size–urban affinity relationships across families, rather than precise effect sizes, we believe our approach provides a robust and conservative test of whether body size consistently predicts urban affinity across taxa. We highlight this point in the limitations section of our manuscript:

      “One important limitation of our synthesis is the heterogeneity in how body size is measured across taxa, including differences among mean, maximum, and sex-specific estimates. While our analytical framework explicitly accounts for this variation through transformation, scaling, and hierarchical modeling with random intercepts (see Methods), residual measurement noise may still obscure weak size–urban affinity relationships. This challenge is inherent to large-scale trait syntheses that integrate data from disparate sources, and highlights the need for continued efforts to standardize trait databases and expand the availability of harmonized organismal trait data across the tree of life.”

      Reviewer #2 (Public review):

      I have completed a thorough review of this paper, which seeks to use the large datasets of species occurrences available through GBIF to estimate variation in how large numbers of plant and animal species are associated with urbanization throughout the world, describing what they call the "species urbanness distribution" or SUD. They explore how these SUDs differ between regions and different taxonomic levels. They then calculate a measure of urban tolerance and seek to explore whether organism size predicts variation in tolerance among species and across regions.

      The study is impressive in many respects. Over the course of several papers, Callaghan and coauthors have been leaders in using "big [biodiversity] data" to create metrics of how species' occurrence data are associated with urban environments, and in describing variation in urban tolerance among taxa and regions. This work has been creative, novel, and it has pushed the boundaries of understanding how urbanization affects a wide diversity of taxa. The current paper takes this to a new level by performing analyses on over 94000 observations from >30,000 species of plants and animals, across more than 370 plant and animal taxonomic families. All of these analyses were focused on answering two main questions:

      (1) What is the shape of species' urban tolerance distributions within regional communities?

      (2) Does body size consistently correlate with species' urban tolerance across taxonomic groups and biogeographic contexts?

      We thank the reviewer for their careful reading of the manuscript and for this generous and accurate summary of the study’s aims, scope, and contributions. We appreciate the recognition of our group’s broader body of work using large biodiversity databases to quantify species’ associations with urban environments, and we are grateful for the reviewer’s acknowledgement that this study extends those efforts to an unprecedented taxonomic and geographic scale. We agree with the reviewer’s articulation of the two core questions motivating the paper, and we have revised the manuscript to ensure that these questions are stated clearly and addressed consistently throughout.

      Overall, I think the questions are interesting and important, the size and scope of the data and analyses are impressive, and this paper has a potentially large contribution to make in pushing forward urban macroecology specifically and urban ecology and evolution more generally.

      Thanks! We see this work as an effort to move beyond species-by-species descriptions of urban responses toward a community- and distribution-level perspective, where the shape of species’ urban associations themselves becomes an object of study. By framing species’ distributions along an urbanization gradient as a collective property of regional species pools, our approach opens a complementary way of thinking about how urbanization filters biodiversity.

      Despite my enthusiasm for this paper and its potential impact, there are aspects that could be improved, and I believe the paper requires major revision.

      Some of these revisions ideally involve being clearer about the methodology or arguments being made. In other cases, I think their metrics of urban tolerance are flawed and need to be rethought and recalculated, and some of the conclusions are inaccurate. I hope the authors will address these comments carefully and thoroughly. I recognize that there is no obligation for authors to make revisions. However, revising the paper along the lines of the comments made below would increase the impact of the paper and its clarity to a broad readership.

      We appreciate the detailed comments provided and have addressed each point in turn - see detailed responses below. We took these concerns seriously and undertook a substantial revision of the manuscript. In summary, we clarified the conceptual framing of “urban tolerance” (now referred to as “urban affinity”), explicitly defined the metric and its interpretation, added equations and a step-by-step methodological roadmap, and expanded justification for our regional stratification. Where appropriate, we refined language in the Results and Discussion to ensure conclusions are tightly aligned with what the metric can and cannot support. We agree that these revisions materially improve the clarity, rigor, and interpretability of the study, and we appreciate the reviewer’s perspective on how doing so strengthens the paper’s contribution and accessibility to a broad readership.

      Major Comments:

      (1) Subrealms

      Where does the concept of "subrealms" come from? No citation is given, and it could be said that this sounds like an idea straight out of Middle Earth. How do subrealms relate to known bioclimatic designations like Koppen Climate classifications, which would arguably be more appropriate? Or are subrealms more socio-ecologically oriented? From what I can tell, each subrealm lumps together climatically diverse areas. It might be better and more tractable to break things in terms of continents, as the rationale for subrealms is unclear, and it makes the analyses and results more confusing. The authors rationalized the use of subrealms to account for potential intraspecific differences in species' response to urbanization, but that is never a core part of the questions or interpretation in the paper, and averaging across subrealms also accounts for intraspecific variation. Another issue with using the subrealm approach is that the authors only included a species if it had 100 observations in a given subrealm, leading to a focus on only the most common species, which may be biased in their SUD distribution. How many more species would be included if they did their analysis at the continental or global scale, and would this change the shape of SUDs?

      We thank the reviewer for raising this point and agree that the rationale for using subrealms required clearer explanation. Next to allowing potential intraspecific differences in urban affinity across regions, our subrealm-based approach also provides a practical way to partition global biodiversity into ecologically meaningful regional assemblages while maintaining sufficient sample sizes for analysis. Urban affinity is likely to vary geographically within species due to differences in climate, habitat availability, urban form, and evolutionary history. By calculating urban affinity within subrealms rather than globally, our approach allows species to exhibit region-specific urban affinities while ensuring that comparisons are made among species co-occurring within the same regional ecological context. We have substantially revised the Methods to explicitly define subrealms, cite their origin, and clarify why this spatial stratification is appropriate for our study:

      “Accounting for geographic context through subrealm stratification

      To account for geographic heterogeneity in both species’ distributions and the baseline levels of urbanization, we stratified our analyses by global biogeographic subrealms (N=52; Fig. S1). Subrealms represent an intermediate hierarchical level within the One Earth [82] (https://www.oneearth.org/bioregions/) bioregionalization framework, grouping the 185 terrestrial bioregions into broader units that reflect shared species pools and ecological contexts while maintaining meaningful regional structure. This scale represents a practical compromise between analyzing data at the finer bioregion level (which would result in many regions with insufficient observations for robust analysis) and broader classifications such as continents or the 14 biogeographic realms, which aggregate ecologically distinct regions and species pools. This regionalization has been widely used in macroecological and biogeographic research to contextualize species–environment relationships because subrealms capture meaningful gradients in biotic assemblages that are not accounted for by climatic classifications alone [83,84].

      This stratification allows species’ associations with urban environments to be interpreted relative to the environments available within the regions they occupy. This is important, as previous work has shown that species’ responses to urbanization are constrained by biogeographic context, because regional species pools reflect shared evolutionary, ecological, and historical filters [23]. Previous work has also shown that urban associations among species are context-dependent, and interpreting species’ responses without accounting for regional baselines conflates availability of urban environments with species’ affinity to them. This distinction is critical because identical levels of urbanization (e.g., VIIRS radiance) can have different ecological meanings across regions with different species pools and land-use histories. It avoids conflating species’ urban affinity with global differences in urban availability.”

      We chose subrealms rather than Köppen climate classifications or continental units because our objective was not to partition species by climatic similarity per se, but to evaluate species’ associations with urban environments relative to the ecological and biogeographic contexts in which they occur. Climatic classifications such as Köppen are highly effective for addressing climate–species relationships, but they do not explicitly capture differences in species pools, evolutionary history, or land-use legacies that strongly shape how species interact with urbanization. Likewise, continents often aggregate ecologically disparate regions and species pools, potentially obscuring meaningful variation in baseline urbanization and species’ realized distributions.

      Importantly, urban affinity in our framework is a relative, context-dependent metric, explicitly interpreted within regions. Identical levels of urbanization (e.g., VIIRS radiance values) can have different ecological meanings across regions with distinct species pools, land-use histories, and settlement patterns. Stratifying analyses by subrealm therefore avoids conflating species’ affinity to urban environments with global or continental differences in the availability and intensity of urban land cover. We have clarified this distinction and motivation in the revised Methods (see responses below).

      Regarding the concern that requiring ≥100 observations per species per subrealm biases analyses toward common species: we agree that this threshold focuses the analysis on well-sampled species. This choice was intentional and follows previous work showing that such cutoffs are necessary to robustly characterize species’ responses to urbanization using occurrence data. While a global or continental analysis would indeed include additional, rarer species, it would also substantially increase uncertainty and conflate species’ responses across ecologically distinct contexts. Our study is therefore best interpreted as a macroecological synthesis of common species, which are also the taxa that disproportionately structure urban communities and drive the shape of Species Urbanness Distributions (SUDs). We now clarify this scope and limitation more explicitly in the introduction:

      “Our aim is to identify broad, cross-taxonomic patterns in species’ urban affinity at a global scale, rather than to resolve the specific causal mechanisms driving urban success or failure within individual taxa or cities.”.

      As well as in the discussion:

      “Our synthesis complements taxon-specific, presence–absence trait studies by identifying broad, cross-taxonomic patterns that can motivate and contextualize more mechanistic analyses [17,23].”

      Finally, while alternative spatial stratifications are possible, the central patterns we report particularly the skewed shape of SUDs—are robust to the use of regional context rather than absolute global metrics. Exploring how SUDs change under different spatial frameworks (e.g., continents, climate zones) is an interesting avenue for future work, but we feel is beyond the scope of the present study.

      (2) Methods - urban score

      The authors describe their "urban score" as being calculated as "the mean of the distribution of VIIRS values as a relative species specific measure of a response to urban land cover."

      I don't understand how this is a "relative species-specific measure". What is it relative to? Figures S4 and S5 show the mean distribution of VIIRS for various taxa, and this mean looks to be an absolute measure. Mean VIIRS for a given species would be fine and appropriate as an "urban score", but the authors then state in the next sentence: "this urban score represents the relative ranking of that species to other species in response to urban land cover".

      We agree that the wording in the original manuscript was unclear and conflated two distinct steps in the workflow. We have now revised the Methods to clearly distinguish between (i) the urban score, which is an absolute, descriptive summary of the mean VIIRS radiance associated with a species’ occurrence locations, and (ii) urban affinity, which is the relative, region-specific metric derived from the urban score. Specifically, we rewrote the methods to have distinct steps as subheadings, as follows: (1) urban score; (2) subrealms and why; (3) urban affinity. In the revised Methods, we explicitly define the urban score:

      “an absolute descriptive summary of the urbanization levels associated with a species’ occurrence locations within a given subrealm”.

      We no longer describe the urban score itself as “relative” or as a ranking among species. Relative comparisons among species arise only in the subsequent step, where species-specific urban scores are expressed relative to the regional background level of urbanization within each subrealm to derive urban affinity.

      We refer the Reviewer to the revised version which we feel is much clearer (lines 428-479)!

      That doesn't follow from the description of how this is calculated. Something is missing here. Please clarify and add an explicit equation for how the urban score is calculated because the text is unclear and confusing.

      The previous response, where we discuss the description, hopefully clarifies this. Further, we have revised the Methods to clearly define the urban score and to include an explicit equation. In the revised manuscript, the urban score for species s is calculated as the mean VIIRS radiance across all occurrence locations of that species:

      where n<sub>s</sub>is the number of GBIF occurrence records for species s, and L<sub>i</sub> is the VIIRS nighttime lights radiance value extracted at the location of occurrence i. We also clarify in the Methods that this urban score is an absolute summary statistic of observed urbanization at species occurrence locations

      (3) Methods - urban tolerance

      How the authors are defining and calculating tolerance is unclear, confusing, and flawed in my opinion.

      Tolerance is a common concept in ecology, evolution, and physiology, typically defined as the ability for an organism to maintain some measure of performance (e.g., fitness, growth, physiological homeostasis) in the presence versus absence of some stressor. As one example, in the herbivory literature, tolerance is often measured as the absolute or relative difference in fitness of plants that are damaged versus undamaged

      (e.g., https://academic.oup.com/evolut/article/62/9/2429/6853425?login=true).

      On line 309, after describing the calculation of urban scores across subrealms, they write: "Therefore, a species could be represented across multiple subrealms with differing measures of urban tolerance (Fig. S4). Importantly, this continuous metric of urban tolerance is a relative measure of a species' preference, or affinity, to urban areas: it should be interpreted only within each subrealm". This is problematic on several fronts. First, the authors never define what they mean by the term "tolerance". Second, they refer to urban tolerance throughout the paper, but don't describe the calculation until, where they write (text in [ ] is from the reviewer): "Within each subrealm, we further accounted for the potential of different levels of urbanization by scaling each species' urban score by subtracting the mean VIIRS of all observations in the subrealm (this value is hereafter referred to as urban tolerance). This 'urban tolerance' (Fig. S5) value can be negative - when species under-occupy urban areas [relative to the average across all species] suggesting they actively avoid them-or positive-when species over-occupy urban areas [relative to the average across all species] suggesting they prefer them (i.e., ranging from urban avoiders to urban exploiters, respectively). They are taking a relativized urban score and then subtracting the mean VIIRS of all observations across species in a subrealm. How exactly one interprets the magnitude isn't clear and they admit this metric is "not interpretative across subrealms".

      This is not a true measure of tolerance, at least not in the conventional sense of how tolerance is typically defined. The problem is that a species distribution isn't being compared to some metric of urbanness, but instead it is relative to other species' urban scores, where species may, on average, be highly urban or highly nonurban in their distribution, and this may vary from subrealm to subrealm. A measure of urban tolerance should be independent of how other species are responding, and should be interpretable across subrealms, continents, and the globe.

      We thank the reviewer for this careful and important critique. We agree that the term “tolerance” is commonly used to describe the ability of an organism to maintain performance (e.g., fitness, growth, physiological homeostasis) in the presence of a stressor, and that our metric does not measure tolerance in this mechanistic or fitness-based sense. To address this directly and unambiguously, we have revised the manuscript to explicitly define the term “urban affinity” as opposed to urban tolerance. 

      In the revised Methods, we also reorganized and clarified the calculation of urban affinity, introduced explicit notation, and provided a formal equation. Specifically, we now define urban affinity for species s in subrealm r as:

      where U<sub>s,r</sub>is the mean VIIRS radiance across all occurrence locations of species s within subrealm r, and Ū<sub>r</sub>is the mean VIIRS radiance across all occurrence records of all species in that subrealm. This transformation centers species’ urban scores on the regional background level of urbanization, yielding a relative measure of spatial association with urban environments.

      We agree with the reviewer that this metric is not interpretable as an absolute measure of affinity, and we now state this explicitly. Urban affinity values are, by construction, relative measures, interpretable only within subrealms, and they quantify whether a species tends to occur in more or less urbanized environments than is typical for that region. The magnitude of the metric therefore reflects deviation from the regional baseline, not a universal or global scale of urbanization, and is not intended to be compared directly across subrealms.

      We respectfully disagree, however, that this makes the metric flawed. Rather, it reflects a deliberate analytical choice aligned with our research questions. Our goal was not to estimate absolute urban exposure or physiological performance, but to compare species’ realized spatial associations with urban environments within shared biogeographic contexts. Because baseline urbanization levels, settlement history, and species pools vary strongly across regions, a globally absolute metric would conflate species’ affinities with regional availability of urban environments. By contrast, a relative, region-centered metric allows meaningful comparisons among species that coexist within the same ecological and biogeographic setting. This approach follows a growing body of macroecological work that infers species’ environmental affinities from spatial distributions rather than direct performance measures (e.g., Callaghan et al. 2020; 2021; 2023), and we now cite these studies explicitly.

      I propose the authors use one of two metrics of urban tolerance:

      (i) Absolute Urban Tolerance = Mean VIIRS of species_i - Mean VIIRS of city centers Here, the mean VIIRS of city centers could be taken from the center of multiple cities throughout a subrealm, across a continent, or across the world. Here, the units are in the original VIIRS units where 0 would correspond to species being centered on the most extreme urban habitats, and the most extreme negative values would correspond to species that occupy the most non-urban habitats (i.e., no artificial light at night). In essence, this measure of tolerance would quantify how far a species' distribution is shifted relative to the most highly urbanized habitat available.

      (ii) % Urban Tolerance = (Mean VIIRS of species_i - Mean VIIRS of city centers)/MeanVIIRS of city centers * 100%

      This metric provides a % change in species mean VIIRS distribution relative to the most urban habitats. This value could theoretically be negative or positive, but will typically be negative, with -100% being completely non-urban, and 0% being completely urban tolerant.

      Both of these metrics can be compared across the world, as it would provide either absolute (equation 1) or relative (equation 2) metrics of urban tolerance that are comparable and easily interpretable in any region.

      In summary, the definition of tolerance should be clear, the metric should be a true measure of tolerance that is comparable across regions, and an equation should be given.

      We thank the reviewer for this thoughtful and constructive suggestion, which raises an important conceptual issue regarding how “urban tolerance” should be defined and quantified. We agree that any such metric must be clearly defined, interpretable, and accompanied by an explicit equation, and we have revised the manuscript accordingly to clarify both our definition and its intended interpretation.

      The alternative metrics proposed by the reviewer anchoring species’ distributions to city centers or to the most highly urbanized habitats represent a valid and intuitive absolute framing of urban tolerance. Indeed, a closely related approach was explored and evaluated in Callaghan et al. (2020; https://doi.org/10.1016/j.ecolind.2020.106905), where species’ occurrence-based urbanness scores derived from VIIRS night-time lights were compared against abundance-based estimates of urban tolerance using explicit urban–non-urban contrasts. That study further demonstrated that urbanness scores depend on the choice of spatial baseline (e.g., regional buffers around cities versus continental extents), and showed that different baselines capture complementary, but not identical, aspects of species–urban associations.

      In the present study, we deliberately adopt a relative, regionally contextualized metric (now referred to as urban affinity), expressing each species’ mean VIIRS association relative to the background urbanization of the biogeographic subrealm in which it occurs. This choice reflects our goal of comparing species’ relative affinities to urban environments within shared ecological and biogeographic contexts. Importantly, identical VIIRS values can correspond to very different ecological conditions across regions, and anchoring all species to city centers or global urban maxima risks conflating species’ affinities with regional differences in urban availability and infrastructure.

      We now make this distinction explicit throughout the manuscript, including by (i) defining urban affinity as a relative, occurrence-based measure of urban affinity (rather than physiological or fitness-based tolerance), (ii) providing an explicit equation for its calculation, and (iii) clarifying that these values are interpretable within, but not across, biogeographic subrealms. We view absolute, city-center–anchored metrics and relative, regionally normalized metrics as complementary approaches, each suited to different questions; the latter is most appropriate for the macroecological, comparative analyses pursued here.

      (4) Figure 1: The figure does not stand alone. For example, what is the hypothesis for thermophily or the temperature-size rule? The authors should expand the legend slightly to make the hypotheses being illustrated clearer.

      We now expanded the legend so that the figure and hypotheses presented can be understood based on just the figure and its legend; we did so by explaining the illustrated hypotheses as requested by the Reviewer. The figure legend now reads as follows:

      “Fig. 1: Conceptual framework illustrating hypothesized mechanisms linking urban affinity to interspecific body-size shifts. These include dispersal and mobility constraints under habitat fragmentation [44,45], thermophily and the temperature–size rule driven by the urban heat island effect [15,30], size-biased competition and survival [94,95], and size-biased human preferences [64]. Urban fragmentation of habitat resources can select for increased mobility (e.g., larger butterflies) or reduced mobility (e.g., larger seeds) depending on isolation severity. Elevated urban temperatures favor thermophily, which often negatively correlates with size as it affects the heat balance via thermal inertia. Similarly, these higher temperatures generally favor smaller-bodied adult ectotherms because they accelerate development and reduce time available for growth (i.e., temperature-size rule). In plants, the increased CO<sub>₂</sub> and nutrient availability associated with anthropogenic environments due to heating- and traffic-related CO2 emissions and eutrophication provides a competitive advantage to larger plant species, and human preferences too may favor larger species (e.g., tree-lined streets), whereas smaller species may be advantaged in colonizing built infrastructure.”

      (5) SUDs: I don't agree with the conclusion given on line 83 ("pattern was consistent across subrealms and several taxonomic levels") or in the legend of Figure 2 ("there were consistent patterns for kingdoms, classes, and orders, as shown by generally similar density histograms shapes for each of these").

      The shapes of the curves are quite different, especially for the two Kingdoms and the different classes. I agree they are relatively consistent for the different taxonomic Orders of insects.

      We agree that our original wording overstated the similarity of distributions across taxa and regions. We have revised the text to clarify that the consistency we refer to pertains primarily to central tendencies rather than identical distributional shapes. To address this directly, we conducted additional analyses comparing urban affinity distributions across subrealms for taxonomic groups with the largest sample sizes. These results, now presented in new Supplementary Figures (Fig. S2-S4), show that while distributional shapes vary among higher taxonomic groups, median values and overall spread are broadly similar within comparable taxonomic levels. We have updated the Results text and the Figure 2 legend accordingly to reflect this more precise interpretation. 

      “These patterns in central tendency were broadly consistent across subrealms and taxonomic levels, although distributional shapes varied among higher taxonomic groups (Fig. 2).”

      “To evaluate this more formally, we compared distributions across subrealms for groups with the largest sample sizes and found that while distributional shapes varied among higher taxa, median values and overall spread were broadly similar within comparable taxonomic levels (Fig. S2–S4).”

      Figure 2 caption: “There were consistent patterns for kingdoms, classes, and orders (B) as shown by similar central tendencies despite variation in distributional shape.”

      We refer the Reviewer to the revised manuscript and supplementary material, but show the kindom level in Fig S2.

      More broadly, our goal in introducing Species Urbanness Distributions (SUDs) is not to argue that their exact shapes are invariant, but rather to provide a generalizable framework for describing how assemblages are structured along an urbanization gradient. In this respect, SUDs are conceptually analogous to Species Abundance Distributions (SADs), where the precise functional form has long been debated, yet the framework itself has proven extremely valuable for ecology. We therefore emphasize the utility of SUDs as a descriptive and comparative tool for quantifying community-level responses to urbanization, rather than as a claim about strict uniformity in distributional shape across taxa or regions.

      Reviewer #3 (Public review):

      Summary:

      This paper reports on an association between body size and the occurrence of species in cities, which is quantified using an 'urban score' that can be visualized as a 'Species Urbanness

      Distribution' for particular taxa. The authors use species records from the Global Biodiversity Information Facility (GBIF) and link the occurrence data to nighttime lighting quantified using satellite data (Visible Infrared Imaging Radiometer Suite-VIIRS). They link the urban score to body size data to find 'heterogeneous relationship between body size and urban tolerance across the tree'. The results are then discussed with reference to potential mechanisms that could possibly produce the observed effects (cf. Figure 1).

      We thank the reviewer for this clear and accurate summary of the study. We agree that the primary contribution of this work lies in the scale and taxonomic breadth of the analysis, and in introducing a framework (Species Urbanness Distributions) for quantifying species’ relative affinities to urban environments using globally available data. We have revised the manuscript to further clarify the scope of inference and the distinction between descriptive macroecological patterns and mechanistic explanations.

      Strengths:

      The novelty of this study lies in the huge number of species analyzed and the comparison of results among animal taxa, rather than in a thorough analysis of what traits allow species to persist under urban conditions. Such analyses have been done using a much more thorough approach that employs presence-absence data as well as a suite of traits by other studies, for example, in (Hahs et al. 2023, Neate-Clegg et al. 2023). The dataset that the authors produced would also be very valuable if these raw data were published, both the cleaned species records as well as the body sizes. The paper could strongly add to our understanding of what species occur in cities when the open questions are addressed.

      We appreciate highlighting the novelty of the taxonomic breadth and scale of our analysis. We agree that our approach is complementary to more detailed, taxon-specific trait studies based on presence–absence data. In response, we have further emphasized this distinction in the Discussion:

      “Our synthesis complements taxon-specific, presence–absence trait studies by identifying broad, cross-taxonomic patterns that can motivate and contextualize more mechanistic analyses17,23.”

      We also agree that the cleaned occurrence data and body size information represent a valuable resource, and all data will be made available, with the exception of some body size datasets which we are not able to make available.

      Weaknesses:

      I value the approach of the authors, but I think the paper needs to be revised.

      In my view, the authors could more carefully validate their approach. Currently, any weakness or biases in the approach are quickly explained away rather than carefully explored. This concerns particularly the use of presence-only data, but also the calculation of the urban score.

      The vast majority of data in GBIF is presence-only data. This produces a strong bias in the analysis presented in the paper. For some taxa, it is likely that occurrences within the city are overrepresented, and for other taxa, the opposite is true (cf. Sweet et al. 2022). I think the authors should try to address this problem.

      We thank the reviewer for raising this important point. We fully agree that GBIF occurrence data are subject to well-known sampling biases, including uneven geographic coverage, observer effort, and taxonomic focus. These limitations are now more explicitly acknowledged in the revised manuscript. At the same time, GBIF currently represents the only global biodiversity database that allows the scope of analysis undertaken here, spanning thousands of species across multiple taxonomic groups and regions. Systematic monitoring datasets that provide presence–absence data are typically restricted to particular taxa (often vertebrates or plants) and are geographically concentrated in the Global North, which would substantially limit the taxonomic and geographic breadth of our analysis.

      Importantly, our objective was not to estimate absolute species-specific responses to urbanization, but rather to examine relative patterns of urban affinity across species and families within comparable regional contexts. To address this, we structured our analyses at the subrealm level, which aggregates observations across large spatial extents and reduces sensitivity to fine-scale sampling biases associated with individual cities or urban–rural gradients. In addition, we restricted analyses to species with ≥100 observations per subrealm to focus on well-sampled taxa and reduce the influence of extremely sparse occurrence records. While these steps cannot fully eliminate sampling biases inherent to occurrence data, they substantially mitigate their influence when examining broad comparative patterns.

      Recent work has also evaluated the performance of GBIF data in urban biodiversity contexts. For example, Sweet et al. (2022) compared GBIF-derived species richness patterns with independent state-level biodiversity databases across cities and surrounding regions, finding that GBIF provided comparable or broader coverage across taxa and spatial extents. Their analysis showed that species richness was consistently higher in the surrounding region than in the city itself, suggesting that GBIF data capture broad urban–regional biodiversity gradients rather than systematically overrepresenting urban occurrences. Although our analysis differs in design, these results support the use of GBIF as a valuable resource for examining large-scale biodiversity patterns.

      More broadly, occurrence databases such as GBIF have become widely used for analyzing species–environment relationships at macroecological scales. While they may be insufficient for estimating precise species-specific environmental tolerances, they are informative for identifying broad patterns across taxa and regions. Our goal here is therefore to identify large-scale comparative patterns in urban affinity and generate hypotheses about trait– urbanization relationships, which can subsequently be tested with more structured monitoring datasets where available.

      Another important consideration is that our analyses focus on comparative differences among species within shared taxonomic and geographic contexts, rather than absolute estimates of urban affinity. Sampling biases in occurrence databases are often structured by observer behaviour (e.g., detectability, accessibility, or taxonomic interest), meaning that species recorded by similar observer communities are likely subject to similar sampling biases. Under these conditions, relative differences among species are expected to be preserved even when absolute occurrence frequencies are biased. This logic is consistent with the widely used target-group background approach in presence-only species distribution modelling, where species recorded by similar observer groups (often within the same taxonomic group) are used to control for shared sampling bias. Previous work by Callaghan et al. (2021; https://doi.org/10.1111/gcb.15670) performed additional validation analysis comparing our distribution-based urban affinity metric with estimates derived from occupancy modelling using well-sampled European butterflies (see Fig. S5 from the Callaghan et al. 2021 paper). The strong positive relationship between these approaches suggests that the broad patterns identified here are unlikely to arise solely from sampling artifacts.

      Finally, in the revised manuscript we now include additional comparisons among well-sampled taxonomic groups (see responses to other comments throughout our response document for details), which show substantial variation in urban affinity even among taxa with extensive sampling. These results suggest that the patterns reported here are unlikely to arise solely from sampling artifacts, but instead reflect meaningful ecological variation in how species interact with urban environments.

      The authors should compare their results to studies focusing on particular taxa where extensive trait-based analyses have already been performed, i.e., plants and birds. In fact, I strongly suggest that the authors should compare their results to previous studies on the relationship between traits, including body size and occurrences along a gradient of urbanisation, to draw conclusions about the validity of the approach used in the current study, which has a number of weaknesses.

      We agree that explicitly situating our findings within the existing trait-based urban ecology literature strengthens both interpretation and validation of our approach. We had already referenced several relevant studies (e.g., Hahs et al. 2023 and others) in the Introduction and Discussion, but we recognize that these comparisons were not sufficiently explicit. We have now added text to the Discussion directly comparing our results with previous trait-based studies across taxa:

      “Our results are broadly consistent with prior taxon-specific trait-based studies (eg., Hahs et al.[17]), but also highlight that relationships between body size and urbanization vary across taxa and analytical frameworks. For example, global syntheses and regional studies have reported positive, negative, or null size–urbanization relationships depending on clade and spatial scale. A recent global analysis that compiled empirical occurrence data for multiple terrestrial faunal taxa across cities worldwide reported broadly similar body-size responses to urbanization [17]. For four of the five groups that overlap with our analysis—amphibians, bats, bees, and birds—the direction of the body-size relationship with urbanization was consistent between studies. The only exception was carabid beetles, which tended to be smaller-bodied in highly urbanized environments in that analysis, whereas we detected no significant size effect for this family. Studies on birds, for example, have found mixed results, including positive associations to urbanization in some regional assemblages [45], no global relationship in others [46] or an overall negative relationship globally [23], and negative relationships in particular clades such as raptors [40]. Such discrepancies likely arise because different studies quantify urbanization differently, focus on different spatial grains, or analyze different components of species responses (e.g., presence– absence, abundance, or occurrence distributions). Additionally, a study on multiple taxa including butterflies and moths found a positive relationship in butterfly and moth community-weighed mean body size with increases in urbanization level, similar to our findings [31]. Researchers have also found that smaller-bodied dung-associated beetles potentially benefit from urban environments, which is similar to the negative association we found between urbanization and body size in beetles [47]. Our approach complements these studies by estimating occurrence-based urban associations across thousands of taxa simultaneously, allowing comparison of how consistently body size predicts urban affinity across taxonomic groupings rather than within a single lineage. In this sense, variation among published results does not contradict our findings but instead reinforces the conclusion that body size is a context-dependent filter whose direction and strength depend on ecological setting, taxonomic scope, and the urbanization metric used.”

      These additions highlight that published relationships between body size and urbanization vary widely across taxa, spatial scales, and analytical approaches. For example, prior studies have reported positive, negative, or null size–urbanization relationships depending on clade, geographic extent, and how urbanization or occurrence is quantified. Even within birds alone, the literature spans positive regional relationships, null global relationships, and negative relationships in particular clades such as raptors. We now explicitly discuss these contrasts and clarify that such discrepancies are expected because different studies measure different components of species’ responses (e.g., presence–absence vs. abundance vs. occurrence distributions), use different spatial grains, or focus on different taxonomic subsets.

      We emphasize that our analysis is not intended to replace taxon-specific trait studies, but rather to complement them by providing a macroecological synthesis across thousands of species simultaneously. Importantly, the heterogeneity we observe among families is itself a key biological result, indicating that body size is not a universal predictor of urban affinity but instead a context-dependent filter whose direction and strength vary across ecological and phylogenetic settings. We now state this interpretation more clearly in the revised manuscript.

      They should be be more careful in coming up with post-hoc explanations of why the pattern found in this study makes sense or suggests a particular mechanism. This reviewer considers that there is no way in which the current study can disentangle the different possible mechanisms without further analyses and data, so I would suggest pointing out carefully how the mechanisms could be studied.

      We agree that our study cannot disentangle the causal mechanisms underlying species’ responses to urbanization. Our intent in discussing potential mechanisms was not to claim definitive explanations, but rather to situate our findings within existing ecological theory and to highlight plausible, non-exclusive pathways that may generate the observed patterns. To make this clearer, we have revised the Discussion to explicitly frame these interpretations as hypotheses rather than conclusions, and to emphasize that testing the underlying mechanisms will require additional data and approaches, such as targeted trait datasets, experimental manipulations, and longitudinal or within-city studies:

      “Because our synthesis is correlative and macroecological in nature, the mechanisms discussed above are best viewed as hypotheses that can be evaluated through future work combining experimental, trait-based, and longitudinal data.”.

      Additionally, we modified our overall goal to make it clear that this is not inherently a mechanistic study per se:

      “Our aim is to identify broad, cross-taxonomic patterns in species’ urban affinity at a global scale, rather than to resolve the specific causal mechanisms driving urban success or failure within individual taxa or cities.”.

      More details should be given about the methodology. The readers should be able to understand the methods without having to read a number of other papers.

      We have substantially revised and expanded the Methods section to ensure that all analytical steps can be understood directly from the manuscript without requiring consultation of prior publications. In particular, we now (i) provide a clear conceptual roadmap of the workflow at the start of the Methods, (ii) define all key metrics explicitly, including equations for both the urban score and urban affinity, and (iii) clarify the interpretation, assumptions, and limitations of each step. We also added text explaining the rationale for subrealm stratification and the intended interpretation of relative values. Together, these revisions make the methodological framework fully transparent and self-contained (see revised Methods and related responses above and below).

      References:

      Hahs, A. K., B. Fournier, M. F. Aronson, C. H. Nilon, A. Herrera-Montes, A. B. Salisbury, C. G. Threlfall, C. C. Rega-Brodsky, C. A. Lepczyk, and F. A. La Sorte. 2023. Urbanisation generates multiple trait syndromes for terrestrial animal taxa worldwide. Nature Communications 14:4751.

      Neate-Clegg, M. H. C., B. A. Tonelli, C. Youngflesh, J. X. Wu, G. A. Montgomery, Ç. H. Şekercioğlu, and M. W. Tingley. 2023. Traits shaping urban tolerance in birds differ around the world. Current Biology 33:1677-1688.

      Sweet, F. S. T., B. Apfelbeck, M. Hanusch, C. Garland Monteagudo, and W. W. Weisser. 2022. Data from public and governmental databases show that a large proportion of the regional animal species pool occur in cities in Germany. Journal of Urban Ecology 8:juac002.

      We have incorporated these (and additional new references) into our revised manuscript.

      Recommendations for the authors:

      Reviewing Editor Comments:

      As you see from the general comments above and the specific recommendations below, the reviewers are impressed by your comprehensive data set and the analytic approach. However, they ask you to clarify your measures of organism size, occurrence data (vs. presence/absence and corresponding sample-bias caveats), urbanness (lighting differences between cities and regions?), urban tolerance (measure should not be relative to other species and particular regions), and region ("subrealm" vs. more commonly used defintions of world regions such as continents). They also encourage you to compare your general results with more detailed local studies to better justify using size as the only, easily available trait.

      We thank the Editor for this clear synthesis of the key priorities for revision. We have carefully addressed each point and substantially revised the manuscript to improve clarity, methodological transparency, and interpretability. In particular:

      We clarified how body size data were compiled, harmonized, and modeled, including explicit description of how different measurement types (mean, maximum, sex-specific) were retained and statistically accounted for through scaling and hierarchical modeling. We now state these procedures explicitly in the Methods.

      We expanded the Methods and Discussion to clarify that our analyses rely on occurrence data rather than presence–absence or abundance data, and we now explicitly discuss the implications and limitations of presence-only datasets, including potential sampling biases and how these may influence inference.

      We strengthened justification for using VIIRS night-time lights as a continuous proxy for urbanization, added supporting citations, and clarified that spatial heterogeneity in lighting primarily introduces additional variance rather than systematic bias. We also explicitly describe how urbanization values were calculated and interpreted.

      We substantially revised the manuscript to clearly define urban affinity at the outset (including in the Abstract), distinguish it from physiological definitions of tolerance, and provide explicit equations and step-by-step descriptions of how both urban score and urban affinity are calculated and interpreted. We now emphasize that the metric is a relative, region-contextualized measure of occurrence-based urban affinity.

      We added full justification, citations, and methodological explanation for the use of biogeographic subrealms, clarified how they differ from continents or climate zones, and explained why this stratification is appropriate for the ecological questions addressed. We also clarified the scope of inference and limitations of this approach.

      We expanded the Discussion to explicitly compare our results with prior trait-based urban ecology studies across taxa (including birds and other groups), highlighting where results converge, diverge, and why such variation is expected across spatial scales, taxa, and analytical frameworks.

      Reviewer #1 (Recommendations for authors):

      (1) Abstract

      (a) Please define how tolerance is being used here

      We now use affinity throughout and it is defined in various places (see responses to other comments here).

      (b) The abstract should clarify at what taxonomic scale body size is assessed. It is unclear in the abstract as to whether the reader expects intraspecific measures and interspecific, and at what resolution.

      We have revised the abstract by adding one sentence explicitly stating the scale body size was assessed:

      “We then assessed whether body size, an integrative ecological trait fundamental to space use, mobility, metabolism, and environmental sensitivity, showed consistent associations with urban affinity among species and across 371 taxonomic families. Analyses were conducted at the interspecific level and focused primarily on variation among taxonomic families (provided with this paper is an accompanying application to view results).”

      (2) Results/Discussion

      (a) The species urbanness distribution and comparison with the species abundance distribution is an interesting and conceptually useful contribution to urban ecology and underscores how urbanization functions on biodiversity at scale.

      We thank the reviewer for this positive assessment and are encouraged that they view the Species Urbanness Distribution (SUD) as a conceptually useful contribution to urban ecology. We see SUDs as a flexible framework that can be extended in several important directions, including comparisons across additional traits, cities of differing size and configuration, and temporal analyses that track how urbanness distributions shift with ongoing urban expansion or restoration. More broadly, we hope that SUDs can provide a framework to think about a macroecological understanding of how urbanization filters biodiversity.

      (b) In our Lambert et al. (2023) study that you reference, we suggest that 'exaptation' may be valuable to explore in urban areas. Although body size wasn't the trait we were considering at that time, it may be worth putting your discussion around pre-adaptation in this context.

      We agree that exaptation provides a valuable conceptual lens for interpreting species’ responses to urban environments. We have revised the Discussion to explicitly frame species’ urban success in this context:

      “Such traits “pre-adapted” to urban conditions allow for some species to not only persist but thrive in urban environments where most species cannot. Framing these patterns through the lens of exaptation may be particularly useful, as traits that evolved under non-urban selective pressures may incidentally confer advantages in urban environments without having arisen in response to urbanization per se (sensu Lambert et al.[4]). We therefore speculate that the skewed shape of SUDs may reflect the uneven distribution of exaptive traits across species pools, rather than widespread adaptive evolution to urban conditions. 

      Consistent with this interpretation, if exaptive traits that facilitate urban persistence are unevenly distributed across species pools, most species would be expected to exhibit avoidance rather than affinity of urban environments. Indeed, we found that the median urban affinity is most often below one, indicating widespread avoidance among species.”.

      (c) Given the family-scale effect, it would be helpful to discuss how often species within a family co-occur in a given geographic region, how much other traits covary with size, etc. Do we have an a priori reason to expect family to be the taxonomic resolution at which body size seems to be most varied?

      Our exploratory and preliminary analyses revealed that variation in the body size– urban affinity relationship was strongest at the family level, which prompted us to focus our main analyses at this taxonomic resolution. (But we also present results on order as well). Families represent a biologically meaningful intermediate scale in taxonomy: species within families typically share broad morphological, ecological, and life-history characteristics, yet still exhibit substantial variation in body size and ecological strategies. Indeed, body size is well known to covary with multiple traits—including dispersal ability, metabolism, and space use—making it an integrative trait that captures several ecological dimensions simultaneously within and among families. These correlated traits likely contribute to the heterogeneous responses to urbanization observed among families.

      Using the family level also provides a practical balance between biological relevance and statistical robustness. Many families contain sufficient numbers of species to allow independent model estimation while avoiding the strong data imbalance that would arise at higher taxonomic levels. In addition, family is a commonly used unit in macroecological trait analyses (e.g., Roy et al. 2009; Smith et al. 2004), and it often reflects major morphological and ecological similarities among species, as reflected in taxonomic identification frameworks.

      Regarding co-occurrence, our analytical framework already accounts for geographic context by estimating urban affinity within subrealms. This ensures that species are compared within the same regional species pools and environmental contexts, rather than across globally disparate assemblages. Consequently, family-level effects emerge from comparisons among species that co-occur within shared biogeographic settings rather than from global taxonomic aggregation.

      We have added a short clarification in the manuscript to emphasize that body size functions as an integrative trait that covaries with multiple ecological attributes, and that family-level analyses represent a balance between ecological interpretability and data availability:

      “Because body size covaries with multiple ecological traits (e.g., dispersal ability and metabolic rate), we focused on family-level analyses to capture shared ecological strategies while still allowing sufficient variation among species to detect trait– environment relationships [39]”.

      (d) The result that body size shows a stronger effect in plants perhaps could suggest that plant records in GBIF are more sensitive to potential collection bias, perhaps due to detectability differences or preferences for where botanists and citizen scientists collect plant data? You mention ornamental plants late, but it may be worth discussing this here, too.

      We agree that this is a possible mechanism, which likely conflates detectability and ecological signal. We have expanded this point in the discusssion to better address this:

      “These human-driven preferences may also influence detectability and recording effort, as larger and more conspicuous plant species are more likely to be planted, maintained, and documented in urban environments, and thus be available in GBIF for our analyses. However, we suggest that this is not purely a sampling artifact, but such processes likely interact with ecological filtering to shape the realized size structure of urban plant communities.”.

      (e) I appreciate the additional taxonomic layering to the discussion. Seeing patterns at the family and order levels is helpful for generating new theory and predictions about how urbanization structures biodiversity at different taxonomic scales.

      We agree that examining patterns across multiple taxonomic scales is particularly valuable for generating testable hypotheses about how urbanization structures biodiversity, as different mechanisms may emerge or break down depending on the resolution of analysis. We hope this multi-scale perspective helps stimulate new theory and predictions about the ecological processes shaping urban biodiversity across the tree of life.

      (3) Methods

      (a) The methodology provides a scalable, consistent, and reasonable measure of both urbanness and species-level urban tolerance. The urban tolerance measure will, of course, not be useful for certain types of research (e.g., animal behavior), but it is appropriate for the resolution of this study.

      We agree that the urban affinity metric presented here is intended for broad-scale, comparative analyses and is not designed to capture fine-scale processes such as individual behavior or short-term demographic responses. Our goal was to develop a scalable and consistent measure that enables cross-taxon and cross-region comparisons at a global extent, which we believe is appropriate for addressing the questions posed in this study. We have sought to be explicit about this scope throughout the manuscript (e.g., to better alleviate Reviewer #1 concerns) and emphasize that the framework is complementary to, rather than a replacement for, more mechanistic or organism-focused approaches.

      (b) I'm concerned that the authors were not able to constrain their dataset to mean, median, or maximum, not potentially sex variability in sizes. Later in the methods, the authors state that they selected the measure of size that was most common within a family. Does this mean that species within a given family that didn't have that measure of body size were removed from the analysis?

      We appreciate this important point and agree that heterogeneity in how body size is measured (e.g., mean, maximum, or sex-specific estimates) is a real and unavoidable challenge in large-scale trait syntheses. Our analytical approach was explicitly designed to minimize the influence of this heterogeneity while retaining as many species as possible, rather than excluding species based on inconsistent trait metadata.

      Specifically, species within a family were not removed based on the availability of a particular body size definition. All species with at least one body size estimate were retained. When multiple measures existed for a species, we selected the measurement type that was most commonly available within each family to maximize comparability while preserving sample size. Remaining heterogeneity among measurement types (including units, measurement detail, and whether values reflected means, maxima, or sex-specific estimates) was explicitly accounted for through log-transformation and metadata-aware centering and scaling, with measurement metadata included as random intercepts in the hierarchical models. We have clarified this point in the Methods:

      “Importantly, this procedure did not result in the exclusion of species lacking a particular body size definition; rather, all species with at least one available body size estimate were retained, with measurement heterogeneity explicitly accounted for through metadata-aware scaling and hierarchical modeling.”

      In addition, our taxonomic modeling strategy was intentionally hierarchical. Species belonging to families that did not meet the minimum threshold for family-level modeling (≥10 species) were not discarded; rather, they were included in higher-level taxonomic analyses (e.g., order- or class-level models), ensuring that available information was retained wherever statistically appropriate. This approach reflects our broader goal of maximizing data inclusion while matching inference to the resolution supported by the data.

      Reviewer #2 (Recommendations for the authors):

      (1) Overlap between VIIRS and GBIF data: While it would have been nice for the GBIF records and VIIRS timescales to match, the degree of mismatch isn't overly large (2010-2021 vs 2015-2021), and any bias or inaccuracies should be minimal. I am mainly making this comment as a potential counterpoint to a possible criticism from other reviewers.

      We thank the reviewer for this helpful observation and agree with their assessment. While the temporal coverage of GBIF occurrence records (2010–2021) and VIIRS night-time lights data (2015–2021) does not perfectly overlap, the mismatch is relatively small and unlikely to introduce substantial bias, particularly given our focus on broad, global patterns of urban affinity rather than fine-scale temporal dynamics. We appreciate the reviewer highlighting this point as a potential counterargument to concerns about temporal alignment.

      (2) Line 87: "only a select few species seem to possess traits that enable them to thrive in urban...".

      This seems like an odd statement, given how many of these species have positive urban tolerance measures.

      Agreed that this was oddly worded. We have revised for clarity, focusing on the magnitude of urban affinity:

      “Similarly, much like the skewed distributions observed in SADs [24,26], the skewed shape of SUDs indicates that while many species exhibit some degree of urban affinity, a relatively small subset of species attain high levels of urban affinity and dominate urban environments.”

      (3) Line 81: "skewed shape of SUDs suggests that traits enabling species to tolerate urban environments are both rare and specific".

      Again, based on the shape of some of these curves, I'm not convinced that it is rare, and there is nothing about these curves that suggests it is something "specific". Indeed, urban tolerance could be very multivariate, and the authors' own results suggest this is indeed the case.

      We have revised the sentence to retain a focus on traits while avoiding overinterpretation of adaptation from the distributional patterns alone. The revised wording emphasizes the uneven expression of high urban affinity across species without implying rarity or trait specificity:

      “The skewed shape of SUDs suggests that traits enabling species to tolerate urban environments are unevenly expressed, given that only a handful of species show extreme urban affinity values, but our results suggest this is geographically widespread across taxa.”.

      We also agree with the likelihood that it is multivariate, and return to this in the conclusion in a stronger sense:

      “Although body size emerged as a predictor of urban affinity, we found not only substantial heterogeneity across families and orders, but also that body size filtering alone is unlikely to explain the consistently skewed SUD shape. Taken together, these patterns suggest that urban affinity likely emerges from multiple trait combinations rather than a single, universally advantageous trait, and that strong affinity to urban environments is not uniformly expressed across taxa, despite occurring broadly across regions.”.

      (4) Line 100: "UHI", avoid abbreviations unless absolutely necessary.

      We have removed this abbreviation throughout.

      (5) Body size: focusing on one trait seems like a shot in the dark, and so it isn't too surprising that this didn't reveal a strong or consistent pattern. However, I also recognize that collecting consistent trait data across so many taxa is challenging, and size is a low-hanging fruit that correlates with multiple traits. Perhaps discuss more the range of traits you think are most likely to predict urban tolerance.

      Body size is indeed the ‘easiest’ to collect, but we acknowledge that there are other traits which could be important, and body size correlates with multiple traits. We revised our discussion to be more comprehensive to discuss some of the additional traits, and be explicit about the shortfalls of body size:

      “Ultimately, the heterogeneous and sometimes weak relationships between body size and urban affinity suggests that body size alone cannot explain the emergence of extreme urban exploiters and the skewed shape of SUDs. Focusing on body size as a focal trait necessarily represents a simplification of the multidimensional processes underlying species’ responses to urbanization, driven in part by data availability when conducting a taxonomically-broad synthesis. Instead, urban affinity likely depends on multivariate trait combinations [17,58] that vary among taxa [59] and ecological contexts [60]. Traits that are likely to correlate with urban affinity include dispersal capacity, behavioral flexibility, diet breadth, reproductive strategy, thermoregulatory ability, and, in plants, life history traits such as growth form, clonality, phenology, and seed size. The diversity of trait pathways through which species may persist or thrive in urban environments is consistent with the pronounced taxonomic heterogeneity we observe and helps explain why body size alone does not yield a universal pattern.”

      (6) Figure S2: This figure and analysis appear to 'come out of nowhere'. I think this is distracting and tangential, and it should be removed. I have the same thoughts about Figure S3. While I do think a discussion of other traits to measure is well warranted and needed, the inclusion of "preliminary' results that aren't motivated by clear questions, appropriate context, and rigorous analysis should be discouraged.

      We have removed Figure S2 and Figure S3 in response to this comment.

      I hope the authors find my constructive comments useful in their revision process.

      This was a very thorough and thoughtful review. We are greatly appreciative of the opportunity and guidance to improve our work!

      Reviewer #3 (Recommendations for the authors):

      Here is a list of a number of further points that the authors may want to address:

      (1) Figure 1 somehow misses the fact that humans simply do not want very large animals in the city. We kill large predators if they come too close to cities, and the same for large herbivores such as wild boar or deer.

      We agree that direct human persecution and management of large-bodied species can influence which species occur in urban environments, particularly for large predators and herbivores. Such processes represent important mechanisms shaping urban species assemblages and represent an entire field of socio-ecological dynamics. We have now clarified this point in the Discussion by noting that human–wildlife conflict, management, and persecution could contribute to observed size–urbanization relationships for some taxa, and that disentangling these mechanisms represents an important direction for future research. We added some text to highlight this point):

      “Similarly, human–wildlife conflict and active management of large-bodied animals in cities may influence which species persist in urban environments, potentially constraining the upper end of the body size distribution. Taken together, these examples illustrate the importance of considering the socio-ecological context of urban species assemblages [65]”.

      (2) Line 270. So you removed all data from the grid-based survey?

      We did not remove all data originating from grid-based surveys or gridded products. Rather, we retained GBIF point-occurrence records and applied a standard spatial filtering step, removing only those individual observations with reported coordinate uncertainty greater than 1 km. This was done to ensure reliable alignment between species occurrence points and remotely sensed environmental layers. We have clarified this distinction in the Methods to avoid confusion:

      “Due to uncertainty in matching observations with remotely-sensed products, any GBIF observation with a coordinate uncertainty > 1 km was removed. This filtering step removed individual observations with high spatial uncertainty, rather than excluding entire datasets or survey types.”.

      (3) Line 278. Human population density?

      Yes, we have added ‘human’ here (and elsewhere in this section) to make this clearer to the reader.

      (4) Line 284. What is a pixel?

      We have modified the text to make this clearer:

      “VIIRS Stray Light Corrected Nighttime Day/Night Band Composites product, representing monthly composites, (i.e., this dataset in Google Earth Engine: NOAA/VIIRS/DNB/MONTHLY_V1/VCMSLCFG) with a native resolution of ~500 m<sup>2</sup>. We took the median of all monthly composites for each pixel (i.e., a single grid cell of the night-time lights raster representing a fixed ground area) to calculate a pixel-level urbanization value, measured in average radiance, and used imagery from January 2015 to January 2021 to calculate this median”.

      (5) Line 292. It seems to me that lighting is different in different types of cities with the same level of impervious surface, depending on local customs of how many lights are installed, left switched on, etc. I guess that petrol stations and strongly lit industrial areas both produce high levels of light, while for the industrial areas, there could be lawn or other vegetation?

      We thank the reviewer for this thoughtful observation and agree that night-time lighting can vary across cities with similar levels of impervious surface due to differences in land use, infrastructure, and cultural lighting practices. We do not interpret VIIRS night-time lights as a direct measure of any single urban feature, but rather as a continuous, integrative proxy for urbanization that captures the combined footprint of human activity, infrastructure intensity, and energy use. VIIRS radiance has been repeatedly shown to correlate strongly with human population density, built infrastructure, and urban extent, while being negatively correlated with vegetation cover (e.g., EVI). It is repeatedly used in remote sensing and urban sustainability literature. This approach is widely supported in the literature, for example:

      Panić et al. used night-time lights were to map spatial and temporal patterns of artificial lighting as a proxy for human population distribution and activity, distinguishing areas of urban and rural occupancy.

      (https://www.ceeol.com/search/article-detail?id=1035395)

      Zhou et al. used night-time light observations were to develop a globally consistent time series of annual urban extent, delineating urban clusters and quantifying global urban growth over decades. (https://doi.org/10.1016/j.rse.2018.10.015)

      Chakraborty & Stokes used night-time light time series with machine learning to detect and quantify urban change processes—identifying deviations from expected radiance trends to monitor diverse urban transitions.

      (https://doi.org/10.1016/j.rse.2023.113818)

      Zhao et al. reviewed night-time light remote sensing was for its broad capacity to quantify human activities and socioeconomic dynamics—such as urbanization, economic change, and environmental impacts—across scales.

      (https://doi.org/10.3390/rs11171971)

      Zheng et al. used VIIRS nightime lights across 30 global megacities to produce a classification scheme to disentangle urban land changes into five categories, and assess global urbanization processes. (https://doi.org/10.1016/j.isprsjprs.2021.01.002)

      Zhao et al. argue that nighttime lights provide a consistent dataset to model and interpret urbanization dynamics and use this to track urban dynamics in Southeast Asia. (https://doi.org/10.1016/j.rse.2020.111980)

      While localized mismatches may occur (e.g., brightly lit industrial areas with surrounding vegetation), such heterogeneity is expected to introduce additional variance rather than systematic bias in the measure of urbanization, making our inference conservative. We have clarified this interpretation and added additional supporting references in the Methods:

      “Previous work has shown that VIIRS night-time lights is negatively correlated with greenness measured through the Enhanced Vegetation Index (EVI) and positively correlated with human population density [69,71]. Although night-time light intensity can vary among cities with similar impervious surface due to differences in land use, infrastructure, and cultural lighting practices, at broad spatial scales it functions as an integrative proxy of urbanization [75,76,77,78,79,80], with localized heterogeneity contributing primarily to additional variance rather than systematic bias.”

      (6) Line 295. How did you reconcile the spatial uncertainty of >1km with an urbanization pixel of 150m2? For how many species did you have a higher uncertainty than pixel size? In my experience, your ca. 39m accuracy is a strong assumption for GBIF data.

      We would like to clarify that we do not assume species occurrence accuracy at the scale of the geohash blocks (i.e., tens of meters), and we do not interpret GBIF records as having ca. 39 m positional accuracy. The use of geohash7 (~150 m blocks) reflects a computational indexing choice, not an assumption about biological or observational precision. All GBIF observations with reported coordinate uncertainty greater than 1 km were removed prior to analysis, ensuring that retained occurrences were compatible with the effective spatial resolution of the remotely sensed urbanization data. Importantly, the effective spatial resolution of our urbanization metric remains that of the VIIRS night-time lights product (~500 m). Geohash encoding at a finer resolution was used solely to efficiently associate point occurrences with the appropriate VIIRS pixel while avoiding redundant extraction or averaging across adjacent pixels. This approach does not increase the effective spatial precision of the analysis, nor does it imply sub-pixel inference. We have clarified this in the Methods:

      “The VIIRS night-time lights data, with a native resolution of ~500 m<sup>2</sup>, was then matched to these blocks by assigning each geohash7 block the average VIIRS radiance value that intersects it. We do not assume positional accuracy at the scale of the geohash blocks, but geohash encoding was used solely for computational indexing, while the effective spatial resolution of the urbanization metric is that of the VIIRS data (~500 m). This approach allows us to avoid unnecessary redundancy in the data while maintaining the original VIIRS resolution”.

      (7) Line 296. Why this high resolution in the species data when your light data is 500m2?

      The apparent mismatch in resolution reflects a distinction between data handling resolution and analytical resolution. Species occurrence records were retained at their native point-level precision to avoid premature spatial aggregation and to ensure that each observation could be accurately matched to the appropriate VIIRS night-time lights pixel. The finer-resolution geohash encoding does not imply that species data were analyzed at that scale, nor does it increase the effective spatial resolution of the analysis. We note, however, that the reported spatial uncertainty of some GBIF records may approach or exceed the resolution of the VIIRS data. Retaining such records represents a deliberate trade-off between spatial precision and data coverage, and is necessary to maximize taxonomic and geographic representation in a global analysis of this scope. Importantly, any residual spatial uncertainty is expected to introduce additional noise rather than systematic bias, making our estimates of species–urban affinity relationships conservative.

      (8) If you could show how your results match the results of Hahs et al and others with respect to occurrence and traits, this would strengthen your approach.

      We agree that explicitly comparing our findings with prior trait-based studies strengthens the interpretability of our approach. We have now added text to the Discussion that directly compares our results with published analyses, including Hahs et al. (2023) and other taxon-specific studies. In particular, we highlight where our occurrencebased estimates recover similar body size–urbanization relationships (four of five taxa in Hahs et al.) and where they differ (e.g., carabids), and we discuss how such differences likely arise from variation in spatial grain, response variables, and definitions of urbanization. These additions clarify how our framework aligns with, complements, and extends existing trait-based work rather than replacing it.

      (9) I wonder whether you could run your analysis with simplified data. In the end, you do not talk much about how high the urban score is, so you may also aggregate values to "highly lighted", "lighted", "some light" and "dark" and re-do the analysis, after checking how these scores correlate with e.g. impervious surface in a slightly larger area than what you used (maybe 50x50m).

      Our analytical framework—and the concept of Species Urbanness Distributions (SUDs) in particular—relies on retaining the continuous nature of the underlying urbanization metric. Discretizing night-time light values would necessarily introduce arbitrary thresholds, reduce information content, and obscure subtle but ecologically meaningful variation in species’ relative affinities to urban environments. Because we focus on relative affinity patterns rather than absolute urbanization classes, maintaining a continuous metric is central to both our methodological approach and conceptual contribution. That said, we agree that exploring how continuous urban affinity scores relate to categorical urban classes or alternative urbanization proxies (e.g., impervious surface at different spatial grains) represents a valuable direction for future work. Such analyses could be particularly informative for translating continuous affinity metrics into applied conservation or urban planning contexts.

    1. Author response:

      The following is the authors’ response to the original reviews.

      eLife Assessment

      This important study investigates how the brain categorizes written words from different writing systems (e.g., alphabetic vs. non-alphabetic), shedding potential light on the neural basis of language's social‑categorization function. Overall, the evidence supporting the authors' claims is solid, though some analyses and key interpretations would benefit from fuller justification.

      Thank you for handling our manuscript! We’ve modified the manuscript according to the reviewers’ comments and suggestions.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study demonstrates, through a series of EEG and MEG experiments, that the human brain automatically categorizes words from alphabetic and non-alphabetic languages, and it unpacks the neural mechanisms of this process from multiple angles. The work examines not only univariate repetition-suppression (RS) effects, but also how repeating or alternating languages influences the representational similarity of words within and across language categories.

      Strengths:

      The univariate RS effects across multiple experiments lend support to some of the main conclusions

      Weaknesses:

      I have reservations about the logic underlying the multivariate analyses, and I believe the implications of the control experiments merit fuller discussion.

      (1) Question 1: Logic of the multivariate analyses

      The original text states:

      "The processing of intra-language similarity was quantified as correlation distances between neural responses to two words of the same language, which occurred more frequently and would be inhibited in the Rep-Cond (vs. Alt-Cond) due to habituation (Fig. 1c)...".

      I argue that this passage conflates two levels. Building a representational dissimilarity matrix (RDM) is a data-analysis step; it cannot be equated with a cognitive computation. Hence, there is no sense in which this computation occurs "more frequently" in one condition. RDM construction rests on the pairwise similarity of activity patterns, so even if a task engaged no cognitive computation of representational similarity, we could still compute an RDM. Conversely, if a task factor alters the RDM, we must explain how that factor changes the underlying neural patterns, not claim that it triggers specific cognitive processing. Therefore, I neither understand what "more frequent processing" the authors refer to, nor accept their account of the multivariate results.

      The multivariate result pattern, briefly, is that distances between words, both within and across languages, are larger under the repetition condition. One plausible interpretation is that a word representation comprises two parts: language-type (alphabetic vs. non-alphabetic) and fine-grained identity features (visual shape, orthography, semantics, phonology, etc.). Repetition of language type may, via RS, reduce the weight of the first component, thereby increasing the relative contribution of fine-grained features and amplifying inter-word differences. This could explain the multivariate findings.

      Thank you for these insightful comments regarding the logic of the multivariate analyses. In the revision, we’ve elaborated the rationale underlying our experimental design. Specifically, we’ve explained why the processing of intra-language similarity is expected to occur more frequently in the repetition condition (Rep-Cond) than in the alternation condition (Alt-Cond) whereas the reverse is true for the processing of inter-language difference. Importantly, we’ve clarified that the processing of intra-language similarity was assessed rather than defined by conducting the multivariate analyses. The multivariate analyses were conducted to assess correlation distances between neural responses to pairs of words, either within the same language or across different languages. We explained what smaller intra-language correlation distances and larger inter-language correlation distances mean for language-base categorization of words (see Page 7-8).

      We appreciate the alternative account of the observed neural repetition suppression (RS) effects in terms of language-type versus fine-grained identity (visual shape, orthography, semantics, phonology, etc.) feature processing. We included a paragraph in the revised Discussion to discuss how possible the early neural RS effect can be attributed to the processing of the fine-grained identity features of visual words. This discussion allowed us to clarify that the early neural RS effects related to visual words of familiar and unfamiliar languages highlight the early spontaneous language-based categorization as a unique process of visual words of alphabetic and non-alphabetic languages. However, our results do not exclude the possibility that the processing of the linguistic properties of visual words may contribute to the long-latency RS effect (see Page 37-38).

      Page 7-8

      “The processing of intra-language similarity occurs when two words of the same language are perceived repeatedly with short interstimulus intervals. Because words of the same language were repeatedly presented in the Rep-Cond and words of two different languages were displayed in the Alt-Cond, the processing of intra-language similarity occurred more frequently and would be inhibited in the Rep-Cond (vs. Alt-Cond) due to habituation (Fig. 1c). By contrast, the processing of inter-language difference takes place when two words of different languages are perceived with short interstimulus intervals. Since words of different languages appeared more frequently in the Alt-Cond (vs. Rep-Cond), we would expect RS of the processing of inter-language difference in the Alt-Cond (vs. Rep-Cond). The neural processing of intra-language similarity was quantified as correlation distances between neural responses to two words of the same language whereas the neural processing of inter-language difference was assessed as correlation distances between neural responses to two words of two different languages. The correlation distances from the multivariate analyses were further employed to assess how words of one language are clustered and how far words of two languages are separated in a two-dimensional (2D) space during language-based word categorization. Enhanced language-based word categorization is associated with smaller intra-language correlation distances, which reflect more densely clustered words of the same language, and larger inter-language correlation distances, which manifest further separated words of two different languages.”

      Page 37-38

      “How possible are the early neural RS effects within 200 ms after word onset observed in our study related to the processing of low-level perceptual features or high-level linguistic (e.g., orthography, semantics, phonology) properties of visual words? Our analyses of the ERPs to scrambled Chinese and English words in Experiment 2 did not show significant RS effect. Because only low-level visual features were preserved in the scrambled words, the ERP results provided no evidence that the early RS effects on the neural response to words can be attributed to habituation of perception of the low-level perceptual features. Furthermore, we found that the RS effects on the neural response to radicals and letters in Experiment 3 took place in a delayed time window and exhibited different scalp distributions (i.e., over the central region for radicals and occipital regions for letters) compared with the neural RS effects related to words. Thus the early RS effects on the neural response to words cannot be interpreted as habituation of perception of the middle-level units of Chinese and English words (i.e., radicals and letters) either. In addition, the early neural RS effects were similarly observed for both familiar (i.e., Chinese and English) and unfamiliar (i.e., Korean and Italian) languages and occurred earlier than the time window in which the processing of the linguistic properties of visual words takes place (Marinkovic et al., 2003; Hodgson et al., 2021; Zhu et al., 2022). Therefore, the early neural RS effects identified in our work were unlikely to be associated with the processing of the linguistic (e.g., orthography, semantics, phonology) properties of visual words since these properties of unfamiliar languages were unknown to the participants. Taken together, our findings of the early neural RS effects highlight an early word-level representation of alphabetic vs. non-alphabetic languages which distinguishes words from letters/radicals but is similar for familiar or unfamiliar languages. Our results, however, do not exclude the possibility that the processing of the linguistic properties of visual words may contribute to the long-latency RS effect around 300 ms after word onset. Further processing of the linguistic properties of visual words of familiar languages may follow the early language-based categorization of visual words, though this should be tested in future research.”

      (2) Question 2:

      For unlearned languages, people cannot distinguish lexical from sub-lexical levels. What, then, determines (i) the RS-effect difference between letters and radicals in familiar languages and words in unlearned ones, and (ii) the similarity of repetition effects between words in unlearned and familiar languages? An explicit account is needed.

      Thank you for this suggestion. In the revised manuscript, we’ve included a dedicated paragraph addressing these two issues. Specifically, we’ve provided a more precise account of the differences in repetition suppression (RS) effects between words and letters/radicals in familiar languages, as well as the similar RS effects observed for unlearned and familiar languages. We believe that our findings of the early neural RS effects highlight an early word-level representation of alphabetic vs. non-alphabetic languages which distinguishes words from letters/radicals but is similar for familiar or unfamiliar languages (see Page 37-38).

      Page 37-38

      “How possible are the early neural RS effects within 200 ms after word onset observed in our study related to the processing of low-level perceptual features or high-level linguistic (e.g., orthography, semantics, phonology) properties of visual words? Our analyses of the ERPs to scrambled Chinese and English words in Experiment 2 did not show significant RS effect. Because only low-level visual features were preserved in the scrambled words, the ERP results provided no evidence that the early RS effects on the neural response to words can be attributed to habituation of perception of the low-level perceptual features. Furthermore, we found that the RS effects on the neural response to radicals and letters in Experiment 3 took place in a delayed time window and exhibited different scalp distributions (i.e., over the central region for radicals and occipital regions for letters) compared with the neural RS effects related to words. Thus the early RS effects on the neural response to words cannot be interpreted as habituation of perception of the middle-level units of Chinese and English words (i.e., radicals and letters) either. In addition, the early neural RS effects were similarly observed for both familiar (i.e., Chinese and English) and unfamiliar (i.e., Korean and Italian) languages and occurred earlier than the time window in which the processing of the linguistic properties of visual words takes place (Marinkovic et al., 2003; Hodgson et al., 2021; Zhu et al., 2022). Therefore, the early neural RS effects identified in our work were unlikely to be associated with the processing of the linguistic (e.g., orthography, semantics, phonology) properties of visual words since these properties of unfamiliar languages were unknown to the participants. Taken together, our findings of the early neural RS effects highlight an early word-level representation of alphabetic vs. non-alphabetic languages which distinguishes words from letters/radicals but is similar for familiar or unfamiliar languages. Our results, however, do not exclude the possibility that the processing of the linguistic properties of visual words may contribute to the long-latency RS effect around 300 ms after word onset. Further processing of the linguistic properties of visual words of familiar languages may follow the early language-based categorization of visual words, though this should be tested in future research.”

      Reviewer #2 (Public review):

      Summary:

      This study investigates how the human brain categorizes visual words from distinct writing systems (alphabetic vs. non-alphabetic) as a neural basis for the social-categorization function of language. Using a repetition suppression paradigm combined with electroencephalography and magnetoencephalography, the authors conducted nine experiments with independent participants to identify the neural network underlying language-based categorization, characterize its temporal dynamics, and test whether this process operates independently of linguistic properties such as semantic meaning and pronunciation.

      Strengths:

      (1) The study employs a well-validated design with clear control conditions and systematically manipulates key variables, including writing system, language familiarity, and native language background. The use of nine experiments with independent participant samples strengthens the reliability and replicability of the results.

      (2) The work combines EEG and MEG, cross-validating findings across imaging modalities to support the reported neural effects. A combination of univariate, multivariate, and connectivity analyses is used to characterize neural responses and network interactions.

      (3) Results are consistent across multiple language groups and for both familiar and unfamiliar languages, supporting the generalizability of the identified neural mechanism beyond specific languages or prior experience.

      Weaknesses:

      The authors provide compelling evidence that the identified neural network supports the categorization of words by language, including computations of intra-language similarity and inter-language difference. However, the conceptual framing of this finding as directly reflecting the social-categorization function of language may be premature. While the task captures spontaneous language categorization, it does not involve social evaluation or intergroup processes. The connection to social categorization is inferred from prior literature rather than demonstrated within the current experimental design. Clarifying this distinction would strengthen the conceptual precision of the manuscript.

      Thank you for this important comment. In the revised Introduction and Discussion, we’ve clarified several related issues. First, prior research suggests that language can serve as a socially relevant category cue. Second, these findings imply that rapid categorization of words by language may occur in the human brain. Third, although our results identify a neural network supporting such rapid language-based categorization of visual words, they do not directly test how this process relates to social categorization of people (see Page 3-4; Page 39). Highlighting these points help delineate the scope of our findings and point to important directions for future research.

      Page 3-4

      “The social-categorization function of language revealed in these behavioral studies implicates that rapid categorization of words of different languages may occur in the human brain. Furthermore, the findings of infant studies (e. g., Liberman et al., 2017b) suggest that the neural process involved in categorization of words of different languages may develop even prior to the processing of linguistic properties (e.g. semantic meanings) of words. Nevertheless, up to date, there has been little neuroimaging research examining the neural mechanisms underlying automatic and fast categorization of words of different languages.”

      Page 39

      “Finally, it should be noted that the current work was initiated by the previous behavioral findings which suggest that language can serve as a socially relevant category cue but focused on the neural mechanisms underlying rapid language-based categorization of visual words. Although the previous findings suggest that the language-based categorization of visual words provides a cognitive basis of social categorization of people, our work did not directly test whether and how the neural processes involved in the language-based categorization of visual words are linked to social evaluation or intergroup processes which are critical for social categorization of people. To clarify this issue should promote deep comprehension of the neural mechanisms underlying the social-categorization function of language but is beyond the scope of the current study. Future research should investigate the connection between language-based categorization of words and social categorization based on other social cues (e.g., faces), which is pivotal to understanding of social interactions in real-world situations.”

      Recommendations for the authors:

      Reviewer #2 (Recommendations for the authors):

      (1) Revise the conceptual framing to clarify the relationship between the experimental results and the proposed social-categorization function of language. If the authors wish to retain the emphasis on social categorization in the title or discussion, they should explicitly explain how the observed neural mechanisms of language-based word categorization link to social evaluation, intergroup processes, or real-world social categorization. This clarification would strengthen the conceptual coherence and justify the use of social categorization within the current study's scope.

      Thank you for this and the following suggestions. In the revised Introduction and Discussion, we’ve clarified the following point: First, the findings of prior behavioral studies suggest a social-categorization function of language. Second, based on these behavioral findings, we predicted automatic and fast categorization of words by language. Our study tested this prediction using neuroimaging and investigated the neural mechanisms of language-type-based categorization of visual words. This is the main goal of our work. Third, to examine how the observed neural mechanisms of language-based word categorization link to social evaluation, intergroup processes, or real-world social categorization is important but beyond the scope of the current work. However, this is a very important question. Future research should test the connection between the neurocognitive processes involved in social categorization of people and the neural categorization of visual words by language revealed in our study. Consistently, the title of our paper “Neural categorization of visual words of alphabetic and non-alphabetic languages” and Discussion focus on contributions of our findings to understanding of the neural categorization of visual words by language rather than its connection to social categorization of people. Above all, we’ve clarified in the revision that our study was initiated by the findings of social function of language but was limited to the neural processing of visual words (see Page 3-4; Page 39). Thanks again for this comment.

      Page 3-4

      “The social-categorization function of language revealed in these behavioral studies implicates that rapid categorization of words of different languages may occur in the human brain. Furthermore, the findings of infant studies (e. g., Liberman et al., 2017b) suggest that the neural process involved in categorization of words of different languages may develop even prior to the processing of linguistic properties (e.g. semantic meanings) of words. Nevertheless, up to date, there has been little neuroimaging research examining the neural mechanisms underlying automatic and fast categorization of words of different languages.”

      Page 39

      “Finally, it should be noted that the current work was initiated by the previous behavioral findings which suggest that language can serve as a socially relevant category cue but focused on the neural mechanisms underlying rapid language-based categorization of visual words. Although the previous findings suggest that the language-based categorization of visual words provides a cognitive basis of social categorization of people, our work did not directly test whether and how the neural processes involved in the language-based categorization of visual words are linked to social evaluation or intergroup processes which are critical for social categorization of people. To clarify this issue should promote deep comprehension of the neural mechanisms underlying the social-categorization function of language but is beyond the scope of the current study. Future research should investigate the connection between language-based categorization of words and social categorization based on other social cues (e.g., faces), which is pivotal to understanding of social interactions in real-world situations.”

      (2) Clarify the consistency between the reported model order (5 ms lag) and the sampling rate after downsampling (250 Hz, corresponding to 4 ms per time point). If a discrepancy exists, clearly explain how the time-series data were processed.

      We clarified in the revision (see Page 53) that “because down-sampling was not applied to the GCA analyses, a 5-ms lag was used for prediction of the neural activity in one brain region using the neural activity in another brain region”.

      (3) For the representational similarity analysis (RSA), report reliability measures for the representational dissimilarity matrices (e.g., split-half reliability) to verify that the observed effects are stable given the number of trials per condition.

      Following this suggestion, we’ve conducted split-half reliability analyses and reported the results in the revised supplementary materials. The reliability analyses are also mentioned in the revised Discussion (see Page 40).

      Page 40

      “In conclusion, our EEG and MEG results revealed robust RS effects in the early neural responses to visual words of the same language. The reliability of these RS effects was confirmed across words of different familiar and unfamiliar languages, in samples of speakers with different native languages, and through split-half reliability analyses (see Supplementary Materials, Fig. S19). These effects were supported by the bilateral neural networks whose activity reflected computations of correlation distances between word pairs, capturing both intra-language similarity and inter-language differences during the categorization of visual words in alphabetic and non-alphabetic languages. Together, these findings advance our understanding of spontaneous, language-based neural categorization of visual words as a key basis of the social-categorization function of language.”

      (4) Provide complete statistical information for all significant results reported in the supplementary materials, including relevant test statistics (e.g., t-values, cluster p-values) in figure legends or a supplementary results table to improve transparency.

      Complete statistical information has been provided in the revised supplementary materials (see Tables S4 and S5).

      (5) Streamline the presentation of the nine experiments in the main text to emphasize the core conceptual and methodological logic, potentially using a schematic overview or flowchart to improve readability.

      As suggested, we’ve included an overview of the nine experiments in the revised Introduction. This overview helps understanding of the core conceptual and methodological issues in our work (see Page 6).

      Page 6

      “In nine experiments we recorded EEG/MEG signals from Chinese, English, and German speakers when viewing words of an alphabetic language and a non-alphabetic language (English and Chinese words, or Italian and Korean words) or of two alphabetic languages (English and German) in the Rep-Cond and Alt-Cond. We recorded EEG signals from Chinese participants to examine temporal neural dynamics of spontaneous language-based word categorization in Experiment 1. The similar paradigm was employed in Experiments 2 and 3 to investigate whether perceptual features or radical/letters of words are sufficient to generate spontaneous language-based categorization of visual words. The results in Experiment 1 were replicated in native English and German speakers in Experiments 4 and 5, respectively. Neural dynamics of categorization of words of two unlearned languages were further investigated in Chinese participants in Experiment 6. Finally, the neural networks supporting the spontaneous categorization of words of two learned or unlearned languages were localized using MEG in Chinese and English speakers in Experiments 7-9, respectively.”

      (6) Strengthen the transition between the discussion of the social-categorization function of language and the neural mechanisms of visual word categorization in the introduction.

      Following this suggestion, we’ve modified the Introduction to strengthen the transition between the discussion of the social-categorization function of language and research on neural mechanisms of visual word categorization (see Page 3-4).

      Page 3-4

      “The social-categorization function of language revealed in these behavioral studies implicates that rapid categorization of words of different languages may occur in the human brain. Furthermore, the findings of infant studies (e. g., Liberman et al., 2017b) suggest that the neural process involved in categorization of words of different languages may develop even prior to the processing of linguistic properties (e.g. semantic meanings) of words. Nevertheless, up to date, there has been little neuroimaging research examining the neural mechanisms underlying automatic and fast categorization of words of different languages.”

      (7) Briefly define the repetition suppression (RS) paradigm when first mentioned (i.e., reduced neural response to repeated stimuli from the same category, reflecting categorical processing) to improve accessibility for non-specialist readers.

      The RS paradigm is now defined in Introduction when being mentioned for the first time in the manuscript (see Page 5-6).

      Page 5-6

      “The present study investigated neural dynamics of categorization of visual words of two different (an alphabetic versus a non-alphabetic, or two different alphabetic) languages by combining EEG/MEG with a repetition suppression (RS) paradigm adopted from previous studies of social categorization of faces (Zhang et al., 2023b; Zhou et al., 2020). RS refers to the attenuation in neural responses to a repeated occurrence of stimuli that engage common neuronal populations or processes due to habituation (Grill-Spector et al., 2006). The RS paradigm consisted of an alternating condition (Alt-Cond), in which visual words of two different languages were presented alternately, and a repetition condition (Rep-Cond), in which words of one language were presented repeatedly (Fig. 1a). Neural responses to stimuli of the same category were attenuated in the Rep-Cond compared to Alt-Cond due to habituation and this RS effect disentangles the neural activities underlying categorization of faces and body silhouettes of a specific social group.”

      (8) Report detailed participant demographic information, including exact age range/mean age and gender ratio for each experiment, to meet standard reporting practices in neuroscience.

      We’ve modified Table S1 to include the information about exact age range/mean age and gender ratio in each experiment.

      (9) Correct minor typographical and grammatical errors, including These finding (line 59) and Chinse (line 223).

      These and other grammatical errors have been corrected in the revision.

    1. Author response:

      The following is the authors’ response to the original reviews

      Summary of revision for all referees:

      We thank referees for their constructive comments. To address their concerns, we now performed additional statistical analyses integrating both paired and unpaired data, performed positive controls for comparisons between NH- and CI- evoked iEEG measurements, developed tools for measuring and collected new experimental data on forward masking ECAP measurements in CI implanted rats (N=3), and reworked both manuscript text and figures to improve clarity. These most significant changes are summarized here, and a complete list of responses to reviewers and corresponding changes will follow.

      Summary of major changes to revised manuscript:

      (1) Statistical treatment of paired vs unpaired recordings using mixed-effects models (updates to all manuscript figures that compare NH vs CI); this largely confirmed the results reported in our original submission.

      (2) New analysis, controlling for information-theoretic cross-modality comparison (i.e., training with tone- and testing with cochlear implant-evoked iEEG measures, Fig. 8).

      (3) Clarification of methods (Supplemental Fig. 2 & manuscript text)

      (4) Additional experiments testing peripheral tuning of our 8-channel CI rodent model via forward masking ECAP measures across 3 animals (N=3, Supplemental Fig. 1)

      (5) Detailed response addressing robustness of tonotopy in NH and CI animals

      Public Reviews:

      Reviewer #1 (Public Review):

      Strengths:

      The study poses a timely, clinically relevant question with clear implications for CI strategy. The analytical toolkit is appropriate: µECoG captures mesoscale patterns; TCA offers a transparent separation of spatial and temporal structure; and mutual-information decoding provides an interpretable measure of single-trial discriminability. Within-subject recordings in a subset of animals, in principle, help isolate modality effects from inter-animal variability. Where analyses are most direct, the acoustic condition yields higher single-trial decoding accuracy, which is a meaningful and clearly presented result.

      We appreciate the comments on the strengths of our analytic approaches.

      Weaknesses:

      Parts of the statistical treatment do not match the data structure: some comparisons mix paired and unpaired animals but are analysed as fully paired, raising concerns about misestimated uncertainty.

      Please see our response to specific comment #2 above. In short, we agree with this critique of our original analyses, and in our revised manuscript we re-analyzed all NH vs. CI comparisons using linear mixed effects models that incorporate both paired and unpaired observations within a single framework. This allows us to include all animals, account for within-animal dependence for paired experiments (normal hearing and cochlear implant data from the same animal when available), and to align the statistical tests with the data shown in the figures. In almost every case, the mixed effects models confirm our original conclusions. Two comparisons that were previously nonsignificant now reach criterion for statistical significance (Fig. 2E, p=0.048 and Fig. 6F, p=0.027). We updated the manuscript to report these values and to clarify the use of mixed effects modeling in the methods under the section titled, “Linear mixed effects modeling.”

      Methodological reporting is incomplete in places; essential parameters for both acoustic and electrical stimulation, as well as objective verification of implantation and deafening, are not described with sufficient detail to support confident interpretation or replication.

      Please see our response to comment #5 below. We have revised our manuscript to now include this information in the methods.

      Figure-level clarity also undermines the message. In Figure 2, non-significant slopes for CI, repeated identification of a single "best channel," mismatched axes, and unclear distinctions between example and averaged panels make the assertion of spatial organisation unconvincing; importantly, the normal-hearing panels also do not display tonotopy as clearly as expected, which weakens the key contrast the paper seeks to establish.

      This is an important point, thanks- please see responses to comment #1 above. We note that conventional tonotopic maps in auditory cortex are characteristic frequency maps, i.e., maps of topographic organization for responses to lowest-threshold stimuli (often presented around 20-50 dB SPL). Our maps were constructed from stimuli presented at 70 dB SPL, thus blunting crisp tonotopy to some degree. Furthermore, we quantified spatial organization using a previously published method from the Polley lab (Romero & Hight et al. 2020), in which local tonotopic gradient vectors (magnitude and direction) were computed from GCaMP responses at each pixel and projected onto a unit circle. Mean vector strength across all pixels was then compared to a shuffled distribution as a measure of tonotopic organization. We applied the same procedure to our iEEG best-frequency and best-channel maps. Both map types yielded mean vector strengths that were substantially larger than those derived from shuffled maps (p < 10<sup>-10</sup>), indicating that our maps have a consistent tonotopic (for BFs) or cochleotopic (for CI channels) organization that is highly unlikely to arise by chance. This is now included in our revised manuscript.

      Finally, the decoding claims would be strengthened by simple internal controls, such as within modality train/test splits and decoding on raw ERP/high-gamma features to demonstrate that poor cross-modal transfer reflects genuine differences in the underlying responses rather than limitations of the modelling pipeline.

      Please see our response to comment #12 below. In short, we have now included this analysis in revised Figure 8.

      Reviewer #2 (Public Review):

      Strengths:

      The study includes interesting analyses of the sound and cochlear implant representation structure based on decoders.

      We appreciate the comment on how interesting our analyses are, thanks!

      Weaknesses:

      The observation that responses to cochlear implant stimulation (stimulation) are spatially organized is not new (e.g., Adenis et al. 2024).

      We agree that it is not particularly novel to report that there is spatial organization to cochlear implant stimulation. However, we believe that our direct comparisons (when possible, within animal) between normal-hearing and cochlear implant modality maps is unusual in the literature, including asking how decoders based on one set of responses might apply to responses evoked from the other modality. Adenis et al. (2024) is a fantastic study of pulse shape and monopolar vs bipolar stimulation modes with a 6-channel implant in guinea pig, but as far as we can tell this study does also not compare normal hearing maps prior to deafening and implantation to the cochlear implant maps in the same animals.

      The claim that spatial and temporal dimensions contribute information about the sound is also not new; there is a large literature on this topic. Moreover, the results shown here are extremely weak. They show similar levels of information in the spatial and temporal dimensions, and no synergy between the two dimensions. This is however, likely the consequence of high measurement noise leading to poor accuracy in the information estimates, as the authors state.

      Good point, please see our response to comment #1 below.

      The main claim of the study - the mismatch between cochlear implant and sound representation - is not supported. The responses to each modality are measured in different animals. The authors do not show that they actually can compare representations across animals (e.g., for the same sounds). Without this positive control, there is no reason to think that it is possible to decode from one animal with a decoder trained on another, and the negative result shown by the authors is therefore not surprising.

      Good point, thanks- please see our response to comment #2 below, where we describe this new control we have added.

      Reviewer #3 (Public Review):

      Strengths:

      The model combining micro-eCoG and cochlear implantation and the methodology to extract both the Event Related Potentials (ERPs) and High-Gammas (HGs) is very well designed and appropriately analyzed. Likewise, the PCA-LDA and TCA-LDA are powerful tools that take full advantage of the information provided by the cortical ensembles. The overall structure of the paper, with a paced and exhaustive progress through each step and evolution of the decoder, is very appreciable and easy to follow. The exploration of single-trial encoding and stimulus identity through temporal and spatial domains is providing new avenues to characterize the cortical responses to CI stimulations and their central representation. The fact that single trials suffice to decode the stimulus identity regardless of their modality is of great interest and noteworthy. Although the authors confirm that iEEG remains difficult to transpose in the clinic, the insights provided by the study confirm the potential benefit of using central decoders to help in clinic settings… the reviewer wants to reiterate that the study proposed by Hight et al. is well constructed, relevant to the field, and that the overall proposal of improving patient performances and helping their adaptation in the first months of CI use by studying central responses should be pursued as it might help establish new guidelines or create new clinical tools.

      We thank the Reviewer for the positive comments about the thoroughness of our analyses and clear organization of our manuscript.

      Weaknesses:

      The conclusion of the paper, especially the concept of distinct cortical encoding for each modality, is unfortunately partially supported by the results, as the authors did not adequately consider fundamental limitations of CI-related stimulation. First, the reviewer assumed that the authors stimulated in a Monopolar mode, which, albeit being clinically relevant, notoriously generates a high current spread in rodent models.

      Thanks, this is an important potential concern. Please see our response to comment #5 of Referee 1 and responses to comment #3 below. We agree that monopolar stimulation would be expected to be less spatially specific than bipolar or multipolar modes. However, we chose monopolar stimulation because it is the main clinical configuration in human CI users and therefore most relevant for translational purposes. For our revised manuscript, we made new ECAP measurements of peripheral (spatial and temporal) tuning via a forward masking paradigm and demonstrate that monopolar is effectively tuned (Supplemental Fig. 2). Together with additional single-animal maps in Supplementary Figure 3, together with our vector-strength analysis (Response Fig. 2), demonstrate that even under acute monopolar stimulation we observe structured cochleotopic organization in cortex, rather than the extremely low-pass patterns one might expect if monopolar spread was a major contaminant.

      Second, comparing the averaged BF maps for iEEG (Figure 2A, C), BFs ranged from 4 to 16kHz with a predominance of 4kHz BFs. The lack of BFs at higher frequencies hints at a potential location mismatch between the frequency range sampled at the level of the cortex (low to medium frequencies) and the frequency range covered by the CI inserted mostly in the first turn-and-a-half of the cochlea (high to medium frequencies). Looking at Figure 2F (and to some extent 2A), most of the CI electrodes elicited responses around the 4kHz regions, and averaged maps show a predominance of CI-3-4 across the cortex (Figure 2C, H) from areas with 4kHz BF to areas with 16kHz BF. It is doubtful that CI-3-4 are located near the 4kHz region based on Müller's work (1991) on the frequency representation in the rat cochlea.

      Please see our responses to comment #3 below.

      Taken together with the Pearsons correlations being flat, the decoder examples showing a strong ability to identify CI-4 and 3 and the Fig-8D, E presenting a strong prediction of 4kHz and 8kHz for all the CI electrodes when using a pure tone trained decoder, it is possible that current spread ended stimulating indistinctly higher turns of the cochlea or even the modiolus in a non-specific manner, greatly reducing (or smearing) the place-coding/frequency resolution of each electrode, which in turn could explain the coarse topographic (or coarsely tonotopic according to the manuscript) organization of the cortical responses. Thus, the conclusion that there are distinct encodings for each modality is biased, as it might not account for monopolar smearing. To that end, and since it is the study's main message and title, it would have benefited from having a subgroup of animals using bipolar stimulations (or any focused strategy since they provide reduced current spread) to compare the spatial organization of iEEG responses and the performances of the different decoders to dismiss current spread and strengthen their conclusion.

      Please see our responses to comment #4 below as well as our responses related to monopolar vs bipolar stimulation. We agree that for future studies, it will be important to do a heads-on comparison of the differences between bipolar and monopolar stimulation depending on electrode location and stimulation intensity.

      Recommendations for the authors:

      Reviewer #1 (Recommendations for the authors):

      We thank the reviewer for commenting on the strengths of our manuscript, including appreciating the power and timeliness of our approach.

      (1a) Figure 2 does not convincingly support the claim that "tone-evoked and CI-evoked iEEG measurements are spatially organized," particularly for CI data: Figure 2C repeatedly highlights the same "best channel," and the slopes in Figures 2B and 2G are non-significant; there are also discrepancies between panels (A vs. C, F vs. H) and mismatched frequency ranges (0-16 kHz vs. up to 32 kHz), which should be clarified as exemplar versus averaged displays and harmonized in scale.

      (First we note that Reviewer 3 also raised related concerns about the robustness of tonotopy in our iEEG data.) We address these by comparing our maps to previously published tonotopic maps, and using an established quantitative analysis of tonotopic strength from Romero & Hight et al. (2020).

      First, to place our tone-evoked iEEG maps in context, we overlaid them on the same spatial scale and orientation as both single-unit tonotopy in rat primary auditory cortex (A1) from Polley et al. (2006) and iEEG maps obtained with the same surface array in Insanally et al. (2016). The rostral–caudal and dorsal–ventral axes and cortical extents are matched across panels. Our best-frequency maps (Figure 2C) qualitatively recapitulate the high-to-low frequency gradient and spatial layout reported in both of these prior studies, supporting our claim that tone-evoked iEEG captures canonical mesoscale tonotopy. We have updated the manuscript results section to directly reference these two studies, “The area and orientations of tone-evoked maps qualitatively match those published from single unit recordings (Polley et al. 2006) and published using similar iEEG arrays (Insanally et al. 2016).”

      Second, to quantify tonotopy in a way that is directly comparable to previous work, we reproduced the analysis of Romero & Hight et al. (2020), who examined tone-evoked GCaMP signals (Romero & Hight et al. (2020)). In that paper, local tonotopic gradient vectors (magnitude and direction) were computed at each pixel and projected onto a unit circle; the mean vector strength across all pixels was then compared to a shuffled distribution as a measure of tonotopic organization. We applied the same procedure to our iEEG best-frequency and best-channel maps (Fig. 2C-E). Both map types yielded mean vector strengths that were substantially larger than those derived from shuffled maps (p < 10<sup>-10</sup>), indicating that our maps have a consistent tonotopic (for BFs) or cochleotopic (for CI channels) organization that is highly unlikely to arise by chance. We cite this paper for these analyses related to Figure 2.

      (1b) Figure 2C repeatedly highlights the same ‘best channel’

      We agree that many CI-evoked maps are dominated by a single channel, as seen in our exemplar and in the additional animals shown in new Supplemental Fig. 3. In Fig. 2C, channel 5 emerges as the dominant best channel, as CI-evoked activity in this animal is broad and is strongest for channel 5 (Fig. 2A). This reflects a feature of iEEG signals rather than a plotting artifact. Biophysically, iEEG reflects spatially summed local field potentials that low-pass filter underlying neural activity; these far-field signals aggregate excitatory and inhibitory processes and are not expected to show the sharp single-neuron tuning seen in spike recordings. As a result, broad peaks centered on the most strongly driven channels are expected. We have added text in the results section discussing these limitations, overall maps reduced from iEEG responses were similar in size and orientation compared to single unit maps, “albeit at coarser gradients likely due to aggregate recordings of excitatory and inhibitory activity and low-pass filtering due to potentials originating far from recording sites.” We also added in the results section the comparison of spatial correlations (Fig. 2B,G) at the extremes of stimulus separation “electrode separations (CI 1 vs ≥5 electrodes, ERP: p=0.01, HG: p=0.04)” as analyzed by linear mixed effects models.

      (1c) Mismatched frequency ranges

      We constricted the range of frequencies plotted in some panels (e.g., Fig. 2C from 1.4-32 kHz to 1.4-16 kHz) to emphasize the compressed range of tonotopic gradients and patterns.

      (1d) The slopes in Figures 2B and 2G are non-significant

      We agree that non-significant group-level slopes indicate that CI-evoked tonotopy is weaker than tone-evoked tonotopy, and we now emphasize this point. At the same time, the data exhibit systematic structure: for both ERP and HG, mean spatial correlations decline monotonically with increasing CI channel separation (Fig. 2B,G). We also directly compared spatial correlations at the extremes of stimulus separations (1 vs. ≥5-channel separation) and found a significant difference. This is updated in the manuscript as: “At the extremes, the spatial correlations were always higher for small vs. large tone separations (NH 0.5 vs ≥3.5 octaves, ERP: p<10<sup>-4</sup>, HG: p<10<sup>-4</sup> Student’s one-tailed t-test) and electrode separations (CI 1 vs ≥5 electrodes, ERP: p=0.01, HG: p=0.04).”. Together with the strong deviation from shuffled maps in the vector-strength analysis (Fig. 2E), we argue that analysis of spatial correlations indicates that CI-evoked maps are not random but reflect a coarse underlying gradient. In addition, as tone-evoked maps exhibit tonotopy, we asked if CI stimulation itself is at least spatially tuned in the periphery. Using ECAPs with a forward-masking paradigm (new Supplemental Fig. 1), we show that probe-evoked ECAPs are significantly more suppressed by adjacent than by distant maskers (N = 3), demonstrating functional spatial tuning of CI electrodes in the cochlea. We have also replotted these results in comparison with the same measurements from a human CI user (Author response image 1). This supports the interpretation that peripheral input is spatially specific and that the weaker cortical cochleotopy likely reflects the properties and resolution of iEEG and acute CI stimulation rather than a complete absence of spatial organization. Overall, the new comparative figures and analyses are intended to make transparent that (i) iEEG robustly captures tonotopy for acoustic tones, and (ii) CI-evoked CI-evoked responses exhibit coarser, but statistically non-random, cochleotopic organization.

      Author response image 1.

      Here, we compare data from the new Supplemental Figure 1C,D with human data (N=1) for spatial & temporal tuning in the periphery, as assessed by forward masking ECAP measurements. A) Spatial tuning functions were averaged across all probe electrodes and 3 animals (left) and 1 human subject (right) (black, mean; gray: s.e.m..; orange, average of individual subjects). B) Temporal tuning functions were averaged across all probe electrodes and 3 animals (left) and 1 human subject (right) (black, mean; gray, s.e.m.; orange, average of individual subjects). Note: human subject is the first-author, a long-term cochlear implant user (>10 years) with significant open set speech perception.

      (2) The statistical approach is inappropriate where pairing is incomplete: a Student's paired two-tailed t-test is used despite not all data being paired; a linear mixed-effects model would be more suitable, whereas an unpaired test risks reduced power.

      We agree with this suggestion. As the reviewer notes (also raised by Reviewer 3), our original analyses did not fully exploit the partially paired structure of the data. In the initial submission we used paired t-tests when animals contributed both normal-hearing (NH) and CI measurements, which meant that animals with only NH or only CI data were excluded from those tests.

      To address this, we have re-analyzed all NH vs. CI comparisons using linear mixed-effects models that incorporate both paired and unpaired observations within a single framework. This approach allows us to (i) include all available animals, (ii) appropriately account for within-animal dependence when both conditions are present, and (iii) align the statistical tests with the data shown in the figures. In nearly all cases, the mixed-effects models confirm our original conclusions. Two comparisons that were previously non-significant are now significant in the positive direction: Fig. 2E (p = 0.048) and Fig. 6F (p = 0.027, linear mixed-effects models). We have updated the manuscript to report these values and to clarify the use of mixed-effects modeling in the methods under the section titled, “Linear mixed effects modeling.”

      (3a) Given the surgical complexity, objective verification of implantation and deafening is needed (e.g., eABRs for implant function and post-deafening ABR thresholds)”

      We agree that objective verification of both implant placement and deafening is critical, particularly given the surgical complexity of multichannel CI implantation in rats. Note that we previously extensively documented deafness in our cochlear implant rats with eABRs, histology of hair cell counts, and behavior (turning the implant off and seeing performance drop to chance). As we argued in Glennon et al. Nature 2023, the primary outcome measure and definition of deafness is behavioral, as anatomical and physiological markers are correlates of functional deafness but ultimately deafness must be defined in terms of behavioral performance. This is described in more detail below.

      We agree that objective verification of both implant placement and deafening is critical, particularly given the surgical complexity of multichannel CI implantation in rats. Note that we previously extensively documented deafness in our cochlear implant rats with eABRs, histology of hair cell counts, and behavior (turning the implant off and seeing performance drop to chance). As we argued in Glennon et al. Nature 2023, the primary outcome measure and definition of deafness is behavioral, as anatomical and physiological markers are correlates of functional deafness but ultimately deafness must be defined in terms of behavioral performance. This is described in more detail below.

      Implant placement: Our primary concern during surgery is to ensure that the CI array is correctly positioned along the cochlear spiral toward the apex. As shown in Author response image 2, once the bulla is opened and the cochleostomy is made at the junction of the temporal bone and the stapedial artery, the orientation of the cochlear spiral is clearly visible under the surgical microscope. We advance the 8-channel array only in the apical direction, and we require that all 8 electrodes pass through the cochleostomy. A complete insertion of all 8 electrodes cannot be achieved with a basal-ward trajectory, so full insertion provides a strong anatomical confirmation that the array is directed apically. The white band on the array, visible just basal to the cochleostomy (Author response image 2), serves as a consistent visual marker of complete insertion. We have added text and this figure to the Methods to clarify these criteria, “We required that all eight electrodes pass through the cochleostomy, confirming that the array was inserted in the direction of the apex.”

      Verification of deafening: We also share the reviewer’s concern about confirming profound hearing loss, particularly because some CI animals were presented acoustic tones to drive individual channels. We used the same mechanical-only deafening procedure described and validated in our previous work (King et al., 2016; Glennon et al., 2023), which was chosen to minimize systemic side-effects and maximize post-surgical survival, validated in three ways:

      - Histology: In N=4 deafened animals, inner hair cell loss was ~50% and outer hair cell loss was near complete at almost 100% in all animals.

      - Physiology: For N=14 rats, acoustic ABRs were substantial before deafening but statistically similar to baseline noise after deafening.

      - Behavior: For N=16 deafened rats, behavioral performance with implant on was d′: 1.7±0.1, but when implant was turned off in a subset of sessions, performance dropped to chance (d′: −0.05±0.1, P < 0.0001).

      Author response image 2.

      Visual confirmation of a successful electrode insertion. The direction of an 8-channel array being implanted toward the apex is clear under microscope. Full insertion of all 8 channels is further confirmed by the white band’s (located after basal electrode) proximity to the cochleostomy.

      This combination of histological, physiological, and behavioral evidence indicates that the mechanical-only deafening protocol produces profound hearing loss, with no functionally relevant residual hearing at intensities equal to or greater than those used in our study (70 dB SPL). Given this prior validation under identical surgical and experimental conditions, we are confident that our CI animals were effectively deafened and that the iEEG responses we report are driven by the implant rather than by residual acoustic hearing. We now clarify this in the Methods and explicitly cite our validation: “(mechanical only, as described and validated in Glennon et al. 2023).

      (3b) One CI animal did not learn the task (Fig. 1C), potentially reflecting implantation efficacy.

      Good point, thanks. For both humans and rats, cochlear implant performance can be highly variable, reflecting a number of factors in terms of device performance, training efficacy and motivation, or other technical or biological sources of heterogeneity. We note however that not all animals included in this study were behaviorally trained, and wanted to show the full range of variable performance for the subset of animals that were trained (N=4 typical hearing and N=3 cochlear implant rats, one of the 4 trained animals lost the implant before it could be re-trained on the cochlear implant version of the task). We now highlight this range of performance variability in the results section and explain why N=4 normal-hearing and N=3 cochlear implant rats.

      (4) The behavioural paradigm and cohort accounting are unclear: Figure 1C shows four NH-trained rats, yet subsequent analyses include only two NH-trained animals, which is confusing.

      We have now clarified the relation between the behavioral cohort and the iEEG cohort in the revised manuscript. The key point is that the animals in Figure 1C are defined by their behavioral training history (NH vs CI training), whereas inclusion in the iEEG analyses is defined by the specific stimuli collected during acute recordings, and these two categorizations are not always the same. In total, four rats underwent both iEEG recordings and behavioral training. Of these four, three were subsequently deafened, implanted with chronic CIs, and trained on the CI-driven task (Fig. 1C). With respect to the acute iEEG experiments, we obtained tone-only iEEG in 1 animal, CI-only iEEG in 2 animals, and both tone- and CI-evoked iEEG in 1 animal.

      Thus, the “NH-trained” label in Figure 1C refers to behavioral training status, not to the stimulus conditions used during iEEG recordings. All iEEG measurements were acute and performed immediately after surgery (for CI animals) or in the normal-hearing condition, before any CI behavioral training. Consequently, the behavioral cohort in Figure 1C is larger than the subset of animals that contributed to specific iEEG contrasts in later figures, which explains why some panels include only two NH animals.

      To clarify this, we have added a new Supplementary Figure 2 that provides a timeline for each animal, indicating when behavioral training occurred, when deafening and implantation occurred, and which stimulus conditions (tones vs CI) were used for each iEEG recording. We kept this figure in the Supplementary section because the focus of the manuscript is on evoked iEEG measurements rather than behavior, but the revised text now explicitly refers to this schematic when describing the cohorts “The combinations of animals that underwent behavioral training and acute iEEG measurements are shown in Supplemental Fig. 2.”

      (5) Methods lack essential details: specify acoustic stimulus types and intensities, CI stimulation parameters (e.g., current/charge per phase, phase width, rate, loudness setting), and the recording state (awake vs. anaesthetised), which is only implied in the discussion.

      We agree that these details are essential, and Reviewer 3 raised similar concerns about methodological clarity. We have now expanded the Methods to specify the acoustic stimuli, CI stimulation parameters, and recording state.

      Acoustic stimuli: We now describe the acoustic stimulus set in the Methods, which references Insanally et al. (2016). Briefly, tones were pure sinusoids spanning frequencies from 1.4 to 32 kHz (half octave spaced), presented at 70 dB SPL with a duration of 50 ms with 2ms cosine-squared ramps and at a pseudorandom sequence of 1.25 Hz. These parameters are now updated in the methods under “Stimulus presentation for cortical sensory mapping in normal hearing rats.”

      CI stimulation parameters: CI stimulation used standard clinical-style monopolar mappings. We now specify in the Methods that pulses were biphasic, charge-balanced, with 8 µs interphase gaps and 25 µs /phase (total pulse width = 58 µs); stimulation rate was 900 pulses per second (pps); and current amplitude (and thus charge per phase) was set individually for each electrode based on its ECAP threshold. All stimulation levels were within normal and safe limits: charge densities remained below the Shannon limit and within the electrochemical “water window.”

      Loudness setting: In this study, CI stimuli were presented primarily at a single level—each electrode was stimulated at its ECAP threshold level for the tone-to-CI mapping experiments. We have added these details in the methods under the “Stimulus presentation for cortical sensory mapping in cochlear implanted rats” subsection.

      Recording state: All iEEG recordings reported in the manuscript were acute and performed under anesthesia. This is now stated explicitly at the start of the Methods section.

      (6) Plasticity and training effects warrant further consideration: although the manuscript reports no difference between naïve and trained rats, Figure 3 suggests greater across-trial variability for CI than NH that is not evident in the trained subset; examining relationships among behavioural performance, decoder performance, across-trial variability, and training duration would strengthen interpretation.

      We agree that plasticity and training effects are central questions for cochlear implant research and that iEEG is well suited to study how cortical representations evolve with CI use. However, the current dataset was collected mainly to compare cortical encoding of acoustic versus CI stimulation under matched, acute conditions (not necessarily after behavioral training with the implant, and we note that most studies of physiological responses to cochlear implant function in non-human species also do not incorporate aspects of training). All CI-evoked iEEG recordings were obtained immediately after implantation, before any CI-based behavioral training. As a result, any training effects reflected in the iEEG data can only arise from prior normal-hearing training, not from experience with CI stimuli themselves. Only a small subset of animals (N = 3 of 10) underwent behavioral training with cochlear implants, and their training histories (duration, performance levels, CI hardware status) are not uniform. This yields insufficient statistical power to meaningfully examine correlations among behavioral performance, decoder performance, across-trial variability, and training duration. While we note the reviewer’s observation that across-trial variability appears qualitatively different in the small, trained subset, we do not believe the current data justify strong conclusions about training-related plasticity.

      (7) Differentiating the CI rats stimulated directly or through the microphone of the speech processor -at least in the figures - would be useful to allow the reader to assess whether both stimulation strategies give rise to similar results.

      We agree that it is important to distinguish between rats stimulated directly via CI hardware and those stimulated acoustically through a speech processor. We now show in new Supplementary Figure 2, which animals received direct electrical stimulation and which were driven acoustically through the processor microphone. We also now plot tonotopic and cochleotopic maps for all CI animals in Supplementary Figure 3, with the stimulation mode indicated for each animal. As discussed in our response to comment #2 of Reviewer 3, we also provide validation that acoustic tones can be used to selectively drive individual electrodes via the speech processor. However, the sample sizes for the two stimulation strategies are small (N = 4 rats with direct CI stimulation, N = 3 rats with acoustic CI stimulation). For this reason, we have chosen not to draw strong statistical conclusions about differences between direct vs acoustic CI stimulation in the present manuscript.

      (8) Typographical error at the end of the introduction ("To this end we have designed and manufactured..."), and in the first paragraph of the Discussion ("...that both that...").”

      Thanks, we have updated the manuscript accordingly.

      (9) Inconsistent terminology: use a single form (e.g., "normal-hearing") throughout.

      Good suggestion, thanks. We have updated all main manuscript to only use normal-hearing. We found and changed two instances in which we used the acronym NH in lieu of normal-hearing, once early in the results section and once in the legend for Figure 3.

      (10) In Figure 3D (temporal), there appears to be an extra data point for the NH-trained group.

      Thank you for flagging this mis-labeling, which Reviewer 3 also pointed out. We have switched the appropriate data point in Figure 3D from ‘trained’ to ‘naïve’.

      (11) In Figure 4D, the yellow line is not defined; based on Figure 6D, it likely represents shuffled/chance performance and should be labeled accordingly (including beneath the chance line on the plots).

      We have updated Figure 6 to indicate that the yellow line does indeed reflect shuffled/chance.

      (12) Figure 8 would benefit from a control demonstrating that poor cross-modal decoding reflects train-test distribution differences rather than weak decoders (e.g., train on a subsample of NH and test on held-out NH), and from reporting decoding on raw ERP/HG features in addition to TCA-derived data.

      Good suggestion, thanks; we have now added this control. We agree that a positive control is necessary to show that poor tone→CI decoding reflects differences of underlying representations rather than a failure of the decoder or modeling approach. (Reviewer 2 raised the same point.)

      To validate our cross‑modal analysis pipeline, we re‑implemented the full procedure used in Figure 8, but instead of training on tone‑evoked responses and testing on CI‑evoked responses, we trained and tested on independent sets of tone‑evoked trials from the same animals (tone→tone). For each tone in each animal, we withheld 10 trials as a test set. Using the remaining trials, we fit the original TCA model to obtain spatial and temporal factors (Fig. 8A). We then fixed these factors and re‑optimized only the trial factors on the withheld tone‑evoked trials (Fig. 8B). The LDA decoder was trained on the trial factors from the original TCA fit and tested on the re‑optimized trial factors from the withheld trials, using the same classification pipeline as in the main analysis.

      As shown in the top panels of Figure 8C,D, this positive control yielded robust tone→tone generalization: predicted tone frequencies closely matched the actual tones, decoder performance was significantly above chance, and prediction errors were tightly clustered around the true stimulus, indicating that the decoder was tuned to tone frequency. In contrast, when we trained on tone‑evoked responses and tested on CI‑evoked responses, information transfer was markedly reduced (Fig. 8E-G).

      These results demonstrate that the TCA+decoder pipeline can reliably transfer information across independent tone‑evoked datasets, confirming that the method captures shared structure when it exists. The poor cross‑modal transfer between tone‑ and CI‑evoked activity therefore is unlikely to be due to a weak decoder or to a failure of the modeling pipeline, but instead reflects a genuine mismatch between CI and sound representations in auditory cortex. We have updated Figure 8 and the Results section to describe this positive control analysis and clarify the interpretation.

      (13) Perception and interpretation of signals are mentioned several times in the introduction, although perception is not explored in the manuscript (only neuronal processing). This might be confusing.

      We appreciate the need to distinguish between neuronal encoding and perception. We also feel we have been careful not to invoke relationships to perception when presenting analyses on iEEG measurements, but we did identify an opportunity to further clarify this distinction between neuronal processing and perception by adding text in the intro, as follows “for the auditory system to interpret patterns of evoked neural activity and inform downstream auditory areas.”

      (14) Figure 1C. Why is the performance of CI rats so much lower than what was previously published (Glennon et al., 2023)? Did the training duration change?

      The three animals that were behaviorally trained on the normal-hearing (pre-deafening) and cochlear implant task (post-deafening) are within the distribution of the full set of animals from Glennon et al. (2023). However, we note that for Glennon et al. (2023), as one of our behavioral criterion was days to d’ > 1, animals were trained daily until reaching that level and not included in the initial data set if they did not reach that level. However, as we were including animals in this study of iEEG responses that were not trained at all, we felt it appropriate to include this third animal as well, that was trained just for 3 days before recordings were made. The two other animals were trained for 9 and 13 days. We have now included this information in the methods.

      (15) The p-values = 0.5 should be given with an additional digit.

      We previously rounded to the nearest single decimal digit, for all p-values greater than 0.10. We have updated the figures and manuscript text to ensure precision at least to the second digit.

      Reviewer #2 (Recommendations for the authors):

      We thank the Reviewer for their thoughtful comments on our study.

      (1) Less noisy recording methods based on spike detection would provide stronger claims.

      We agree that spike recordings, particularly isolated single-unit activity, are powerful for testing hypotheses about sensory encoding in auditory cortex, and we plan to incorporate such approaches in future work. However, our decision to use iEEG arrays in the present study was deliberate and central to the scientific and translational goals of the project.

      First, iEEG and related population-level approaches such as scalp EEG (e.g., Lalor and Foxe, 2010; O’Sullivan et al., 2015) and fNIRS (e.g., Bortfeld et al., 2009; Peelle, 2017) are widely used in humans and have been highly successful in decoding sound- and speech-evoked responses, revealing fundamental principles of how sound and speech are encoded in the human brain. Because speech is uniquely human and cochlear implants are primarily designed to restore speech perception, aligning our recordings with clinically relevant, human-used modalities enhances the translational relevance of our work.

      Second, iEEG arrays provide distinct advantages over modern multi- and single-unit electrophysiology. Even with high-density probes, the spatial sampling of neuronal activity does not match the coverage of the 60-channel iEEG arrays used here, which span large extents of auditory cortex. One might instead consider optical methods such as calcium imaging to interrogate topographical encoding at single-neuron and mesoscale resolutions, as has been done in normal-hearing mice (Romero and Hight et al., 2019). However, calcium signals are intrinsically slow, limiting access to the temporal precision that is critical for CI encoding, and these tools are unlikely to be available in humans in the foreseeable future, substantially reducing their translational value.

      Using iEEG arrays, we show that CI-evoked responses are topographically organized, consistent with prior work (Klinke et al. 1999, Bierer and Middlebrooks 2002, Middlebrooks and Bierer 2002, including Adenis et al., 2024 now referenced in the manuscript). Our study extends these findings by exploiting simultaneous recordings across both spatial and temporal domains, which are essential for several key analyses (Figs. 3-8), including quantification of trial-by-trial variability, decoding of stimulus identity from single trials, and cross-modal comparisons between normal-hearing and CI-evoked iEEG responses.

      Thus, we believe that the strength of this study is due to, rather than in spite of, its use of iEEG arrays. This approach uniquely allows us to test hypotheses about CI encoding across cortical topography and time using a modality that is directly translatable to human research and clinical practice. In response to the reviewer’s concern, we have also (i) improved the statistical treatment of our data (by adopting linear mixed-effects models that incorporate both paired and unpaired observations), (ii) added additional positive controls (see response to comment #2), and (iii) collected new data that further validate our rodent CI model. Together, these additions strengthen the support for our conclusions while preserving the key advantages of the iEEG-based approach.

      (2) A positive control is necessary to claim the mismatch between CI and sound representations.

      We agree. We now have added a positive control specifically designed to validate our cross-modal analysis pipeline in our revised manuscript. As also suggested by Reviewer 1, the goal was to test whether our method can successfully transfer information when the training and test datasets are matched in modality (tone→tone), thereby ensuring that the observed failure of cross-modal transfer (tone→CI) is not an artifact of the analysis.

      To do this, we re-implemented the full pipeline used in Figure 8, but instead of training on tone-evoked responses and testing on CI-evoked responses, we trained and tested on independent sets of tone-evoked trials from the same animals. For each tone in each animal, we withheld 10 trials as a test set. Using the remaining trials, we fit the original TCA model to obtain spatial and temporal factors (Fig. 8A). We then fixed these factors and re-optimized only the trial factors on the withheld tone-evoked trials (Fig. 8B). The LDA decoder was trained on the trial factors from the original TCA fit and tested on the re-optimized trial factors from the withheld trials, using the same classification pipeline as elsewhere in the manuscript.

      As shown in the top panels of Figure 8C,D, this positive control yielded robust tone→tone generalization: predicted tone frequencies closely matched the actual tones, decoder performance was significantly above chance, and prediction errors were tightly clustered around the true stimulus, indicating that the decoder was tuned to tone frequency. In contrast, when we trained on tone-evoked responses and tested on CI-evoked responses, information transfer was markedly reduced and not different from shuffled controls (Fig. 8E-G).

      These results demonstrate that the TCA+decoder pipeline can reliably transfer information across independent tone-evoked datasets, confirming that the method captures shared structure when it exists. The poor cross-modal transfer between tone- and CI-evoked activity therefore cannot be attributed to a failure of the modeling pipeline but instead reflects a mismatch between CI and sound representations in auditory cortex. We have updated Figure 8, the methods, and the results section to include this new important analysis.

      Reviewer #3 (Recommendations for the authors):

      We thank reviewer 3’s appreciation for study design and the appropriateness of analyses taken. We also appreciate the recognition of noteworthiness, specifically that stimulus identity can be decoded on a single-trial basis and of the potential benefit of using central decoders in clinical settings.

      (1a) Animal heterogeneity: It is difficult to keep track of the animals used in this study, and some received a different protocol of stimulation (sounds through the speech processor vs. direct stimulation) and were also trained in a behavioral task using different target stimuli (4kHz vs. 22.6kHz, also no mention of the CI electrode used as a target).

      We have now clarified the animal cohorts and stimulation protocols in our revised manuscript. We added a new Supplementary Figure 2 that schematizes, for each animal if it underwent behavioral training with pure tones in the normal-hearing condition, if tone-evoked iEEG measurements were collected, if CI-evoked iEEG measurements were collected (and whether stimulation was direct or via the speech processor), and if it subsequently received CI-based behavioral training. Regarding the behavioral targets, we now specify in the Methods that for normal-hearing training, the target stimulus was a 22.6-kHz pure tone. For CI-trained animals, the target was either CI channel 3 (n = 2 rats) or CI channel 4 (n = 1 rat). Details about stimuli targets during behavior have been added to the methods section under “Behavioral training for tone and implant channel detection.”

      (1b) There is no comparison of the CI maps from rats tested with the speech processor and directly stimulated. How different were they? Was the frequency allocation of each electrode the same for each animal? Since data might already have intrinsic variability because of the grid placement, the mechanical deafening, and the cochlear implantation in each animal, such heterogeneity in the 'background' and stimulation protocol might blur the authors' results.

      Our study focuses on cortical encoding of single-channel CI stimulation, so it is indeed important to ensure that the stimuli are effectively delivered by a single electrode, regardless of whether they are driven acoustically via the speech processor or by direct electrical stimulation.

      Stimulation mode and frequency allocation: The project began with single-channel stimulation achieved by presenting pure tones to the speech processor (N=3 animals) and later transitioned to direct programmatic control of individual electrodes (N=4 animals) to simplify the experimental setup. In both cases, the goal was to activate only one CI channel at a time.

      For the programming speech-processor animals, the validation protocol described in Glennon et al. (2023) is as follows:

      - Set the number of active channels in the processor to 1 (the clinical default is 8) to avoid spectral spread across electrodes.

      - Disabled all additional signal-processing strategies (e.g., Scan, ASC, ADRO, SNR-NR, WNR).

      - Used customized frequency allocation tables that mapped narrow frequency bands to individual electrodes, as shown in Glennon et al., 2023, Extended Data Fig. 2.

      To confirm that a given tone drove only the intended electrode, we recorded tone-evoked electrodograms—measurements of the output at each electrode—and verified that only the targeted channel was active (Glennon et al., 2023, Extended Data Fig. 2). Thus, although the initial CI drive was acoustic, the effective stimulation at the array was restricted to a single electrode with a well-defined frequency allocation.

      For the direct-stimulation animals, we used the same underlying frequency allocations to choose which electrode to stimulate, but the pulses were delivered programmatically rather than via the speech processor. In both modes, the center frequency associated with each electrode was therefore defined consistently across animals, and stimulation was confined to one channel at a time.

      Comparison of maps across stimulation modes: We now explicitly indicate the stimulation mode (speech-processor vs direct) for each CI animal in Supplementary Figure 2 and plot the maps for all animals in Supplementary Figure 3. Qualitatively, the spatial organization of CI-evoked maps is similar across the two stimulation strategies; we do not observe systematic differences in map structure that would suggest large biases introduced by the stimulation mode. However, the sample sizes for each group are small (N = 3 speech-processor, N = 4 direct). For this reason, we have not performed formal between-mode statistics and instead treat stimulation mode as a source of minor heterogeneity, alongside inevitable variability from grid placement, mechanical deafening, and cochlear insertion. Given the electrodogram validation (Glennon et al., 2023, Extended Data Fig. 2) and consistent frequency allocation tables, we are confident that both approaches produce single-channel activation with comparable effective frequency assignments.

      (1c) The number of animals used is also confusing. The authors report 7 NH and 7 CI animals (14 total), 4 NH and 3 CI were trained before being implanted (so 3 naïve NH and 4 naïve CI remain). Figure 1C reports that only 3 trained NH performed with the CI (let us call them 3 NH->CI). But then Figure 1E reports only 1 trained NH->CI and only 1 trained NH and 3 naïve NH that got implanted later. On the other hand, Figure 1E reports only 1 true naïve CI animal, the 3 others being naïve NH that got implanted. For the sake of clarity, I would encourage the authors to provide a timeline of the procedures/stimulation protocols coupled with a schematic distribution of the animals.

      To address this, we have added a new Supplementary Figure 2 that provides, for each individual animal a chronological timeline (NH recordings, deafening, implantation, CI recordings); if it was behaviorally trained in the NH condition, the CI condition, or both; if CI stimulation was delivered via the speech processor or by direct electrical stimulation; and which stimulus conditions (tone-evoked iEEG, CI-evoked iEEG) were collected. This schematic makes it clear how the reported totals arise (7 NH and 7 CI for iEEG; 4 NH-trained and 3 CI-trained behaviorally) and shows which specific animals contribute to each panel in Figure 1 and to the later iEEG analyses. We now reference Supplementary Figure 2 in the Results when introducing the cohorts to guide readers through animal accounting.

      (2a) Methods and statistics: Deafening is only mechanical, with no direct or postmortem proof that deafening was complete. The authors cite previous studies, but that would have been a good control to have since mechanical deafening isn't as accepted as the chemical deafening, like Neomycin, especially when some of your animals were stimulated with pure tones through the speech processor.”

      We agree that rigorous verification of deafening is essential, particularly when some CI animals are driven acoustically through the speech processor. Ototoxic approaches (e.g., systemic or local neomycin) are one established method, but their effectiveness can be sensitive to dose and delivery, and they introduce systemic side-effects that can complicate long-term survival and recovery.

      Our laboratory has used the mechanical deafening procedure since it was first described in King et al. (2016) and more recently in Glennon et al. (2023). In King et al., mechanical and ototoxic methods were combined, and we found that ototoxic methods provided no more additional robustness in deafening compared to mechanical lesion. Instead, the additional time required for ototoxic drug application reduced survival times in what was already a very complex and long surgical procedure for bilateral deafening and unilateral cochlear implantation.

      In Glennon et al. (2023) we intentionally employed mechanical-only deafening to minimize side-effects while still achieving profound hearing loss in implanted animals. Glennon et al. (2023) provides an extensive validation of this mechanical-only protocol under the same surgical and experimental conditions as the present study. As we mentioned in our response to comment #3a of Referee 1, we assessed deafness through three measures:

      Histology: In N=4 deafened animals, inner hair cell loss was ~50% and outer hair cell loss was near complete at almost 100% in all animals.

      Physiology: For N=14 rats, acoustic ABRs were substantial before deafening but statistically similar to baseline noise after deafening.

      Behavior: For N=16 deafened rats, behavioral performance with implant on was d′: 1.7±0.1, but when implant was turned off in a subset of sessions, performance dropped to chance (d′: −0.05±0.1, P < 0.0001).

      This convergent anatomical, physiological, and behavioral evidence demonstrates that the mechanical procedure produces profound deafness, with no functionally relevant residual hearing at levels ≥90 dB SPL. Also as we mentioned in response to comment #3a of Referee 1, we believe that the behavioral criterion is most essential and also least common in the literature. Because the tones used to drive the speech processor in the current study were presented at 70 dB SPL, we have no reason to believe that residual acoustic hearing contributed to any of the CI-evoked responses we report.

      We now cite these validation data explicitly in the methods under the section “Bilateral sensorineural hearing loss” as follows “(mechanical only, as described and validated in Glennon et al. 2023)” to make clear why we consider the mechanical-only approach sufficient for ensuring deafness in the present experiments.

      (2b) What motivated the selection of 15 Principal Components for the PCA? That might need to be justified, maybe by scree plot or variance plot (Eigen Values or CEV), as if too many PCs are selected, you are at risk of losing information. Side comment for TCA: why is it important that the number of latent factors exceeds the number of tones or stimuli? Is there a way to justify this statement?

      We thank the reviewer for raising this point. Our choice of 15 components/latent factors was motivated by both theoretical and empirical considerations, which are now made explicit in the manuscript.

      For the PCA analyses, we selected 15 principal components for two reasons. First, because our decoder must discriminate between 10 tone conditions, we reasoned that providing at least as many dimensions as stimuli would be beneficial, while also allowing for the possibility that some components may carry little or no stimulus-selective information. We therefore chose a modest number of components that exceeded the number of tones (10) but avoided unnecessarily high dimensionality. Second, we empirically examined the variance explained as a function of the number of components. As shown in the new scree plots (Supplemental Fig. 4A), the cumulative variance explained enters a near-linear, low-slope regime beyond ~15 PCs, indicating diminishing returns for including additional components. Thus, 15 PCs capture a substantial fraction of the stimulus-related variance while minimizing the risk of overfitting and retaining a consistent dimensionality across animals.

      For the TCA analyses, we used 15 latent factors to match the dimensionality used in PCA and to ensure that the latent space was sufficiently flexible to represent the 10 tone conditions without being under-parameterized. In practice, increasing the number of TCA components reduces reconstruction error (Williams et al., 2018), but with diminishing improvement beyond a certain point. We therefore systematically evaluated model error as a function of the number of latent factors and found that error decreased rapidly up to ~15 components and then plateaued (Supplemental Fig. 4B). This pattern parallels the PCA scree plots and supports 15 as a reasonable trade-off between model flexibility and parsimony.

      We have updated the Results clarify these choices, as follows “The number of components (15) was chosen based on PCA scree plots (Supplemental Fig. 4A), which showed that explained variance entered a near‑linear, low‑slope regime beyond this point demonstrating a similar plateau in reconstruction error (Supplemental Fig. 4B).”

      (2c) Legend of Figure 2E, J states that a Student's paired t-test was used, meaning that only the 'linked' points of the graph were used (thus, comparing only animals that got tested NH then implanted). This is usually the same across the manuscript. Why not include all the points with an unpaired t-test? Otherwise, why are all the points plotted if they serve no purpose? This choice should be justified.

      We agree with this concern, which was also raised by Reviewer 1. We have revised our statistical approach accordingly in our revised manuscript. In the original submission, we used paired t-tests when animals contributed both normal-hearing (NH) and CI data, which meant that animals with only NH or only CI measurements were excluded from those comparisons even though they were shown in the plots.

      To address this, we have re-analyzed all normal-hearing vs. CI comparisons using linear mixed-effects models that include both paired and unpaired data within a single framework. This approach ensures that every plotted data point contributes to the statistical tests, properly accounts for within-animal dependence when both conditions are present, and avoids the loss of power that would arise from either paired-only or purely unpaired tests.

      The mixed-effects results are consistent with our original interpretations, with two comparisons becoming significant in the updated analysis: Fig. 2E (p = 0.048) and Fig. 6F (p = 0.027). We have updated the Results and figure legends to describe the use of mixed-effects models and to report these revised p-values. Together with the new tonotopy and cochleotopy analyses described above, these changes strengthen the statistical support for our conclusions without altering the overall interpretation of the data.

      (2d) Side comment: There are inconsistencies on the bar plots of Figure 6C (Missing a purple point) and Figure 3D (Temporal has 3 purple points).

      Thank you for flagging this mis-labeling (which Reviewer 1 also noticed). We have correctly updated the appropriate data point from trained to naive for Fig. 3D and from naive to trained for Fig. 6C.

      (3a) Pure tones and CI-evoked responses maps: It is the reviewer's understanding that Figure 2 is an averaged representation for all animals. Why is the tonotopic shift so dim for ERPs? The averaged maps aren't very convincing. How were the gradients on an animal-to-animal basis since Figure 2D is only an example animal? Also, everything has been evaluated at 70dB, where selectivity might not be best. It would have been easier to follow the tonotopic gradient at the CFs where contrasts are higher.

      We agree that the strength and interpretation of tonotopy/cochleotopy in our iEEG data needed to be presented more clearly. Reviewer 1 raised closely related concerns, and we have substantially expanded the analyses and explanations in response. Here we highlight the points that address your specific questions.

      Single-animal vs. averaged maps: We included both exemplar maps and population summaries in Figure 2. The panels analogous to Figure 2D show single-animal best-frequency (BF) or best-channel maps; these were chosen because they exhibit clear, interpretable gradients. In the exemplar shown, there is a local high-frequency (HF) region along the medial edge of the array that transitions to lower frequencies toward the rostral edge. For CI-evoked best-channel maps in the same animal, we observe a parallel pattern in which basal electrodes (e.g., electrode 8, representing higher frequencies) occupy the HF region and apical electrodes (e.g., electrode 1, lower frequencies) occupy the LF region.

      Averaged ERP maps, by contrast, necessarily blur some of this structure because iEEG is a summed field potential and animal-to-animal differences in array placement, cochlear insertion depth, and anatomy introduce variability. We have softened the language in the text to reflect that ERP-based tonotopy is coarse and weaker at the population level, while emphasizing that robust gradients are evident in single animals and in HG-based measures.

      Quantitative assessment across animals: To move beyond visual impressions, we added quantitative analyses that mirror those used in Romero and Hight et al. (2020) for calcium imaging data (Romero and Hight et al. 2020 and Fig. 2). For each map we computed local tonotopic gradient vectors at every pixel and summarized their magnitude/direction on a unit circle, then compared the mean vector strength to shuffled maps. Applied to our BF and best-channel maps, this analysis shows that both are significantly more ordered than shuffled controls (p < 10<sup>-10</sup>), indicating that the maps are tonotopic/cochleotopic rather than random, despite the apparent dimness of the gradients in some averaged ERP plots. These new results are described in the revised manuscript and shown in Romero and Hight et al. 2020 and Fig. 2.

      Effect of intensity (70 dB SPL) and “dim” gradients: We agree that stimulus level influences the apparent sharpness of tonotopy. Higher intensities tend to broaden tuning and compress the dynamic range of BF maps. As we now discuss in more detail (adapted from our response to Reviewer 1), tones were presented at 70 dB SPL, so we expect maps to emphasize mid-frequency regions (around 8 kHz) and to show somewhat broader tuning than maps derived at threshold. For CI stimulation, we used ECAP thresholds to set intensity, which is effective in our preparation because animals can robustly discriminate individual electrodes and these electrodes evoke clear cortical activity (King et al., 2015; Glennon et al., 2023).

      In summary, we clarified which panels in Figure 2 show single-animal exemplars vs population summaries, added quantitative analyses demonstrating spatial correlations are greater for adjacent stimuli compared to far-apart stimuli, and expanded the discussion of how recording modality and stimulus level influence the visibility of tonotopic gradients. These changes are intended to make the evidence for tonotopy/cochleotopy in our iEEG data (and its limitations) more transparent.

      (3b) Since new experiments might not be available, it is the reviewer's suggestion to add a supplementary figure showing a couple of animal examples following the format of Figures 2A and 2C that have more contrasted gradients to strengthen the group data. In the case of the CI-evoked responses map, this might also provide another argument to dismiss the potential monopolar smearing.

      Good suggestion, thanks. We now include a new Supplementary Figure 3 that shows additional single-animal examples for both tone-evoked and CI-evoked maps, following the same format as Figure 2C.

      Regarding monopolar stimulation, we agree that monopolar configurations are expected to be less spatially specific than bipolar or multipolar modes because current returns to an extracochlear reference electrode, potentially broadening the spread of excitation. We nevertheless chose monopolar stimulation because it is the predominant clinical configuration in human CI users and therefore most relevant for translational purposes. We acquired ECAP measurements of peripheral (spatial and temporal) tuning via a forward masking paradigm and demonstrate that monopolar is effectively tuned (Supplemental Fig. 2). Together with additional single-animal maps in Supplementary Figure 3, together with our vector-strength analysis (Romero and Hight et al. 2020 and Fig. 2), demonstrate that even under acute monopolar stimulation we observe structured cochleotopic organization in cortex, rather than the fully smeared patterns one might expect if monopolar spread completely dominated.

      We also note that all CI-evoked iEEG measurements were made acutely, immediately after implantation and before any CI-based behavioral experience. It is possible that with longer-term use and plasticity, cortical cochleotopy could become sharper than what we observe here under acute conditions. In this sense, our data provide a conservative baseline showing that even at the earliest stages of CI use, monopolar stimulation already engages tonotopically selective regions of auditory cortex. A longitudinal comparison of acute versus chronic maps would be an interesting direction for future work but is beyond the scope of the current study.

      (3c) Side comments: The legends of Figures 2D and 2I should mention that this is an animal example and not group data, as the rest of the figures are group data.

      Thank you for this suggestion to improve figure clarity. We have updated all of our figures, where appropriate, to indicate whether data are single or groups of animals.

      (3d) In general, some of the legends should be revised because they are sometimes too "strong". As an example, Figure 3B, D legend states: "Variability of iEEG measurements across trials (root mean square, rms) was consistently higher for cochlear implant-evoked compared to tone-evoked activity", despite three of the statistical tests being non-significant. The manuscript is correct, on the other hand.

      Good point. We revised the legend for Figure 3 to be consistent with the figure and the manuscript.

      (3e) The example spatial map given in Figure 3A for CI might not be the best choice since it is showing a pretty reliable trial-by-trial response, while your group data proves the opposite.

      We understand the reviewer’s concern and agree that the exemplar CI map in Figure 3A appears relatively reliable on a trial-by-trial basis. This example was chosen deliberately from an animal in which we had both NH- and CI-evoked iEEG recordings, so that the reader could visually compare the two conditions within the same preparation. In this animal, as in the group data, the differences between NH and CI trial-by-trial responses are subtle rather than dramatic.

      Our group-level analysis shows that the RMS error across trials is consistently higher for CI-evoked than for NH-evoked responses, but the absolute differences are small (< 0.1) and relatively uniform across animals. The spatial maps plotted in Figure 3A are representative of this pattern: both conditions show reasonably robust evoked responses, with CI responses nonetheless showing slightly greater variability. To avoid implying a stronger qualitative difference than is supported by the data, we have revised the text to emphasize that (i) CI-evoked responses remain clearly detectable on single trials, and (ii) the key effect is a small but consistent increase in variability across animals, as captured by the RMS error metrics, “We noted that the differences were qualitatively subtle (Fig. 3A, right panel), they were consistent across animals (Fig. 3B).”

      (4a) Decoders for CI stimulation Regarding CI stimulation, Pearson's correlations were truncated at a spacing of 5 electrodes. Likewise, none of the LDA classifiers show prediction for channels past CI-6. Again, that choice should be justified, or the missing channels should be presented.

      We truncated the correlation between electrodes at 5 because beyond that, the estimated means are significantly noisy. These estimated means are noisy because the number of data are significantly reduced, also significantly increasing the standard error. For example, for the maximum stimulus spacing, the number of pairwise correlations is at maximum the number of animals tested (i.e., N=7). We believe it’s important to be transparent, so we have included the non-truncated version of the figure here in this public review (Author response image 3). We leave the figures in the manuscript untouched but have updated the Figure 2 legend justify this selection of data.

      Author response image 3.

      Expanded figures for spatial correlations and LDA performance. A) The same data from manuscript Figure 2 are re-plotted but with expanded x-axes to include up to 4.5 octaves and 7 channels. Due to the smaller numbers of data at these points, the estimates for the mean spatial correlations are noisier. In all cases, the mean correlations are significantly higher for the first data point compared to the last 3 (NH, ERP p<0.001; NH, HG p<0.001; CI, ERP p=0.005; and CI, HG p=0.39, linear mixed effects models). B) The same data from manuscript figure 4 are re-plotted but with expanded x-axes to include up to ±3.5 octaves and ±6 channels.

      (4b) Finally, retrained PCA-LDA on spatial-only and temporal-only for CI are absent in Figure 3D. Since the authors were pretty consistent in showing both NH and CI alongside in the rest of the paper, it would be coherent to add the CI counterpart to Figure 3D, or maybe with a supplementary figure.

      We agree that consistency can be improved by including classifiers for CI-evoked measurements, though presumably for Fig. 6C and not Fig. 3D. Figure 6 has been updated accordingly.

    1. Author response:

      The following is the authors’ response to the original reviews.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      This study investigates the impact of Pink1 loss on glial function and neuronal health in a Drosophila model, highlighting the role of mitochondria-organelle contacts and key genes such as Ccz1, Vps13, Mon1, and Rab7. The work provides insights into cellular processes underlying neurodegenerative diseases, with a focus on glia-neuron interactions. While the findings are promising, the study lacks critical controls, detailed mechanistic evidence, and explanatory figures to strengthen its claims.

      Strengths:

      (1) The study addresses an important topic in neuroscience, exploring the mechanisms of Pink1 loss, which has implications for Parkinson's disease and neurodegeneration.

      (2) The focus on mitochondria-organelle contacts and their regulation by Rab7-mediated pathways is novel and provides a potential mechanism for neuronal dysfunction.

      (3) The identification of key genes (Ccz1, Vps13, Mon1, Rab7) and their potential roles in Pink1-related pathways adds valuable knowledge to the field.

      (4) The manuscript uses a combination of genetic tools, Drosophila models, and functional assays to approach the problem from multiple angles.

      Weaknesses:

      (1) Specificity of Mz-Gal4: The study lacks validation of Mz-Gal4 specificity, as it may also drive expression in a few neurons or other types of glia. Additional control experiments using nls-GFP with Elav, Repo, or Draper antibody staining or alternative glial drivers would be helpful.

      We have addressed this issue of Gal4 driver specificity based on new experiments in the revised manuscript.

      (2) DLG staining is central to the story but is not well-supported by high-resolution Z-stack imaging, which should be included in the supplementary figures.

      We have included these in the supplement.

      (3) The manuscript does not confirm whether the candidate RNAi (Ccz1, Vps13, Mon1, Rab7) directly influence Rab7-mediated membrane trafficking or mitochondria-lysosome contacts in Pink1 mutants.

      This is indeed the case. These more mechanistic experiments were not yet performed.

      (4) Using ERG as a readout for EG effects in the antenna is not a direct or appropriate assay. Alternative functional assays relevant to antenna glia should be considered.

      We made the assumption that ensheating glial function is conserved across brain regions and now make this explicit in the reworded manuscript.

      (5) A graphical explanation of the interactions and functions of the candidate genes in Pink1 KO mutants is missing. This would greatly enhance the manuscript's clarity.

      We have included such a scheme in the new manuscript.

      (6) The study lacks details on sample sizes, effect sizes, and reproducibility, which are necessary for robust conclusions.

      We have included these essential data in the reworked document.

      (7) There are repeated words on page 3 ("olfactory Olfactory Receptor Neurons") and a lack of explanation in Figure 3C regarding the most up-regulated and down-regulated genes and the significance of large red dots.

      We have included the requested information.

      Reviewer #2 (Public review):

      Summary:

      This study proposes a novel role for ensheathing glia (EG) in a Pink1-model of Parkinson's disease and shows that this cell population exibits the highest number of DEG in a pre-symptomatic stage. In the olfactory system, there seems to be morphological changes in this cell-type that resembles an 'activated' state and the authors further show that the neuronal loss of Pink1 is responsible for this defect. The authors go on to show that manipulation of Pink1 in EG also leads to some defects in the visual system and in the dopaminergic neurons (DAN) that innervate the mushroom body (MB), and performed a screen based on the 'on-transient' defect of the ERG to identify potential genes that may modulate the function of EG in synaptic regulation. They focus on several genes related to Rab7/Vps13, and performed some additional experiments in the visual system and MB to propose the role of vesicle/lipid trafficking in EG as a important factor for PD pathogenesis.

      Strengths:

      The study proposes functional and mechanistic connections between several genes that have been linked to PD (PINK1, VPS13A/C). I feel that the data presented in Figure 1 and Fig3A-C are performed with rigor and are convincing/novel. The selection of Drosophila to study the questions is also a strength and the lab has extensive experiences in this field and model organism.

      Weaknesses:

      There is one fundamental concern I have with the genetic experiments performed in this paper (especially in Fig 3D and Fig4, see major issue #1), and I feel that there is a bit of a disconnect between the EG 'activation' phenotype the author show in the olfactory system and the other two neuronal systems (visual system, MB DAN) that the authors investigate see major issue #2). Also, there are quite a bit of information that is not provided in the manuscript (see major issues #3 and #4), which makes me difficult to judge the rigor and interpretation of several experiments.

      Major Concern #1: A number of lines used in this study are referred to as "RNAi" lines but when I look at the actual genotypes of reagents listed in the table in the METHODS section, many are actually NOT RNAi lines. Quite a few lines, including lines that the authors use as RNAi against Ccz1, Rab7 and Mon1, are gRNA lines for the TKO (TRiP-CRISPR knockout) system. While these reagents can theoretically knock-out these genes in somatic cells if used in combination with UAS-Cas9, there is no mention that UAS-Cas9 was used in this work throughout the manuscript. Hence, when these lines are just crossed to GAL4 with or without the Pink1 mutant, they shouldn't be having any effects. Similarly, the strongest hit from their screen was a TOE (TRiP-CRISPR Over Expression) gRNA against PIG-A, which could allow overexpression of PIG-A if there is a UAS-dCas9::VP64. However, I also do not see any mention that such activator was introduced into the crossing scheme. Considering that 3 of the 4 'hits' from their screen are not RNAi lines, I am quite skeptical of the study. Similarly, except for Vps13, all reagents used in Fig4 are TKO gRNA lines. Therefore, if this experiment was conducted without an UAS-Cas9, most of the data shown here are problematic. Also, note that several of the 'RNAi' lines listed in the Table in the METHODS section are actually MiMIC alleles. While some MiMIC lines could function as strong LOF alleles (if they are inserted in the exon or in an intron of the gene in the same orientation as the gene), some of the lines are not expected to affect gene function (e.g. FASN2 and CG17712, MiMICs are in introns and face the opposite orientation). Hence, the rationale of including these reagents in the screen doesn't make much sense. The description of the modifier screen should be much more detailed in the RESULTS and METHODS section and if the UAS-Cas9/dCas9::VP64 transgenes were not introduced when the TKO/TOE reagents were utilized, what can be concluded?

      In addition, for the 4 genes that the authors further study in Fig4, there are many other reagents that the authors can use, including mutant alleles, previously characterized RNAi lines (e.g. Vps13) and dominant negative/constitute active lines (e.g. especially for Rab7). The authors should validate their results with independent reagents to really convincingly show that the same conclusions can be drawn for the Vps13/Rab7 related genes since this is the key takeaway message of this paper.

      Also, they do not show whether the manipulation of these genes in a wild-type background (they only show what happens in Pink1 mutants) affect ERG and MB DAN synapse morphology. If these manipulations alone dramatically affect these phenotypes, it would be very difficult to interpret their data.

      We sincerely thank the reviewer for spotting this major oversight regarding the use of the TKO (TRiP-CRISPR knockout) and TOE (TRiP-CRISPR Over Expression) systems and the MiMIC alleles. As the reviewer pointed out, these lines were not used as intended, therefore our results and conclusions regarding the genetic interactions between Pink1 and several genes (PIG-A, Rab7, Ccz1, CG10646, Mon1, FASN2, CG17712), are incorrect and based on a technical mistake. These results were removed from the manuscript. While our mistake compromises the data regarding PIG-A, Rab7, Ccz1, CG10646, Mon1, FASN2, CG17712, it does not affect the results and conclusions for most of the genes of the screening and for Vps13 where we did use RNAi lines.

      Also, in the reworked manuscript, we provide additional evidence that modulation of vesicle trafficking proteins involved in mitochondria–endoplasmic reticulum (ER) membrane interactions, such as Vps13 and Vps35, influences neuronal function and rescues Pink1 mutant phenotypes when selectively downregulated in EG.

      Major Concern #2: In Figure 1, the authors show some morphological evidence that EG are 'activated' in Pink1 mutants, but whether the same phenomenon occurs in the visual system and in the MB is not shown. Since all of the studies in Fig3D and Fig4 are done in the visual system and MB, it is not clear whether the visual system and MB phenotypes are related to 'activation' of EG.

      Also, in the RNA-seq data in Fig1A and Fig3C, is there any molecular evidence that EG are indeed 'activated'? The only evidence that the authors show to state that EG are 'activated' in young Pink1 null animals is based on increased CD8::GFP staining in the olfactory system.

      The authors cannot draw a strong conclusion that indeed EG are 'activated' based on these data (e.g. perhaps the expression level of CD8::GFP is just increased). Additional evidence that the EG are 'activated' could be provided by looking at the increase in Draper intensity (as reported by Doherty et al. and MacDonald et al. that the authors cite), not only in the olfactory system, but also in the visual system and in the MB. It would also be informative if the authors can look at morphology of the EG in the visual system and MB to convincingly that the data shown in Fig4 is relevant to EG 'activation'.

      In line with the identification of DEG across the ensheating glia cluster in our single cell sequencing (where we did not distinguish between EG of different brain regions) we made the assumption that EG-(dys) function is consistent in the Pink1 mutant and conserved across brain regions. Nonetheless, to make clear that we did not consistently analyze EG morphology in the different brain regions that we probed in functional assays, we added a note in the manuscript. Furthermore, we also toned down our conclusion that the EG in Pink1 mutants are in an activated state: we note the similarity in phenotype in Pink1 mutants and situations of neuronal damage (where EG are activated) but added that the phenotype in Pink1 mutants may also be the result of the mere upregulation of GFP expression/fluorescence.

      Major Concern #3: In Fig3, there is no clear explanation why they focus on the ON transients and ignore the OFF transients, and also why the difference in the depolarization is not quantified in Fig4.

      We included this explanation in the reworked manuscript: In the Drosophila ERG, the sustained depolarization primarily reflects phototransduction in photoreceptors (and is defective when photoreceptors degenerate), whereas the ON and OFF transients arise from second-order lamina neurons and are widely used as readouts of signal transfer. We wanted to assess function and focused on the ON transient because in general it provides an onset-locked, more robust readout of function (Vilinsky & Johnson, 2012).

      Major Concern #4: While the authors claim that mz709-GAL4 is a EG specific driver, do the authors know that this is indeed true in the tissues and stages that are studied here? The Ito et al,. paper that is cited in the METHOD section has only looked at the expression of this reporter in embryonic and larval stages. The authors need to that the authors should validate their findings with an additional EG specific driver and/or provide additional data that mz709-GAL4 is indeed specific to EG in the adult fly brain and eye. If mz709-GAL4 is expressed in other cell-types, the interpretation of many of the data in this paper becomes quite questionable. I believe the data in Fig3B is suggesting that mz709-GAL4 is indeed specific to glia cells and not expressed in neurons, but whether this driver is truly specific to EG (and not in other glial types), especially in the visual system (including the lamina as well as in the eye), is not obvious.

      We labelled animals that express UAS-HisTag-eGFP (used also in our paper) under control of MZ709-Gal4 with anti-Elav (a neuronal marker) and find no significant overlap (see below “recommendation for authors”), consistent with MZ709-Gal4 not driving expression in neurons. This is consistent with previous published work: Indeed, MZ709-Gal4 has been amply used in adult flies and shown to be ensheating glia-specific (Doherty et al., 2009; Li et al., 2023; Sehgal et al.,2018). In the lamina neuropil of the Drosophila eye, MZ709-Gal4 is expressed in the marginal glia (Stenesen et al., 2019) which are neuropil-associated glia and are equivalent to generic ensheathing glia (Kremer et al., 2017). MZ709-Gal4 is also expressed also in satellite glia (Stenesen et al., 2019), but these glia enwrap the cell bodies of the lamina neurons and not the neuropil where synapses reside.

      Recommendations for the authors:

      Reviewing Editor Comments:

      We strongly encourage you to very carefully edit this manuscript. The reviewers made many probing comments that you should consider carefully.

      Reviewer #1 (Recommendations for the authors):

      (1) Validate the specificity of Mz-Gal4 by performing experiments with nls-GFP and Elav antibody staining to ensure there is no neuronal overlap. Additionally, consider using alternative glial-specific drivers, such as Repo-Gal4 or WG-Gal4, to confirm the findings.

      We expressed HisTag-eGFP (used also in our paper) under control of MZ709-Gal4 and labelled fly brains with anti-Elav (a neuronal marker). We do not observe significant overlap between the labels indicating MZ709-Gal4 does not express Gal4 in neurons (Supplementary figure 1).

      As indicated, these observations are consistent with previous published work. MZ709-Gal4 has been amply used in adult flies and shown to be ensheating glia-specific (Doherty et al., 2009; Li et al., 2023; Sehgal et al., 2018; Stahl et al., 2018). In the lamina neuropil of the Drosophila eye, MZ709-Gal4 is expressed in the marginal glia (Stenesen et al., 2019) which are neuropil-associated glia and are equivalent to generic ensheathing glia (Kremer et al., 2017). MZ709-Gal4 is also expressed also in satellite glia (Stenesen et al., 2019), but these glia enwrap the cell bodies of the lamina neurons and not the neuropil where synapses reside.

      (2) Include high-resolution Z-stack imaging of DLG staining to strengthen the assessment of synaptic integrity and ensure the robustness of the conclusions. These images should be added to either the main or supplementary figures.

      We included 2 supplementary figures (2 and 3) showing Z stacks that were used to delineate regions of interest at the MBs for the quantification of dopaminergic neuron afferents invasion. Our approach is identical to the one we used in Kaempf et al. 2026 (Kaempf et al., 2026).

      (3) Demonstrate whether the candidate RNAi (Ccz1, Vps13, Mon1, Rab7) directly influence Rab7-mediated membrane trafficking or mitochondria-lysosome contacts in Pink1 mutants. Use an appropriate method to confirm changes in organelle contacts in response to the RNAi treatments.

      Ccz1, Mon1 and Rab 7 were removed due to the technical mistake we made. We did confirm and maintain that Vps35 and Vps13 downregulation in EG rescues neuronal defects in Pink1 mutants. In the reworked manuscript we present a possible mechanism that involves the role of Vps35 and Vps13 in regulating ER-mitochondrial contacts, in line with our previous work (Valadas et al., 2018), while not ruling out possible other mechanisms.

      (4) Provide an alternative functional assay or evidence to support the use of ERG as a readout for EG effects in the antenna. Consider using a more direct assay relevant to antenna glia function.

      We agree that a more direct functional assay of antennal glia would be a nice addition (e.g., single-sensillum recordings or glial/ORN Ca<sup>2+</sup> imaging). However, implementing such assays would require new experimental pipelines and substantial additional data generation that is beyond our current ability and the scope of this revision.

      (5) Add a graphical illustration explaining the proposed mechanism of how Ccz1, Vps13, Mon1, and Rab7 function in Pink1 KO mutants, highlighting their interactions and roles within specific cell types.

      We included a schematic of our working model in Figure 5.

      (6) Clarify Figure 3C by explaining the most up-regulated and down-regulated genes and the significance of the large red dots. This will enhance the interpretability of the data.

      We expanded the legend to this figure: The large red dots represent the genes that rescue Pink1<sup>KO-WS</sup> phenotype when downregulated, the dark green dots are the 50 top most deregulated genes (magnitude of deregulation) in EG in Pink1<sup>KO-WS</sup> compared to controls, while the light green dots represent whole the genes detected in our cell-type specific transcriptomic experiment.

      (7) Correct repeated words on page 3 ("olfactory Olfactory Receptor Neurons") for clarity and consistency.

      Of course, sorry for this.

      (8) Ensure that sample sizes, effect sizes, and the number of replicates are explicitly stated for all experiments. This information is essential for evaluating the robustness and reproducibility of the findings.

      We made sure we consistently added all this information in the revised manuscript.

      (9) Verify and ensure that all data, reagents, and code used in the study are accessible and appropriately documented, in adherence with eLife's publishing policies.

      We made sure all data, reagents and code are available and/or properly described.

      By addressing these recommendations, the authors will significantly improve the clarity, rigor, and reproducibility of the manuscript.

      Reviewer #2 (Recommendations for the authors):

      Minor Points.

      (1) All figures seem to lack titles.

      We fixed this error.

      (2) In the abstract, the authors say that Rab7 and Vps13 are mutated in PD patients but I couldn't find the reference/information for Rab7 (the authors do refer to papers that linked VPS13A/C variants to PD but no mention about RAB7A/B being linked to PD). Please discuss this in the paper or modify the abstract accordingly.

      We removed this statement for rab7 from the paper.

      (3) When referring to the human gene, Pink1 should be written as PINK1 according to the HGNC nomenclature rules.

      We made this change.

      (4) The authors say Vps13 has two mammalian orthologs but actually it has four (VPS13A/B/C/D). I guess two of the four is linked to PD so the authors should modify there statement to reflect this.

      This is a misinterpretation of what we meant and we have clarified our intention: Drosophila possesses 3 paralogues of Vps13 - Vps13, Vps13B, and Vps13D - which we also detected in our screening (Neuman et al., 2025; Velayos-Baeza et al., 2004; Vonk et al., 2017). Among these Vps13 is most similar to human VPS13A and VPS13C (Hanna et al., 2023; McEwan & Ryan, 2022).

      (5) The abbreviation 'CNS' is used in the first page of the intro but I don't see it being spelled out as "central nervous system".

      We have spelled out central nervous system in the first page of the introduction.

      (6) On the top of page 5, the authors state that they confirmed that the 'synaptic area of DAN show a decrease in aged (25 days) animals' but data is not shown. If they want to make a statement like this, I believe such data should be included in supplemental data. Since the phenotype in the aged animal is not relevant to this study, one could remove this statement regarding the aged animals if they prefer not to show the data.

      The decreased synaptic area of DAN in 25-day old Pink1 mutants is shown in figure 2C-D of the manuscript and is consistent with data shown in (Kaempf et al., 2026).

    1. Author response:

      The following is the authors’ response to the original reviews.

      (1) We bioinformatically examined the repeat compositions of MLSs (Figure 3B), which clearly indicated that all MLSs are composed of repetitive sequences to a much greater extent than the rest of the genome.

      (2) We confirmed the blockage of chromosome breakage by the 4R-CBS mutations using a telomere-anchored PCR assay (Figure 5C-E).

      (3) We examined the effect of the 4R-CBS mutations on the expression of genes encoded in 4R-MDS by RNA-seq (Figure 9). This analysis unexpectedly revealed that gene expression from 4R-MDS is not significantly affected in the mutants, allowing us to extend our discussion.

      (4) We added two authors, Alix Lemoine and Tomoko Noto, who performed the experiments for these revisions.

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, Nagao and Mochizuki examine the fate of germline chromosome ends during somatic genome differentiation in the ciliate Tetrahymena thermophila. During sexual reproduction, a new somatic genome is created from a zygotic, germline-derived genome by extensive programmed DNA elimination events. It has been known for some time that the termini of the germline chromosomes are eliminated, but the exact process and kinetics of the elimination events have not been thoroughly investigated. The authors first use germline-specific telomere probes to show that the loss of these chromosome ends occurs with similar timing as other DNA elimination events. By comparative analysis of the assembled germline and somatic genomes, the authors find that the ends of each of the germline chromosomes are composed of a few hundred kilobases of micronuclear limited sequences (MLS) that are removed starting around 14 hours after the start of conjugation, which initiates sexual development. They then develop an in situ hybridization assay to track the fate of one end of chromosome 4 while simultaneously following the adjacent macronuclear destined sequence (MDS) retained in the new somatic genome. This allows the authors to more clearly show that these adjacent chromosomal segments are initially amplified in the developing genome before the terminal MLS is eliminated. Finally, they mutate the chromosome breakage sequence (CBS) that normally separates the MLS terminus from the adjacent MDS region, to show that strains that develop with only one mutant chromosome can produce viable sexual progeny, but it appears that both the MLS and the MDS from the mutant chromosome are lost. If both chromosome copies have the CBS mutation, the cells arrest during development and do not eliminate many germline-limited sequences and fail to produce viable progeny. Overall, this study provides many new insights into the fate of germline chromosome ends during somatic genome remodeling and suggests extensive coordination of different DNA elimination events in Tetrahymena.

      Strengths:

      Overall, the experiments were well executed with appropriate controls. The findings are generally robust. Importantly, the study provides several novel findings. First, the authors provide a fairly comprehensive characterization of the size of the MLS at the end of each germline chromosome. I'm not sure whether this has been published elsewhere. Second, the authors develop a novel method to study the fate of chromosome termini during development and use it to conclusively track the elimination of these termini. Third, the authors show that the elimination of these termini appears to occur concurrently with most other DNA elimination events during somatic genome differentiation. And fourth, the authors show that failure to separate these eliminated sequences from the normally retained chromosome alters the fate of these adjacent MDS and the loss of the cells' ability to produce viable progeny.

      Weaknesses:

      It appears the authors did extensive analysis of the MLS chromosome ends, but did not provide too much information related to their composition. If this has not been published elsewhere, it would be useful to describe the proportion of unique and repetitive sequences and provide more information about the general composition of the chromosome ends. Such information would help the reader understand the nature of these MLS and how they may or may not differ from other eliminated sequences.

      We now calculated the proportions of unique and repetitive sequences for each MLS, and these data are included in Figure 3B and described in the main text of the revised manuscript. A more comprehensive analysis of chromosome-end composition, including detailed characterization in the context of the complete MIC genome assembly, is beyond the scope of the current study and will be presented in a future publication.

      Although the development of the novel FISH probes for large chromosome ends allowed for these novel discoveries, the signal in several images was visible, but often quite faint. I'm not sure there is anything the authors could do to improve the signal-to-noise ratio, but one needs to stare at the images carefully to understand the findings.

      We have submitted higher-resolution images for the revised manuscript, which we believe much improve the visibility of faint signals.

      One main weakness in the opinion of this reviewer is that the authors did very little to understand why, when a terminal MLS and the adjacent MDS fail to get separated because of failure in chromosome breakage, both segments are eliminated. The authors propose that possibly essential genes in the MDS get silenced, and the resulting lack of gene expression is the issue, but this and other possibilities were not tested. The study would provide more mechanistic insight if they had tried to assess whether the MDS on the CBS mutant chromosome becomes enriched in silencing modifications (e.g., H3K9me3). Alternatively, the authors could have examined changes in gene expression for some of the loci on the neighbouring MDS.

      The 4R-CBS mutation causes two distinct defects that should be considered separately: (1) co-elimination of 4R-MLS and the adjacent 4R-MDS during uniparental transmission of the 4R-CBS mutation; and (2) a global block of DNA elimination during biparental transmission of the 4R-CBS mutation.

      For the first defect, 4R-MLS and 4R-MDS may simply co-segregate into the nuclear compartment where DNA elimination occurs when the chromosome break that normally separates 4R-MLS from 4R-MDS is blocked. In this scenario, no additional process, such as spreading of scnRNA production, heterochromatin formation, or gene silencing, would be required to induce co-elimination. This point was not clearly stated in the previous manuscript, and we have now added a discussion of it to the revised manuscript.

      The possibility of gene silencing within 4R-MDS was raised as a potential explanation for the second defect. To test this possibility, we performed RNA-seq analysis of wild-type and 4R-CBS mutant cells to determine whether gene expression from 4R-MDS is affected by mutations at 4R-CBS. Contrary to our expectations, we found that genes in 4R-MDS are not significantly down-regulated in 4R-CBS mutant cells compared with other genes. This result suggests that the DNA elimination defect in these cells cannot be explained by silencing of genes located within 4R-MDS. We have added these RNA-seq data to Figure 9 and described them in the Results section. We have also revised the Discussion to propose alternative possibilities that may guide future investigations.

      The other main weakness is that since the authors only mutated the end of one germline chromosome, it is not clear whether the elimination of the MDS adjacent to the terminal MLS on chromosome 4 when the CBS is mutated is a general phenomenon, i.e., would happen at all chromosome ends, or is unique to the situation at Chromosome 4R. Knowing whether it is a general phenomenon or not would provide important insight into the authors' findings.

      As was described in the manuscript, the short (CBS = 15 nt) target within AT-rich and repetitive regions prevent designing gRNAs specifically targeting some of the chromosome end CBSs. We tried to mutate the CBS sequences of the left end of the chromosome 3 (3L) and the left end of the chromosome 5 (5L) by the strategy we used to mutate 4R-CBS but failed. Therefore, to systematically mutate other chromosome-end CBSs, we need to establish a different strategy, such as combining template-based repairing to CRISPR-induced DSB. We have explained this technical limitation and stated that “Our data support a critical role for 4R-CBS in separating 4R-MLS from 4R-MDS, but it remains unclear whether all MIC chromosome ends are strictly CBS-dependent for their elimination.” in Discussion (Page 12).

      Reviewer #2 (Public review):

      Summary:

      Nagao and Mochizuki investigated how the germline (MIC) telomere was removed during programmed genome rearrangement in the developing somatic nucleus (MAC). Using an optimized oligo-FISH procedure, the authors demonstrated that MIC telomeres were co-eliminated with a large region of MIC-limited sequences (MLS) demarcated on the opposite side by a sub-telomeric chromosome breakage site (CBS). This conclusion was corroborated by the latest assembly of the Tetrahymena MIC genome. They further employed CRISPR-Cas9 mutagenesis to disrupt a specific sub-telomeric CBS (4R-CBS). In uniparental progeny (mutant X WT), DNA elimination of the sub-telomeric MLS was not affected, but the adjacent MAC-destined sequence (MDS) may be co-eliminated. However, in biparental progeny (mutant X mutant), global DNA elimination was arrested, revealing previously unrecognized connections between chromosome breakage and DNA elimination. It also paves the way for future studies into the underlying molecular mechanisms. The work is rigorous, well-controlled, and offers important insights into how eukaryotic genomes demarcate genic regions (retained DNA) and regions derived from transposable elements (TE; eliminated DNA) during differentiation. The identification of chromosome breakage sequences as barriers preventing the spread of silencing (and ultimately, DNA elimination) from TE-derived regions into functional somatic genes is a key conceptual contribution.

      Strengths:

      New method development: Oligo-FISH in Tetrahymena. This allows high-resolution visualization of critical genome rearrangement events during MIC-to-MAC differentiation. This method will be a very powerful tool in this area of study.

      Integration of cytological and genomic data. The conclusion is strongly supported by both analyses.

      Rigorous genetic analysis of the role played by 4R-CBS in separating the fate of sub-telomeric MLS (elimination) and MDS (retention). DNA elimination in ciliates has long been regarded as an extreme form of gene silencing. Now, chromosome breakage sequences can be viewed as an extreme form of gene insulators.

      Weaknesses:

      The finding of global disruption of DNA elimination in 4R-CBS mutant progeny is highly intriguing, but it's mostly presented as a hypothesis in the Discussion. The authors propose that the failure to separate MLS from MDS allows aberrant heterochromatin spreading from the former into the latter, potentially silencing genes required for DNA elimination itself. While supported by prior literature on heterochromatin feedback loops, the specific targets silenced are not identified. While results from ChIP-seq and small RNA-seq can greatly strengthen the paper, the reviewer understands that direct molecular characterization may be beyond the scope of the current work.

      As mentioned in our reply to Reviewer #1’s comment above, we performed RNA-seq on wild-type and 4R-CBS mutant cells at 13.5 hpm and 15 hpm and found that genes in 4R-MDS are not significantly downregulated in 4R-CBS mutant cells (Figure 9), suggesting that the DNA elimination defect in these cells cannot be explained by aberrant heterochromatin spreading. Therefore, the link between the chromosome break at 4R-CBS and general DNA elimination remains elusive and will be a very interesting subject for our future research. We have added these results and revised the discussion in the manuscript.

      Reviewer #3 (Public review):

      Programmed DNA elimination (PDE) is a process that removes a substantial amount of genomic DNA during development. While it contradicts the genome constancy rule, an increasing number of organisms have been found to undergo PDE, indicating its potential biological function. Single-cell ciliates have been used as a prominent model system for studying PDE, providing important mechanistic insights into this process. Many of those studies have focused on the excision of internally eliminated sequences (IES) and the subsequent repair using non-homologous end joining (NHEJ). These studies have led to the identification of small RNAs that mark retained or eliminated regions and the transposons that generate double-strand breaks.

      In this manuscript, Nagao and Mochizuki examined the other type of breaks in ciliates that were healed with telomere addition. They specifically focused on the sequences at the ends of the germline (MIC) chromosomes, which have received relatively less attention due to the technical challenges associated with the highly repetitive nature of the sequences. The authors used the Tetrahymena model and developed a set of new tools. They used a novel FISH strategy that enables the distinction between germline and somatic telomeres, as well as the retained and eliminated DNA near the chromosome ends. This allows them to track these sequences at the cellular level throughout the development process, where PDE occurs. They also analyzed the more comprehensive germline and somatic genomes and determined at the sequence level the loss of subtelomeric and telomere sequences at all chromosome ends. Their result is reminiscent of the PDE observed in nematodes, where all germline chromosome ends are removed and remodeled. Thus, the finding connects two independent PDE systems, a protozoan and a metazoan, and suggests the convergent evolution of chromosome end removal and remodeling in PDE.

      The majority of sites (8/10) at the junctions of retained and eliminated DNA at the chromosome ends contain a chromosome breakage sequence (CBS). The authors created a set of mutants that modify the CBS at the ends of chromosome 4R. CBS regions are challenging for CRISPR due to their AT-rich sequences, making the creation of the 4R-CBS mutants a significant breakthrough. They used the FISH assay to determine if PDE still occurs in these mutant strains with compromised CBS. Surprisingly, they found that instead of blocking PDE, its adjacent retained DNA is now eliminated, suggesting a co-elimination event when the breakage is impaired. Furthermore, in biparental mutant crosses, no PDE occurred, and no viable progeny were produced, indicating that the removal of chromosome ends is crucial for proper PDE and sexual progeny development. Overall, the work demonstrates a critical role for 4R-CBS in separating retained and eliminated DNA.

      We appreciate Reviewer 3’s assessment.

      Recommendations for the authors:

      Reviewing Editor Comments:

      All reviewers agree that this study makes an important contribution to the field; however, they also offered several suggestions for how the manuscript could be improved. In particular, we draw your attention to the comments from Reviewer #1, who suggests that the manuscript could benefit from additional information on the general composition of germline chromosome ends, where available.

      As noted in our response to Reviewer #1 in the Public Reviews above, we have included an analysis of the fraction of repetitive sequences for each MLS as Figure 3B in the revised manuscript, highlighting the highly repetitive nature of MLSs compared with the rest of the genome.

      Reviewer #1 (Recommendations for the authors):

      As mentioned in the weaknesses section, the authors could provide more information regarding the nature of the sequences that make up the terminal MLS. There have been reports that these are highly repetitive; is that the case? Also, did the authors identify common repeats that are not internal to mic chromosomes that could be used to track all terminal segments of the five chromosomes? This would complement their mic-telomere probe.

      As noted in our response to Reviewer #1’s Public Review above, we have added an analysis of the fraction of repetitive sequences for each MLS as Figure 3B in the revised manuscript, which confirms that MLSs are highly repetitive.

      Apart from the moderately conserved Telomere Associated Sequence (TAS), described by Kirk and Blackburn (1995) and of unknown function, we were unable to identify any obvious shared repeats unique to MLSs that could support the development of pan-MLS-specific probes.

      One major weakness is that the authors did little to determine the cause of the elimination of the adjacent MDS along the 4R-MLS when the CBS was mutated. It would really improve the study if the authors could show that:

      (1) Gene expression of genes on the MDS is reduced in 4r-CBS mutant progeny.

      (2) Heterochromatin modifications are unexpectedly acquired on the MDS in mutants relative to wild-type chromosomes.

      (3) Do scnRNA specific to the MDS region appear in the mutant progeny during development, but not in wild-type crosses?

      Any data that would help support the authors' hypothesis regarding how the MDS region is eliminated when the CBS is mutant would definitely strengthen the conclusions of the study.

      As noted in our response to Reviewer #1’s Public Review above, we performed RNA-seq on wild-type and 4R-CBS mutant cells at 13.5 hpm and 15 hpm. Our analysis showed that genes within the 4R-MDS are not significantly downregulated in 4R-CBS mutant cells (Figure 9), suggesting that the DNA elimination defect in these cells cannot be attributed to aberrant heterochromatin spreading. Therefore, the connection between the chromosome break at 4R-CBS and general DNA elimination remains unclear and represents an important avenue for future investigation. We have incorporated these results and revised the discussion accordingly in the updated manuscript.

      The other main weakness is that by mutating the CBS of only one chromosome arm, one can't know whether the loss of the MDS with the MLS in the mutants is generalizable for all chromosome arms or is unique to 4R. The authors noted that they were unable to make any other mutated CBSs. Another way to try to get to this question is to try to rescue the mutant by inserting a new CBS into the 4R arm such that some MLS remains linked to the 4R-MDS and see whether removing the mic telomere is the issue, or would a block of MLS attached to the 4R-MDS be sufficient to cause its elimination. I'm not sure where to exactly put the new CBS, but worth thinking about.

      To introduce a new CBS into 4R-MLS, we would need to insert a CBS-containing construct into the MIC by homologous recombination during conjugation and then select engineered transformants using a drug resistance marker expressed from the derived MAC. However, because 4R-MLS is still eliminated in the progeny of 4R-CBS mutants, the introduced marker would be lost from the MAC even if homologous recombination were successful. Therefore, although the strategy suggested by this reviewer is very interesting, several technical innovations are required to make such experiments feasible, leaving this approach for a future project.

      It seems somewhat curious that the mutation of the CBS completely blocks nuclear development. In Paramecium, the failure to complete internal DNA elimination events can lead to alternative telomere addition. The caveat being that, in Paramecium, telomere addition appears more promiscuous than in Tetrahymena. It would be helpful to know how absolute the failure to produce progeny is in these mutants. Is it zero progeny in 10<sup>6</sup>, 10<sup>7</sup>, 10<sup>8</sup> ..... mated cells? Can the authors provide a possible lowest possible frequency?

      The viability tests were performed using bulk mating of 2.5 × 10<sup>4</sup> cells for each cross. Because ~70-80% of mating pairs complete the conjugation process and produce exconjugants under our standard culture conditions, and because we did not detect any 6-mp-resistant progeny from MUT x MUT crosses, we estimate that the probability of obtaining viable progeny in these crosses was less than 1 progeny per ~2 × 10<sup>4</sup> mating pairs. The number of cells used for the viability assay is described in the “Viability Test of Sexual Progeny” section of Materials and Methods and the estimated frequency of progeny production from the mutants has been mentioned in Results section in the revised manuscript.

      The one implication of the study is that chromosome breakage and DNA elimination, two different events, are coupled. In most mutants that block scnRNA-directed DNA elimination, both IES excision and chromosome breakage occur. In the study by McDaniel, SL. et al (2016). DRH1, a p68-related RNA helicase, is required for chromosome breakage in Tetrahymena. Biology Open pii: bio.021576. doi: 10.1242/bio.021576, germline knockouts of DRH1 could complete IES excision, but not chromosome breakage, indicating that the processes can be uncoupled. It may be useful for the authors to discuss this previous work in relation to their finding that failure in chromosome breakage can lead to DNA elimination of neighboring sequences.

      So far, DRH1 is the only gene reported to be required for chromosome breakage without affecting DNA elimination in Tetrahymena. However, McDaniel SL et al. (2016) examined chromosome breakage at only two CBSs (distinct from 4R-CBS), and thus it remains unclear how broadly chromosome breakage, including that at 4R-CBS, is affected in the absence of DRH1. In addition, McDaniel SL et al. (2016) assessed DNA elimination at three different IESs using PCR, whereas our study examined elimination of the repetitive Tlr1 transposon using FISH. Therefore, without further analysis of the similarities and differences in chromosome breakage and DNA elimination phenotypes between DRH1 knockout cells and 4R-CBS mutants, it is difficult to draw meaningful conclusions. Accordingly, we have limited ourselves to stating the following in the Discussion of the revised manuscript: “Moreover, chromosome breakage can be inhibited without disrupting DNA elimination, as shown in cells lacking zygotic expression of the p68-like RNA helicase Drh1 (McDaniel et al., 2016).”

      Minor corrections:

      Page 7, line 3: the text "......inducing chromosome break" should either be "......inducing chromosome breaks" or "......inducing a chromosome break".

      Corrected as “inducing a chromosome break”.

      Page 13, line 13: "......large block...." should be "......large blocks......".

      Corrected as suggested.

      Reviewer #2 (Recommendations for the authors):

      The authors can experimentally validate that chromosome breakage at 4R-CBS is indeed disrupted by the mutations. A PCR-based assay testing de novo telomere addition is a standard tool. In addition, MLS-linked telomere should only appear transiently during conjugation in WT cells.

      Because it was previously unknown whether de novo telomere addition occurs at the ends of MLSs upon chromosome breakage, we tested this using a PCR-based assay. We detected telomere-added chromosome ends of 4R-MLS and 3L-MLS, which were undetectable until 10.5 hpm, appeared at 12 hpm, and gradually decreased by 18 hpm in wild-type cells (WT × WT cross). Importantly, the appearance of the telomere-added 4R-MLS end, but not the 3L-MLS end, was blocked in 4R-CBS mutants (Mut x Mut crosses), strongly supporting that the 4R-CBS mutations specifically disrupt chromosome breakage at 4R-CBS. These new data are shown in Figure 5C–E and described in the Results section.

      The high FISH background during conjugation may be caused by the abundant presence of dsRNA, which is resistant to RNase A treatment but may be degraded by RNase III.

      The high FISH background was observed in the parental MAC at 9 and 12 hpm (Figure 2, 4, and S2) where dsRNA accumulation was not detected in the previous studies (Woo et al. 2016; Shehzada et al. 2024). In contrast, the MIC at 3 hpm and the new MAC at 9 and 12 hpm, where strong dsRNA accumulation was detected, showed much weaker background FISH signals (Figure 2, 4, and S2). Therefore, we believe that dsRNA is not the main cause of the high FISH background.

      It is likely that the long MIC telomere is treated as IES and targeted for DNA elimination. Indeed, telomere-specific scnRNA is abundantly produced during conjugation (http://www.ncbi.nlm.nih.gov/pubmed/19460867).

      We have cited the suggested literature and the following description has been added in Discussion to relate the reported telomere-derived scnRNAs to the abundant scnRNAs produced from MIC chromosomal ends: “In addition, telomere-complementary scnRNAs were reported to be produced specifically during conjugation (Cao et al. 2009).”

      Global disruption of DNA elimination may be a direct effect (DNA excision machinery affected) or indirect (unrepaired DSB and checkpoint activation).

      It has been reported that unrepaired DSBs caused by loss of Ku80 (Tku80) do not block DNA elimination in Tetrahymena (Lin et al. 2012). Therefore, checkpoint activation by unrepaired DSBs, if it occurs, is unlikely to explain the DNA elimination defect observed in the progeny of 4R-CBS mutants. Nonetheless, this direct-versus-indirect issue would be relevant when considering whether disruption of specific 4R-MDS-encoded genes in 4R-CBS mutants could cause the DNA elimination defect. Our new RNA-seq analysis, however, suggests that this possibility is unlikely. Therefore, we did not add further discussion of this direct-versus-indirect issue.

      Minor points:

      The zoom-in boxes in most images are barely visible.

      We have modified the zoom-in boxes to make them clearer.

      Page 13: scnRNA precursors (Cai et al., 2025) (Cai et al., in press). Is it one paper or two?

      They are two papers and the latter was published reacently. We have updated the citation.

      Reviewer #3 (Recommendations for the authors):

      The manuscript is well-written, with clear data, thoughtful discussion, and concise presentation. I have only a few minor comments below.

      For Figure 4 and others, the right panel shows the stats and percentages, with positive and negative labels. It's a bit confusing at first glance. I think it can be clarified what positive and negative mean in the legend.

      The legends of Figure 4, Figure 6 and Supplementary Figure S2, have been modified as “The presence (Positive) or absence (Negative) of the 4R-MLS FISH signal in new MAC (An) in 50 cells per time point was examined.”

      The quality of the FISH images is low at their current resolution. It is difficult to get a clear view.

      In the initial version, some images were in low resolution when we combined them into a single pdf file for review. In the revised manuscript, the images have been replaced with high-resolution images.

      The co-elimination of neighboring 4R-MDS when 4R-CBS is mutated, can this be viewed as a fail-safe mechanism to ensure the elimination of the chromosome ends? Regardless, the result begs the question of the significance of end removal and remodeling of PDE. Some speculations in the discussion might be helpful.

      Because the neighboring 4R-MDS contains approximately 100 predicted genes, its co-elimination would likely be too risky to evolve as a fail-safe mechanism for ensuring chromosome-end elimination in every generation. Instead, we interpret this as an erroneous process that can still be compensated for through endoreplication of the remaining, normally processed 4R-MDS from the non-mutated copy.

      We further speculate that the connection between chromosome breakage at 4R-CBS and the essential PDE process may serve as an evolutionary pressure to preserve the 4R-CBS locus in a chromosome breakage-competent state. We have added the following discussion to the revised manuscript (Page 15): “The observed link between chromosome breakage at 4R-CBS and the essential DNA elimination process may reflect the biological significance of MLSs and the importance of their removal from the MAC. Coupling these processes may have evolved as a mechanism to ensure that only functional chromosome-end CBS loci are preferentially transmitted to future generations.”

      Figure 1, legend, line 3, "the sexual reproduction process", do you mean "the sexual reproduction proceeds or initiates"?

      We meant “conjugation” = “the sexual reproduction process”. To make this clearer, we have revised the legend as “conjugation, which is the sexual reproduction process of Tetrahymena”.

    1. Author response:

      The following is the authors’ response to the previous reviews

      Public Reviews:

      Reviewer #1 (Public review):

      Summary:

      In this study, the authors propose that HSV-1 infection degrades the class I histone deacetylases HDAC1 and HDAC2. The MDM2 E3 ubiquitin ligase from the DNA damage response pathway is responsible for ubiquitinating these HDACs that are subsequently degraded via proteasomes. The authors hypothesize that HDAC degradation will cause hyperacetylation of viral chromatin and enable viral gene transcription.

      Strengths:

      The ubiquitination of HDAC1 & HDAC2 by Mdm2 and the mapping studies are clear.

      Comments on revised version:

      The authors enhanced their manuscript by more supportive data and providing clarification and the necessary corrections. However, a few more issues pertain:

      (1) In Figure 4j at 2 h post-infection we typically see the input virus and not progeny virus production. The input seems to have about 1-log difference that is expected to impact the results.

      We sincerely appreciate the reviewer's valuable comments regarding the timing notation. It should be noted that the "2 h" indicated in Figure 4j does not refer to two hours after the start of viral infection, but rather to two hours following medium replacement—after the virus has completed adsorption and internalization at 37°C (typically taking 2 hours), with this moment defined as the new time zero point (t = 0 h). Thus, this corresponds to approximately 4 hours post-infection (4 hpi). All subsequent sampling time points (4, 6, 12, and 24 h) are consistently defined according to this same system. This temporal framework aligns with previous studies: Nobe et al. (mBio 2025; DOI: 10.1128/mbio.00280-25) have clearly demonstrated that newly generated viral particles can be detected as early as 4 hours after HSV-1 infection, supporting the possibility of early progeny virus production at this time point in our experiment. We have accordingly revised the figure legend for Figure 4j to explicitly state the time reference ("t = 0 h defined as time of medium replacement post-adsorption") and added detailed procedural descriptions in the Methods section regarding adsorption, medium change, and sample collection time points to ensure clarity and reproducibility of the timing protocol.

      (2) Figs 1A, 1E, 2H it seems unclear why ICP4 becomes detectable at 12 h post-infection in HeLa cells? How about other a-genes? How about other cells? ICP4 is typically detectable within 2-3 h post-infection.

      We sincerely appreciate the valuable comments provided by the reviewers. Regarding the observation that ICP4 was detected only after 12 hours post-infection in HeLa cells, we re-evaluated our experimental conditions and reviewed relevant literature. The results indicate that at a higher multiplicity of infection (MOI = 5), ICP4 can indeed be reliably detected in HeLa cells as early as 2 hours post-infection (Author response image 1). Notably, Fouad S. El-Mayet et al. reported that under MOI = 1, ICP4 could not be detected until 8 hours after HSV-1 infection of mouse neuroblastoma Neuro-2A cells Figure 5A (Fouad S. El-Mayet et al., Antiviral Research, 2024, DOI: 10.1016/j.antiviral.2024.105870), although their early protein VP16 showed positive expression as early as 4 hours post-infection. This time difference is closely related to cell type: Neuro-2A is a highly susceptible neuronal cell line for HSV-1, exhibiting significantly faster viral gene expression kinetics compared to epithelial-derived HeLa cells. In contrast, HeLa cells are human cervical cancer epithelial cells with relatively low efficiency in initial transcriptional activation of HSV-1 and higher baseline expression levels of endogenous antiviral factors (such as interferon-stimulated genes), which may lead to a marked delay in the expression of early immediate-early genes like ICP4.

      Author response image 1.

      (3) In responses 2-2, Fig 5K: An infection without transfection has not been included. This is important to understand kinetics of infection in transfected cells.

      We sincerely appreciate the reviewer's insightful identification of this critical oversight. In all relevant experiments, we have strictly included empty vector transfection controls—serving as a baseline reference for each transfection group to eliminate potential influences from the transfection procedure itself and the vector background on viral replication, gene expression, and signaling pathways. The failure to clearly label this control in previous figure legends and main figures was indeed an omission in our presentation; we have now fully addressed this in the revised manuscript: all figures involving transfections (including Figures 3L, 3M, 5K, etc.) now explicitly indicate the "empty vector" control, and we have added detailed explanations in the figure legends and methods section regarding its role as an internal transfection control and procedural comparator. Once again, we thank the reviewer for their high level of professionalism in helping us enhance the completeness and scientific rigor of our data presentation.

      (4) Why HDAC1 with deleted NES does not accumulate or looks like it is degraded? Why then ICP4 does not accumulate?

      We sincerely apologize for the lack of clear labeling of the FLAG-HDAC1 ΔNES protein band in Author response image 2. This omission may have led reviewers to misinterpret its expression level as abnormal. After re-evaluation and improved annotation, Author response image 2 now clearly indicates the FLAG-HDAC1 ΔNES band its migration position corresponds to the expected molecular weight (slightly smaller than wild-type FLAG-HDAC1), and the band intensity is comparable to that of the empty vector and wild-type groups, indicating stable intracellular expression of this mutant protein without significant degradation. Therefore, its inhibitory effect on HSV-1 replication is not due to protein instability, but rather results from subcellular localization defects caused by the loss of nuclear export signal (NES): the ΔNES mutation causes HDAC1 to abnormally retain within the nucleus, ultimately leading to significant downregulation of ICP4 transcription and impaired protein accumulation.

      Author response image 2.

      Reviewer #2 (Public review):

      Summary:

      The authors discovered that HDAC1/2 are degraded in HSV-1 and PRV infections. They attempted to establish a new mechanism by which HDAC1/2 are translocated to the cytoplasm to be degraded in HSV-1 infection, and the degradation causes changes in histone acetylation to affect the DDR pathway.

      Strengths:

      (1) Interesting findings of HDAC1/2 degradation during HSV-1 and PRV infection, and it may impact more than the virology field.

      (2) Significant work to identify the ubiquitin site in HDAC1/2 and K63 linkage.

      Comments on revised version:

      The authors added experiments to address the previous comments. The added knockdown and overexpression experiments provided sufficient support for the proposed mechanism. The conclusions are now strengthened. However, a few essential controls are still missing.

      (1) Figure 3K: How does the expression level of Flag-HDAC1 variants compare to the endogenous HDAC1 level? The stripe probed by Flag antibody should be reprobed by HDAC1 antibody. Also, how does the K74R mutant affect histone acetylation? Moreover, the numbers between the panels are hard to read and have not been explained.

      We sincerely thank the reviewers for their insightful and constructive feedback. In response to the comment on Figure 3K, we performed antibody re-probing of the Flag-immunoprecipitated or Flag-immunoblotted membranes with a validated HDAC1-specific antibody. Consistent with robust transfection and expression, both wild-type Flag-HDAC1 and its mutants including K74R exhibited markedly elevated total HDAC1 protein levels relative to vector control, confirming efficient exogenous expression and protein stability. To directly assess functional consequences, we evaluated global histone acetylation status in parallel samples and found that the K74R mutant induces significantly greater deacetylation than wild-type Flag-HDAC1, as demonstrated by pronounced reductions in H3K56ac and H4K8 acetylation levels. Finally, to improve clarity and readability, we have revised the lane annotations in Figure 3K—increasing font size, enhancing contrast, and ensuring consistent alignment—and fully documented these modifications in the updated figure legend.

      (2) Figure 3M and 3L: DNA transfection per se frequently stimulates cell reactions that inhibit HSV-1 replication. Is the HSV-1 only sample transfected by empty vector or untransfected?

      We sincerely appreciate the reviewer's insightful identification of this critical oversight. In all relevant experiments, we have strictly included empty vector transfection controls serving as a baseline reference for each transfection group to eliminate potential influences from the transfection procedure itself and the vector background on viral replication, gene expression, and signaling pathways. The failure to clearly label this control in previous figure legends and main figures was indeed an omission in our presentation; we have now fully addressed this in the revised manuscript: all figures involving transfections (including Figures 3L, 3M, 5K, etc.) now explicitly indicate the "empty vector" control, and we have added detailed explanations in the figure legends and methods section regarding its role as an internal transfection control and procedural comparator. Once again, we thank the reviewer for their high level of professionalism in helping us enhance the completeness and scientific rigor of our data presentation.

      (3) Figure 4G-4J: What is the MDM2 knockdown efficiency?

      During the construction of the MDM2 knockdown cell lines, we first systematically validated the knockdown efficiency by qRT-PCR. As shown in Figure 4A, compared to the control group (shCtrl), MDM2 mRNA levels were reduced by approximately 60% in shMDM2 cells, and protein expression also showed a corresponding significant decrease, confirming that the cell line had been successfully established and exhibited stable gene silencing effects.

      (4) Figure 5F and line 400-401: "thereby preventing HDAC1 degradation-markedly impaired HSV-1 replication (Fig. 5F)." However, viral replication is not demonstrated in Figure 5F.

      We sincerely appreciate the reviewer for pointing out the error in the figure legend numbering. Upon verification, the experimental data referred to in lines 400–401 of the original text and in Figure 5F actually correspond to the revised new Figure 5J. We apologize for failing to update the figure references in the main text during the revision process due to an oversight. We have now uniformly corrected all relevant descriptions in the text to "Figure 5J" and conducted a comprehensive review of all figure numbers, table numbers, and cross-references throughout the manuscript to confirm there are no other similar errors.

      (5) Figure 5K: also need a control of empty vector. Furthermore, how does the HDAC1 ΔNES expression affect histone acetylation and DDR responses?

      We sincerely thank the reviewers for their thoughtful and constructive feedback on Figure 5K. With regard to the empty vector control: all pertinent experiments in this study were performed with rigorous inclusion of an appropriate empty vector control (pCMV-Flag or its isogenic backbone), serving as the definitive negative control. The prior absence of this control in the figure representation was unintentional and reflects an oversight in data presentation—not in experimental design—and we offer our sincere apologies. We have now incorporated the empty vector control bands into Figure 5K and revised the figure legend to explicitly identify and describe this control. In addition, per the reviewers’ recommendation, we conducted a comprehensive assessment of HDAC1 ΔNES function, specifically examining its impact on global histone acetylation and canonical DNA damage response (DDR) activation. Quantitative immunoblotting and immunofluorescence analyses revealed that HDAC1 ΔNES expression leads to significantly greater reduction in H3K56ac and H4K8 acetylation compared with wild-type HDAC1. Moreover, upon induction of DNA damage, HDAC1 ΔNES-expressing cells exhibit attenuated DDR signaling, evidenced by diminished γH2AX focus formation, reduced CHK2 phosphorylation (p-CHK2), and blunted p53 stabilization and activation consistent with impaired DDR initiation or propagation (see Author response image 3). Collectively, these data indicate that nuclear retention of HDAC1 due to NES deletion not only potentiates its chromatin-targeted deacetylase activity but also contributes to suppression of DDR signaling, likely through epigenetic modulation of damage-sensing chromatin domains.

      Author response image 3.

      (6) Statements listed below are better moved to discussion after all data being presented. They are quite a stretch when looking at each figure by itself.

      (i) Line 268-270: "Together, these findings indicate that HSV-1 selectively degrades class I HDACs, resulting in widespread histone hyperacetylation that fosters a chromatin state conducive to viral replication". ----may be okay for a statement.

      (ii) Line 291-292: "providing initial evidence that HSV-1 infection promotes DDR activation through downregulation of HDAC1 expression"

      (iii) Line 331-333: "Together, these results indicate that HSV-1 infection promotes K63-linked polyubiquitination of HDAC1/2 at conserved lysine residues, ultimately leading to their proteasomal degradation."

      (iv) Line 334-336 is a repeated sentence.

      We sincerely thank the reviewers for their thoughtful and constructive feedback. As noted, statements of mechanistic interpretation are not appropriate in the Results section; accordingly, we have relocated all such statements to the Discussion section. Furthermore, we have conducted a comprehensive line-by-line review of the manuscript to ensure that (i) every mechanistic inference is directly supported by experimental data presented in the Results, and (ii) integrative interpretations particularly those linking molecular observations to broader biological implications are confined exclusively to the Discussion.